The posting, in Graphcore's own words
archived Oct 7, 2026About Graphcore
Graphcore is a leading innovator in artificial intelligence computing. We develop hardware, software, and data center infrastructure that provide the specialized processing and systems capabilities needed to advance AI while improving the efficiency required for broad adoption. As part of SoftBank Group, Graphcore works alongside companies developing advanced technologies. Our teams bring together AI researchers, silicon designers, hardware and software engineers, and systems architects to solve complex technical problems across the computing stack.
The Opportunity
As Principal Storage Architect, you will define the high-performance storage architecture for Graphcore's AI computing and data center infrastructures. You will lead the design of local and distributed storage tiers that provide the throughput, availability, and predictable tail latency required for large-scale training and inference. Working within Advanced Architecture, you will address storage requirements for saving model states, offloading key-value caches, large datasets, and high-speed data-loading pipelines. You will combine hands-on performance engineering with Principal-level technical leadership across hardware, software, networking, systems engineering, automation, supply chain, and external technology partners.
Read the full posting ↓
What You Will Do
Define the end-to-end architecture and technology roadmap for local, disaggregated, and distributed storage across Graphcore AI server and data center platforms. Design storage topologies that optimize PCIe lane allocation, network-domain placement, and data paths among CPUs, AI accelerators, memory, and storage systems. Lead the architecture of storage control-plane and data-plane solutions, evaluating commercial and open technologies against performance, resilience, manageability, and lifecycle requirements. Profile and tune the Linux storage stack, block layer, I/O schedulers, direct I/O paths, file systems, and drivers to improve IOPS, bandwidth, and 99.99th-percentile latency. Optimize storage for AI workloads, including model checkpointing, key-value cache offload, data ingestion, and direct data movement between NVMe storage and accelerator memory. Set the technical direction for NVMe SSD lifecycle management, including qualification, provisioning, health monitoring, firmware rollout, failure handling, and warranty-return automation across E1.S, E3.S, U.2, and U.3 devices. Define telemetry and alerting requirements for direct-attached and distributed storage, including endurance, wear, drive writes per day, thermals, capacity, performance, and latency anomalies. Provide architecture requirements and technical guidance to automation teams building frameworks that characterize storage performance, reliability, and compatibility across AI platforms. Lead root-cause analysis for complex storage failures and performance degradation across Linux kernels, drivers, PCIe, networks, SSD firmware, and third-party storage systems. Partner with storage vendors, supply chain, networking, systems engineering, and product teams to select, integrate, and deploy storage components and platforms. Evaluate emerging storage, interconnect, and memory-tiering technologies and translate relevant developments into platform requirements and future system designs.