Principal Storage Architect

Posted Yesterday
Be an Early Applicant
Milpitas, CA, USA
Hybrid
Senior level
Artificial Intelligence • Semiconductor
Joining Graphcore gives you a seat at the top-table, shaping the future of Artificial Intelligence.
The Role
Defines end-to-end storage architecture and technology roadmaps for AI servers and data centers. Leads design of local, disaggregated, and distributed storage; optimizes Linux storage performance, NVMe lifecycle management, telemetry, and AI workloads. Diagnoses complex hardware and software issues across kernels, PCIe, networks, firmware, and storage platforms. Provides technical leadership across engineering, automation, supply chain, vendors, and product teams while evaluating emerging interconnect and memory-tiering technologies.
Summary Generated by Built In
About Graphcore

Graphcore is a leading innovator in artificial intelligence computing. We develop hardware, software, and data center infrastructure that provide the specialized processing and systems capabilities needed to advance AI while improving the efficiency required for broad adoption.

As part of SoftBank Group, Graphcore works alongside companies developing advanced technologies. Our teams bring together AI researchers, silicon designers, hardware and software engineers, and systems architects to solve complex technical problems across the computing stack.

The Opportunity

As Principal Storage Architect, you will define the high-performance storage architecture for Graphcore's AI computing and data center infrastructures.   You will lead the design of local and distributed storage tiers that provide the throughput, availability, and predictable tail latency required for large-scale training and inference.

Working within Advanced Architecture, you will address storage requirements for saving model states, offloading key-value caches, large datasets, and high-speed data-loading pipelines.   You will combine hands-on performance engineering with Principal-level technical leadership across hardware, software, networking, systems engineering, automation, supply chain, and external technology partners.

What You Will Do
  • Define the end-to-end architecture and technology roadmap for local, disaggregated, and distributed storage across Graphcore AI server and data center platforms.
  • Design storage topologies that optimize PCIe lane allocation, network-domain placement, and data paths among CPUs, AI accelerators, memory, and storage systems.
  • Lead the architecture of storage control-plane and data-plane solutions, evaluating commercial and open technologies against performance, resilience, manageability, and lifecycle requirements.
  • Profile and tune the Linux storage stack, block layer, I/O schedulers, direct I/O paths, file systems, and drivers to improve IOPS, bandwidth, and 99.99th-percentile latency.
  • Optimize storage for AI workloads, including model checkpointing, key-value cache offload, data ingestion, and direct data movement between NVMe storage and accelerator memory.
  • Set the technical direction for NVMe SSD lifecycle management, including qualification, provisioning, health monitoring, firmware rollout, failure handling, and warranty-return automation across E1.S, E3.S, U.2, and U.3 devices.
  • Define telemetry and alerting requirements for direct-attached and distributed storage, including endurance, wear, drive writes per day, thermals, capacity, performance, and latency anomalies.
  • Provide architecture requirements and technical guidance to automation teams building frameworks that characterize storage performance, reliability, and compatibility across AI platforms.
  • Lead root-cause analysis for complex storage failures and performance degradation across Linux kernels, drivers, PCIe, networks, SSD firmware, and third-party storage systems.
  • Partner with storage vendors, supply chain, networking, systems engineering, and product teams to select, integrate, and deploy storage components and platforms.
  • Evaluate emerging storage, interconnect, and memory-tiering technologies and translate relevant developments into platform requirements and future system designs.
What You Will Bring
  • A bachelor's or equivalent experience or master's degree in computer science, computer engineering, information technology, electrical engineering, or a related field, or equivalent practical experience.
  • Extensive hands-on experience in storage engineering and architecture for AI, high-performance computing, hyperscale, or large data center environments, including technical leadership at Principal, Staff, or Lead Architect scope.
  • Deep knowledge of NVMe, PCIe Gen5 or Gen6 architecture, NAND flash behavior, SSD form factors, endurance, firmware, and device lifecycle management.
  • Deep operational knowledge of Linux storage internals, including the block layer, I/O schedulers, direct I/O, kernel and driver behavior, and file systems such as XFS, ext4, or ZFS.
  • Hands-on experience with storage benchmarking and profiling tools such as fio, blktrace, iostat, and perf, including the ability to design representative workload models and isolate bottlenecks.
  • Programming and automation skills in Python and Bash, with experience using structured data formats and REST APIs to build monitoring, validation, or lifecycle workflows.
  • Demonstrated ability to diagnose complex hardware and software interactions, including kernel failures, PCIe errors, network issues, performance regressions, and SSD firmware defects.
  • Experience defining telemetry signals, alert thresholds, dashboards, and operational response criteria for storage health and performance at fleet scale.
  • Proven ability to lead cross-functional architecture decisions, influence senior engineers and leaders, and communicate technical risks, tradeoffs, and recommendations clearly.
Preferred Qualifications

These qualifications are helpful, not required. We encourage you to apply even if you do not meet every preferred qualification.

  • Experience implementing, testing, or troubleshooting GPU Direct Storage or comparable direct accelerator-to-storage technologies.
  • Experience designing or operating NVMe over Fabrics deployments using RoCEv2 or TCP.
  • Experience with the Storage Performance Development Kit or other user-space storage frameworks.
  • Familiarity with high-performance parallel and distributed file systems such as Weka, VAST Data, Lustre, DAOS, DDN, or comparable platforms.
  • Knowledge of Compute Express Link and its implications for future memory and storage tiering in AI servers.
United States Benefits Overview

Graphcore offers compensation and benefits designed to support employees' health, financial well-being, work-life needs, and professional growth. Benefits and programs for eligible U.S. employees may include:

  • Medical, dental, and vision coverage, with options that may extend to eligible dependents.
  • Mental health, wellness, and employee assistance resources.
  • Retirement savings benefits and company contributions where applicable.
  • Paid vacation, sick time, company holidays, and parental or family leave in accordance with applicable plans and policies.
  • Life insurance and short-term or long-term disability coverage.
  • Flexible working hours and hybrid working arrangements where compatible with the role and team requirements.
  • Professional development resources, learning programs, office amenities, and team-led activities.

Benefits vary by work location, employment status, scheduled hours, and plan eligibility and are subject to the terms of the applicable plans and company policies. This overview is not a contract or guarantee of benefits.

Equal Opportunity and Accommodations

Graphcore is an equal opportunity employer. We consider qualified applicants without regard to race, color, religion, creed, sex, pregnancy, sexual orientation, gender identity or expression, national origin, ancestry, age, disability, genetic information, veteran status, or any other status protected by applicable law.

Graphcore is committed to an inclusive and accessible hiring process. If you need a reasonable accommodation to participate in the application or interview process, please let the recruiting team know.

Candidate Privacy

Personal information submitted during the recruiting process will be handled in accordance with Graphcore's applicable candidate privacy notices.

Skills Required

  • Bachelor's degree or equivalent experience, or master's degree, in computer science, computer engineering, information technology, electrical engineering, or a related field
  • Extensive hands-on storage engineering and architecture experience in AI, high-performance computing, hyperscale, or large data center environments
  • Principal, Staff, or Lead Architect-level technical leadership experience
  • Deep knowledge of NVMe, PCIe Gen5 or Gen6, NAND flash, SSD form factors, endurance, firmware, and device lifecycle management
  • Operational knowledge of Linux storage internals, block layers, I/O schedulers, direct I/O, kernel and driver behavior, and file systems such as XFS, ext4, or ZFS
  • Hands-on experience with fio, blktrace, iostat, perf, storage benchmarking, and performance profiling
  • Programming and automation skills in Python and Bash
  • Experience using structured data formats and REST APIs for monitoring, validation, or lifecycle workflows
  • Ability to diagnose kernel failures, PCIe errors, network issues, performance regressions, and SSD firmware defects
  • Experience defining telemetry signals, alert thresholds, dashboards, and operational response criteria at fleet scale
  • Ability to lead cross-functional architecture decisions and communicate technical risks, tradeoffs, and recommendations
  • Experience with GPU Direct Storage or comparable direct accelerator-to-storage technologies
  • Experience designing or operating NVMe over Fabrics using RoCEv2 or TCP
  • Experience with the Storage Performance Development Kit or other user-space storage frameworks
  • Familiarity with Weka, VAST Data, Lustre, DAOS, DDN, or comparable parallel and distributed file systems
  • Knowledge of Compute Express Link and memory and storage tiering

What the Team is Saying

Monika
Dionysia
Dave

Graphcore Compensation & Benefits Highlights

  • Healthcare Strength — Company materials list private medical insurance, dental coverage, health cash plans, mental‑health support, and day‑one U.S. medical options via Cigna/Kaiser. Feedback suggests these offerings are a meaningful strength, particularly in the UK.
  • Leave & Time Off Breadth — Benefits highlight flexible working hours, unlimited annual leave, generous parental leave, and paid U.S. holidays. Feedback suggests work‑life balance is often described positively alongside these policies.
  • Retirement Support — UK information cites pension matching up to 5%, while U.S. listings describe a 401(k) with a 100% company match up to 6%. Feedback suggests these structured contributions are viewed as competitive elements of the package.

Graphcore Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Bristol, Bristol
903 Employees
Year Founded: 2016

What We Do

At Graphcore, we’re building the future of AI compute. We’re a team of semiconductor, software and AI experts, with deep experience in creating the complete AI compute stack - from silicon and software to infrastructure at datacenter scale. As part of the SoftBank Group, backed by significant long-term investment, we are delivering key technology into the fast-growing SoftBank AI ecosystem. To meet the vast and exciting AI opportunity, Graphcore is expanding its teams around the world. We are bringing together the brightest minds to solve the toughest problems, in a place where everyone has the opportunity to make an impact on the company, our products and the future of artificial intelligence.

Why Work With Us

Our team is at the forefront of the machine intelligence revolution, enabling innovators from all industries to build AI-native products to expand human potential. What we do at Graphcore really makes a difference.

Gallery

Gallery
Gallery
Gallery
Gallery
Gallery
Gallery
Gallery
Gallery
Gallery
Gallery

Graphcore Offices

Hybrid Workspace

Employees engage in a combination of remote and on-site work.

At Graphcore, we value wellbeing and flexibility to support a healthy work/life balance. Our hybrid approach encourages office-based colleagues to work onsite three days a week, with trusted flexibility built on trust and transparency for everyone.

Typical time on-site: 3 days a week
HQHeadquarters
Austin Office
Bengaluru Office
Cambridge Office
Gdańsk Office
Hsinchu Office
London Office
Learn more

Similar Jobs

Graphcore Logo Graphcore

Principal Datacenter Technologist

Artificial Intelligence • Semiconductor
Hybrid
Milpitas, CA, USA
903 Employees

Graphcore Logo Graphcore

Architect

Artificial Intelligence • Semiconductor
Hybrid
2 Locations
903 Employees

Graphcore Logo Graphcore

Director, Sustaining Engineering

Artificial Intelligence • Semiconductor
Hybrid
Milpitas, CA, USA
903 Employees
50K-50K Annually

Graphcore Logo Graphcore

Staff Power Engineer

Artificial Intelligence • Semiconductor
Hybrid
Milpitas, CA, USA
903 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account