Principal Software Engineer, AI Compute Infrastructure

Posted Yesterday
Be an Early Applicant
Seattle, WA, USA
Hybrid
263K-355K Annually
Senior level
Artificial Intelligence • Internet of Things • Semiconductor
Build What the World Depends On
The Role
Designs, builds, and operates large-scale AI compute infrastructure for training, fine-tuning, evaluation, and inference. Responsibilities include managing Kubernetes clusters, enabling CPU and GPU systems, optimizing scheduling and capacity, troubleshooting distributed infrastructure, and automating provisioning, upgrades, monitoring, and maintenance. The role partners closely with AI researchers and engineers to improve infrastructure reliability, scalability, performance, and developer productivity.
Summary Generated by Built In
As a Principal Engineer on the AI Compute Infra team, you will design, build, and operate large-scale infrastructure for AI training, fine-tuning, evaluation, and inference. You will guide work across Kubernetes clusters, accelerator enablement, workload scheduling, high-performance networking, storage, and capacity management, partnering with AI researchers and engineers to improve reliability, performance, scalability, and developer productivity.
Responsibilities:
  • Build and operate Kubernetes clusters while improving workload scheduling, topology-aware placement, capacity use, and recovery.
  • Enable new CPU and GPU systems by integrating and validating drivers, networking, storage, monitoring, and health checks.
  • Investigate performance and reliability issues across applications, cloud infrastructure, clusters, and hardware, then turn findings into lasting improvements.
  • Partner with AI teams to understand their workloads and automate cluster provisioning, upgrades, monitoring, and maintenance around their needs.

Necessary Skills and Experience:
  • 8+ years of experience building or operating cloud, compute, HPC, or distributed infrastructure in a production environment.
  • Programming experience in Go, Python, or another systems language, with an interest in developing reliable infrastructure software.
  • Practical knowledge of Kubernetes, containers, Linux, networking, and storage.
  • Experience supporting GPU, accelerator, or distributed machine-learning workloads.
  • An ability to troubleshoot complex systems and communicate clearly with engineers from different technical backgrounds.

"Preferred" Skills and Experience:
  • Familiarity with Kubernetes scheduling, operators, quotas, or resource management.
  • Experience with NVIDIA technologies such as CUDA, NVLink, NVSwitch, NCCL, EFA, or DCGM.
  • Knowledge of AWS EKS, Terraform, Argo CD, Helm, Prometheus, or Grafana.
  • Familiarity with frameworks such as PyTorch, Ray, vLLM, SGLang, or TensorRT-LLM, or experience qualifying accelerators and tuning distributed workloads.

In Return:
You will join a driven group committed to developing world-class AI compute infrastructure. We provide a cooperative setting where your ideas can come to life. Your efforts will directly impact the success of our AI projects, guaranteeing smooth operations and outstanding results. Join us and help build the future of AI compute infrastructure!
Additional Information
Please note that a relocation package (including visa sponsorship support) is available for this role, for candidates who require it.
We are an equal opportunity employer and value diversity at our company. We do not discriminate on the basis of race, religion, color, national origin, sex, gender, gender expression, sexual orientation, age, marital status, veteran status, or disability status. We will ensure that individuals with disabilities are provided reasonable accommodation to participate in the job application or interview process, to perform essential job functions, and to receive other benefits and privileges of employment. Please contact us to request accommodation.
Salary Range:
$262,700-$355,400 per year
We value people as individuals and our dedication is to reward people competitively and equitably for the work they do and the skills and experience they bring to Arm. Salary is only one component of Arm's offering. The total reward package will be shared with candidates during the recruitment and selection process.
Accommodations at Arm
At Arm, we want to build extraordinary teams. If you need an adjustment or an accommodation during the recruitment process, please email [email protected] . To note, by sending us the requested information, you consent to its use by Arm to arrange for appropriate accommodations. All accommodation or adjustment requests will be treated with confidentiality, and information concerning these requests will only be disclosed as necessary to provide the accommodation. Although this is not an exhaustive list, examples of support include breaks between interviews, having documents read aloud, or office accessibility. Please email us about anything we can do to accommodate you during the recruitment process.
Hybrid Working at Arm
Arm's approach to hybrid working is designed to create a working environment that supports both high performance and personal wellbeing. We believe in bringing people together face to face to enable us to work at pace, whilst recognizing the value of flexibility. Within that framework, we empower groups/teams to determine their own hybrid working patterns, depending on the work and the team's needs. Details of what this means for each role will be shared upon application. In some cases, the flexibility we can offer is limited by local legal, regulatory, tax, or other considerations, and where this is the case, we will collaborate with you to find the best solution. Please talk to us to find out more about what this could look like for you.
Equal Opportunities at Arm
Arm is an equal opportunity employer, committed to providing an environment of mutual respect where equal opportunities are available to all applicants and colleagues. We are a diverse organization of dedicated and innovative individuals, and don't discriminate on the basis of race, color, religion, sex, sexual orientation, gender identity, national origin, disability, or status as a protected veteran.

Skills Required

  • 8+ years of experience building or operating cloud, compute, HPC, or distributed infrastructure in production
  • Programming experience in Go, Python, or another systems language
  • Practical knowledge of Kubernetes, containers, Linux, networking, and storage
  • Experience supporting GPU, accelerator, or distributed machine-learning workloads
  • Ability to troubleshoot complex systems and communicate clearly with engineers from different technical backgrounds
  • Familiarity with Kubernetes scheduling, operators, quotas, or resource management
  • Experience with NVIDIA technologies such as CUDA, NVLink, NVSwitch, NCCL, EFA, or DCGM
  • Knowledge of AWS EKS, Terraform, Argo CD, Helm, Prometheus, or Grafana
  • Familiarity with PyTorch, Ray, vLLM, SGLang, or TensorRT-LLM, or experience qualifying accelerators and tuning distributed workloads

Arm Compensation & Benefits Highlights

  • Healthcare Strength Medical coverage in the U.S. is described as very good, with instances of $0 employee‑only premiums on certain plans, alongside dental, vision, and mental‑health resources. Company materials also highlight globally tailored wellbeing support and Employee Assistance Programs.
  • Retirement Support Arm states a 401(k) match of 100% on the first 6% of pay in the U.S., with retirement or pension plans available in other regions. This creates a clear, predictable savings boost within the core package.
  • Leave & Time Off Breadth Four weeks of vacation in the U.S. plus a fully paid, four‑week sabbatical after four years stand out, with an additional quarterly “Day of Care” for wellbeing. Materials also reference generous PTO, paid holidays, sick time, and bereavement leave.

Arm Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Cambridge, England
8,314 Employees
Year Founded: 1990

What We Do

We bring brilliant people together in a global ecosystem that is sparking the world’s potential. Arm technology enables specialized processing built on the economics, design freedom and accessibility of general-purpose compute that has, so far, led to more than 180 billion chips being shipped by our partners.

Why Work With Us

At Arm, we build the future of computing, powering everything from smartphones to AI. Our 10x mindset drives bold thinking and deep collaboration to solve complex problems together. With a people first culture, flexible work, and strong support for growth and wellbeing, your ideas can make a global impact while your career thrives.

Gallery

Gallery
Gallery
Gallery
Gallery

Arm Offices

Hybrid Workspace

Employees engage in a combination of remote and on-site work.

Typical time on-site: Not Specified
HQCambridge, UK
Galway, Ireland
Budapest, Hungary
Sophia Antipolis, France
Ra'anana, Israel
Bengaluru, India
Noida, India
Yokohama, Japan
Seoul, South Korea
Hsinchu, Taiwan
Taipei, Taiwan
Munich, Germany
Austin, TX
Bristol, UK
Chandler, AZ
Raleigh, NC
Lund, Sweden
Manchester, England
Oslo, Norway
San Diego, CA
San Jose, CA
Sheffield, UK
Trondheim, Norway
Boston, MA
Learn more

Similar Jobs

Arm Logo Arm

Staff Software Engineer

Artificial Intelligence • Internet of Things • Semiconductor
Hybrid
Seattle, WA, USA
8314 Employees
209K-283K Annually

Arm Logo Arm

Principal Software Engineer

Artificial Intelligence • Internet of Things • Semiconductor
Hybrid
Seattle, WA, USA
8314 Employees
263K-355K Annually

Arm Logo Arm

Robotics Engineer

Artificial Intelligence • Internet of Things • Semiconductor
Hybrid
Seattle, WA, USA
8314 Employees
209K-283K Annually

Arm Logo Arm

Staff Software Engineer

Artificial Intelligence • Internet of Things • Semiconductor
Hybrid
Seattle, WA, USA
8314 Employees
209K-283K Annually

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account