Principal System Debug Engineer

Posted Yesterday
Be an Early Applicant
2 Locations
Hybrid
Expert/Leader
Artificial Intelligence • Semiconductor
Joining Graphcore gives you a seat at the top-table, shaping the future of Artificial Intelligence.
The Role
Lead system-level debug and validation for Arm-based server blades and rack platforms. Drive post-silicon bring-up, cross-functional root-cause analysis, debug methodologies, tooling and automation, program metrics, and mentor engineers to ensure POR quality and timely issue resolution.
Summary Generated by Built In
Lead System Debug Engineer – Server & Rack Validation (Principal Level and Above)Position Overview

We are seeking a senior technical leader (Principal Engineer level and above) to lead the bring-up, enablement, and hardware debug of server compute systems and rack-level platforms based on Arm® server architecture.

The successful candidate will be a key member of the System Validation organization, responsible for driving system-level debug activities and facilitating rapid resolution of complex hardware, firmware, and software issues. This role requires close collaboration with engineering teams across the organization to identify root causes, implement corrective actions, and ensure successful program execution.

The ideal candidate will be deeply involved in challenging system debug efforts while developing and executing scalable debug strategies that maximize throughput and ensure Product of Record (POR) quality. In addition, this individual will establish and drive debug methodologies, improve processes, and help create a culture of technical excellence across the organization.

We are looking for a disciplined, dynamic, and highly motivated leader who can thrive in a global environment while fostering strong cross-functional collaboration.

As a Lead System Debug Engineer within Server and Rack Validation, you will drive balanced, scalable, and automated debug solutions that optimize engineering efficiency and product quality. This highly visible role provides the opportunity to innovate and improve debugging capabilities while delivering industry-leading server technologies to market.

Your technical leadership, validation expertise, and problem-solving skills will play a critical role in product development, issue root cause analysis, and resolution. Success in this role requires close collaboration with System Validation, System Architecture, Silicon Engineering, Rack Firmware, and other cross-functional teams.

Primary Responsibilities
  • Develop and drive a Debug Center of Excellence, including scalable debug and triage methodologies, processes, and playbooks for server blade and rack-level issues spanning hardware, firmware, and software integration.

  • Debug issues discovered during server rack bring-up, post-silicon validation, and production phases.

  • Lead complex debug efforts involving silicon, server systems, firmware, and software to determine root causes and drive effective resolutions.

  • Ensure issues are resolved with high quality and within program timelines.

  • Manage and track technical issues, risks, and priorities to remove blockers and achieve key program milestones.

  • Develop and publish debug program metrics and indicators to identify roadblocks and improve overall debug efficiency.

  • Communicate program status, risks, and opportunities to customers, stakeholders, and executive leadership.

  • Drive technical innovation across triage and debug workflows through tool development, scripting, methodology enhancements, and cross-functional engineering initiatives.

  • Mentor engineers and promote best practices in system validation and debug methodologies.

Required Qualifications
  • Strong analytical and problem-solving skills with exceptional attention to detail.

  • Extensive experience in validation and debug roles involving operating systems, firmware, silicon, and hardware issues.

  • Deep understanding of industry-standard server interconnects and software stacks, including PCIe and CXL.

  • Strong knowledge of Arm® CPU or x86 architectures, SoC design, memory subsystems, RAS (Reliability, Availability, and Serviceability), and power management.

  • Extensive experience with system architecture, technical debugging, and validation strategies.

  • Strong understanding of platform-level and system-level debug methodologies, including Operating Systems, Device Drivers, and BIOS interactions.

  • Excellent communication, collaboration, and cross-functional leadership skills.

  • Highly organized and detail-oriented, with the ability to manage multiple priorities and deliver results under tight deadlines.

  • Experience leading technical programs and coordinating cross-functional engineering efforts.

  • Thorough understanding of data center technologies and associated software stacks.

  • Self-motivated with the ability to independently drive tasks from problem identification through resolution.

Preferred Qualifications
  • Master's degree or Ph.D. in Electrical Engineering, Computer Engineering, Computer Science, or a related technical field.

  • Experience with large-scale server platforms, rack-level systems, and hyperscale data center environments.

  • Expertise in automation, scripting, and debug tool development.

  • Experience establishing and scaling debug processes across multiple product generations and engineering organizations.

Skills Required

  • Strong analytical and problem-solving skills with exceptional attention to detail
  • Extensive experience in validation and debug roles involving operating systems, firmware, silicon, and hardware issues
  • Deep understanding of industry-standard server interconnects and software stacks, including PCIe and CXL
  • Strong knowledge of Arm CPU or x86 architectures, SoC design, memory subsystems, RAS, and power management
  • Extensive experience with system architecture, technical debugging, and validation strategies
  • Strong understanding of platform-level and system-level debug methodologies, including Operating Systems, Device Drivers, and BIOS interactions
  • Excellent communication, collaboration, and cross-functional leadership skills
  • Highly organized and detail-oriented, with ability to manage multiple priorities under tight deadlines
  • Experience leading technical programs and coordinating cross-functional engineering efforts
  • Thorough understanding of data center technologies and associated software stacks
  • Self-motivated with ability to independently drive tasks from problem identification through resolution
  • Master's degree or Ph.D. in EE, Computer Engineering, Computer Science, or related field
  • Experience with large-scale server platforms, rack-level systems, and hyperscale data center environments
  • Expertise in automation, scripting, and debug tool development
  • Experience establishing and scaling debug processes across multiple product generations and organizations

What the Team is Saying

Monika
Dionysia
Dave

Graphcore Compensation & Benefits Highlights

  • Healthcare Strength Health coverage includes medical and dental insurance, with US plans through Cigna and Kaiser, HDHP options with employer‑funded HSA contributions, a health cash plan, EAP access, and dedicated mental‑health support. These provisions extend to family options in some regions, reinforcing broad medical and wellbeing support.
  • Retirement Support Retirement programs include a UK pension match up to 5% and a US 401(k) with a 100% company match up to 6% (with a true‑up). This pairing signals strong, predictable long‑term savings support across key locations.
  • Leave & Time Off Breadth Time‑off policies feature “unlimited” holiday in the UK and flexible, generous PTO with paid US holidays. Paid family leave for birthing parents and bonding further broadens time‑away support.

Graphcore Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Bristol
762 Employees
Year Founded: 2016

What We Do

At Graphcore, we’re building the future of AI compute. We’re a team of semiconductor, software and AI experts, with deep experience in creating the complete AI compute stack - from silicon and software to infrastructure at datacenter scale. As part of the SoftBank Group, backed by significant long-term investment, we are delivering key technology into the fast-growing SoftBank AI ecosystem. To meet the vast and exciting AI opportunity, Graphcore is expanding its teams around the world. We are bringing together the brightest minds to solve the toughest problems, in a place where everyone has the opportunity to make an impact on the company, our products and the future of artificial intelligence.

Why Work With Us

Our team is at the forefront of the machine intelligence revolution, enabling innovators from all industries to build AI-native products to expand human potential. What we do at Graphcore really makes a difference.

Gallery

Gallery
Gallery
Gallery
Gallery
Gallery
Gallery
Gallery
Gallery
Gallery
Gallery

Graphcore Offices

Hybrid Workspace

Employees engage in a combination of remote and on-site work.

At Graphcore, we value wellbeing and flexibility to support a healthy work/life balance. Our hybrid approach encourages office-based colleagues to work onsite three days a week, with trusted flexibility built on trust and transparency for everyone.

Typical time on-site: 3 days a week
HQHeadquarters
Austin Office
Bengaluru Office
Cambridge Office
Gdańsk Office
Hsinchu Office
London Office
Learn more

Similar Jobs

Graphcore Logo Graphcore

Technical Program Manager

Artificial Intelligence • Semiconductor
Hybrid
Austin, TX, USA
762 Employees

Graphcore Logo Graphcore

Staff Mechanical HVAC Engineer

Artificial Intelligence • Semiconductor
Remote or Hybrid
2 Locations
762 Employees

Graphcore Logo Graphcore

Senior Systems Engineer

Artificial Intelligence • Semiconductor
Hybrid
Austin, TX, USA
762 Employees

Graphcore Logo Graphcore

Robotics Engineer

Artificial Intelligence • Semiconductor
Hybrid
Austin, TX, USA
762 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account