Hardware Failure Analysis Engineer - Memphis

Posted 27 Days Ago
Be an Early Applicant
2 Locations
In-Office
Junior
Information Technology
The Role
Analyze firmware and hardware releases for compatibility, security, performance, reliability, and safety. Diagnose intermittent and complex hardware failures, manage vendor RMA claims, collaborate with data center technicians, develop monitoring automation, document reliability metrics and root-cause analyses, and participate in on-call incident response for data center hardware issues.
Summary Generated by Built In

SpaceXAI’s mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company’s mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates.

ABOUT THE ROLE:

Determine true root cause of fleet hardware failures — component and system level — and drive fixes through vendors to the manufacturer. Turn "swap it again" into "vendor redesign / firmware fix / manufacturing escape found."

RESPONSIBILITIES:
  • Own named failure classes across GPU trays/baseboards, NIC/DPU, motherboard/PCIe switch, memory, power, thermal, cables/connectors, and rack-scale patterns.
  • Run recurrence analysis and fleet-wide defect clustering; detect systemic patterns before they become fleet-scale loss.
  • Build vendor technical escalation packages with evidence quality that forces action; partner with OEM/ODM/component manufacturers on firmware, bring-up, and manufacturing escapes through committed CAPA.
  • Feed findings into RMA policy, spare strategy, and "do not reseat forever" stop-rules.
  • Work the FA intake queue on rotation; keep queue age within SLA
BASIC QUALIFICATIONS:
  • Bachelor's degree in Systems Engineering, Electrical Engineering, Computer Science, or a related field (or equivalent experience).
  • 2+ years of experience in hardware reliability engineering, preferably in high-performance computing or data center environments.
  • Proven expertise in firmware analysis, hardware specifications review, and release validation.
  • Strong experience with RMA processes, including filing claims, vendor negotiations, and pushing for resolutions outside standard protocols.
  • Demonstrated ability to diagnose and prove complex hardware failures, including grey or intermittent issues, using tools, logic analyzers, or diagnostic software.
  • Familiarity with data center hardware components (e.g., servers, GPUs, networking equipment) and emerging technologies.
  • Proficiency in scripting (Python, Bash) for automation and analysis, plus general experience in at least one systems language (C, C++, Java, Rust, or similar). Not required to be expert in all of them.
  • Excellent problem-solving skills with a data-driven approach to reliability engineering.
  • Ability to work collaboratively with cross-functional teams, including operations technicians.
PREFERRED SKILLS AND EXPERIENCE:
  • Experience in AI/ML infrastructure or supercomputing environments.
  • Knowledge of vendor ecosystems (e.g., NVIDIA, Dell, HP, Supermicro) and supply chain management.
  • Certifications in hardware engineering or reliability (e.g., CRE, CompTIA Server+).
  • Prior work in a fast-paced startup or tech company like SpaceXAI.

SpaceXAI is an equal opportunity employer. For details on data processing, view our Recruitment Privacy Notice.

Skills Required

  • Bachelor's degree in Systems Engineering, Electrical Engineering, Computer Science, a related field, or equivalent experience
  • 2+ years of experience in hardware reliability engineering, preferably in high-performance computing or data center environments
  • Expertise in firmware analysis, hardware specification review, and release validation
  • Experience with RMA processes, vendor negotiations, and resolution escalation
  • Ability to diagnose and prove complex, intermittent, or grey hardware failures using testing tools, logic analyzers, or diagnostic software
  • Familiarity with data center hardware, including servers, GPUs, and networking equipment
  • Proficiency in Python and Bash scripting, plus general experience with a systems language such as C, C++, Java, or Rust
  • Strong data-driven problem-solving and reliability engineering skills
  • Ability to collaborate with cross-functional teams and data center operations technicians
  • Experience in AI/ML infrastructure or supercomputing environments
  • Knowledge of vendor ecosystems such as NVIDIA, Dell, HP, or Supermicro and supply chain management
  • Hardware engineering or reliability certification, such as CRE or CompTIA Server+
  • Experience at a fast-paced startup or technology company
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Palo Alto, CA
96 Employees

What We Do

Understand the Universe

Similar Jobs

CrowdStrike Logo CrowdStrike

Consultant

Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity
Remote or Hybrid
USA
11000 Employees
110K-160K Annually

CrowdStrike Logo CrowdStrike

Consultant

Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity
Remote or Hybrid
USA
11000 Employees
110K-160K Annually

PNC Bank Logo PNC Bank

Technology Solution Center Analyst Lead

Machine Learning • Payments • Security • Software • Financial Services
Remote or Hybrid
USA
55000 Employees
41K-83K Annually

Enverus Logo Enverus

Consultant

Big Data • Information Technology • Software • Analytics • Energy
In-Office or Remote
5 Locations
1800 Employees
95K-110K Annually

Similar Companies Hiring

Axle Health Thumbnail
Artificial Intelligence • Healthtech • Information Technology • Logistics
Santa Monica, CA
25 Employees
NODA AI Thumbnail
Artificial Intelligence • Information Technology • Software • Cybersecurity
Sydney, AU
54 Employees
Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account