Hardware Analytics Engineer

Posted 28 Days Ago
Sunnyvale, CA, USA
Hybrid
214K-225K Annually
Mid level
Artificial Intelligence • Hardware • Software • Semiconductor
The Role
Design and optimize hyperscale data pipelines and ETL for multi-terabyte hardware telemetry. Build ML-driven failure prediction and anomaly detection, create dashboards and real-time telemetry visualizations, lead hardware characterization and A/B studies, perform root cause analysis, and collaborate with cross-functional teams to improve hardware performance, reliability, and energy efficiency for AI platforms.
Summary Generated by Built In

Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference services.
This order of magnitude increase in speed is transforming the user experience of AI applications, unlocking real-time iteration and increasing intelligence via additional agentic computation.
Cerebras works with the leading model labs, global enterprises, and cutting-edge AI-native startups. OpenAI recently announced a multi-year partnership with Cerebras, to deploy 750 megawatts of scale, transforming key workloads with ultra high-speed inference.

Cerebras Systems Inc. has multiple openings for Hardware Analytics Engineer

Title: Hardware Analytics Engineer

Job Duties:

  • Design and optimize scalable data pipeline architectures for multi-terabyte hardware telemetry, reliability analytics, and performance optimization.

  • Architect, develop, and optimize hyperscale data pipeline frameworks and ETL processes to aggregate, process, and analyze multi-terabyte hardware performance and telemetry streams, including utilization, power, thermal, acoustic, and reliability metrics across heterogeneous compute, storage, and AI server platforms, ensuring hardware performance compliance and operational reliability.

  • Design and implement hardware performance analysis and anomaly detection systems using Python, SQL, Tableau, Hive, and Spark to forecast hardware failure curves, identify performance bottlenecks, and generate prescriptive recommendations for hardware and system optimization.

  • Lead hardware characterization experiments and thermal/cooling A/B studies to evaluate operational envelopes, delivering validated strategies that reduce carbon footprint, improve water usage efficiency, and maintain or enhance system reliability.

  • Engineer telemetry ingestion, monitoring, and visualization systems to provide real-time, high-fidelity hardware health data to hardware, firmware, and datacenter operations teams, enabling data-driven decision-making at scale.

  • Define, operationalize, and maintain custom efficiency and reliability metrics; perform root cause analysis of systemic failures using large-scale statistical and machine learning methods; and deploy solutions that improve platform scalability, energy efficiency, and sustainability.

  • Collaborate with cross-functional engineering teams to troubleshoot complex failures, isolate defective components, and implement systemic fixes across CPU, GPU, DRAM, PCIe, networking, and storage subsystems.

  • Support the evolution and optimization of next-generation AI platforms and silicon products, including hardware subsystems (CPU, GPU, DRAM, PCIe, networking, and storage), to meet the performance, scalability, and efficiency demands of large language model training and inference workloads.

Minimum Requirements:

Master’s degree or foreign equivalent degree in Electrical Engineering, Computer Engineering, Computer Science, or a related field and 3 years of experience as Hardware Analytics Engineer, Hardware Engineer, Data Engineer, or a related occupation required.

Required Skills:

  • Large-scale data pipeline architecture and ETL, distributed data processing (Hive, Spark), and dashboard development;

  • Python, SQL, Tableau, Linux, and automation scripting;

  • Design, training, and deployment of machine learning models for hardware performance optimization and failure prediction;

  • Predictive modeling, statistical analysis, A/B testing, anomaly detection, and data visualization in hardware reliability and performance; and

  • Hardware analytics for compute, storage, and AI servers; power and thermal optimization; GPU burn-in efficiency optimization; and reliability modeling for AI hardware systems and components including CPU, GPU, DRAM, and SSD.

 

Additional Information:

Employer’s name: Cerebras Systems Inc.

Job site : 1237 E Arques Avenue, Sunnyvale, CA 94085

Telecommuting permitted

Salary Range: $213,675.00 per year to $225,000.00 per year

If you are interested in applying for this position, please apply online on this web page or mail resume to HR at Cerebras Systems Inc., 1237 E Arques Avenue, Sunnyvale, CA 94085. Please reference Job # 144 on resume or cover letter.

Why Join Cerebras

People who are serious about software make their own hardware. At Cerebras, we have built a breakthrough architecture that is unlocking new opportunities for the AI industry. With dozens of model releases and rapid growth, we’ve reached an inflection point in our business. Members of our team tell us there are five main reasons they joined Cerebras:

  1. Build a breakthrough AI platform beyond the constraints of the GPU.

  2. Publish and open source their cutting-edge AI research.

  3. Work on one of the fastest AI supercomputers in the world.

  4. Enjoy job stability with startup vitality.

  5. Our simple, non-corporate work culture that respects individual beliefs.

Find out more about what it's like to work at Cerebras here!

Apply today and become part of the forefront of groundbreaking advancements in AI!

Cerebras Systems is committed to creating an equal and diverse environment and is proud to be an equal opportunity employer. We celebrate different backgrounds, perspectives, and skills. We believe inclusive teams build better products and companies. We try every day to build a work environment that empowers people to do their best work through continuous learning, growth and support of those around them.

This website or its third-party tools process personal data. For more details, click here to review our CCPA disclosure notice.

Skills Required

  • Master's degree in Electrical Engineering, Computer Engineering, Computer Science, or related field (or foreign equivalent)
  • Minimum 3 years of experience as a Hardware Analytics Engineer, Hardware Engineer, Data Engineer, or related occupation
  • Large-scale data pipeline architecture and ETL design
  • Distributed data processing experience (Hive, Spark)
  • Experience developing dashboards and data visualizations (Tableau)
  • Proficiency in Python
  • Proficiency in SQL
  • Experience with Linux and automation scripting
  • Design, training, and deployment of machine learning models for hardware performance optimization and failure prediction
  • Predictive modeling, statistical analysis, A/B testing, and anomaly detection for hardware reliability
  • Hardware analytics experience for compute, storage, and AI servers (CPU, GPU, DRAM, SSD, PCIe, networking, storage)
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
774 Employees
Year Founded: 2015

What We Do

Cerebras Systems develops wafer-scale semiconductor hardware, AI supercomputers, and software/cloud services for training and inference. Its CS-2 and CS-3 systems help organizations build on-premise AI supercomputers, while pay-as-you-go cloud offerings provide developers and enterprises access to its computing platform. The company focuses on making AI training and inference faster and easier for diverse research and production workloads at scale worldwide.

Similar Jobs

Relativity Space Logo Relativity Space

Development Engineer

Aerospace • Hardware • Robotics • Software • Manufacturing
Easy Apply
In-Office
Long Beach, CA, USA
2200 Employees
148K-222K Annually

CrowdStrike Logo CrowdStrike

Instructional Designer

Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity
Remote or Hybrid
USA
11000 Employees
130K-200K Annually

Optum Logo Optum

Registered Nurse - Field Assessor - Chula Vista, CA

Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
In-Office
Chula Vista, CA, USA
160000 Employees
41-62 Hourly

Optum Logo Optum

Registered Nurse - Field Assessor - San Jose, CA

Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
In-Office
San Jose, CA, USA
160000 Employees
44-67 Hourly

Similar Companies Hiring

Kepler  Thumbnail
Artificial Intelligence • Fintech • Software
New York, New York
9 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Revel.io Thumbnail
Aerospace • Hardware • Robotics • Software
US
50 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account