HPC Engineer

Reposted 16 Days Ago
Be an Early Applicant
Sunnyvale, CA, USA
In-Office
150K-300K Annually
Junior
Information Technology • Automation • Manufacturing
The Role
Provide overnight operational coverage for large-scale GPU clusters: monitor health and performance, triage incidents, support researchers, run recovery procedures, validate deployments and upgrades, track utilization, and build automation and monitoring tools.
Summary Generated by Built In
About MBZUAI
The Institute for Foundation Models (IFM) operates some of the world's largest AI supercomputing environments.
Position Summary
This role provides operational coverage during Abu Dhabi overnight hours and serves as a primary point of contact for infrastructure monitoring, incident triage, researcher support, and production operations.

Responsibilities

    • Monitor health, performance, and availability of large-scale GPU clusters.
    • Respond to incidents and perform first-level triage.
    • Support researchers and troubleshoot job failures.
    • Execute operational runbooks and recovery procedures.
    • Validate cluster deployments, upgrades, and maintenance activities.
    • Track infrastructure utilization and operational metrics.
    • Develop automation and monitoring tools.
    • Contribute to documentation and reporting.

Education

    Bachelor's degree in Computer Science, Computer Engineering, Software Engineering, Information Technology, Electrical Engineering, Mathematics, Physics, or related disciplines.

Experience

    • 2+ years in Linux systems administration, SRE, DevOps, cloud operations, HPC, or infrastructure operations.
    • Strong Linux troubleshooting skills.
    • Experience with scripting using Python or Bash.

Preferred Qualifications

    • Slurm.
    • GPU infrastructure.
    • AWS, Azure, or GCP.
    • Grafana, Prometheus, Datadog, or similar tools.
    • Containers and Kubernetes.
    • AI/ML infrastructure exposure.
    • Research computing environments.

Benefits Include
*Comprehensive medical, dental, and vision benefits 
 *Bonus
*401K Plan
*Generous paid time off, sick leave and holidays
*Paid Parental Leave
*Employee Assistance Program
*Life insurance and disability
 

Skills Required

  • Bachelor's degree in Computer Science, Computer Engineering, Software Engineering, IT, Electrical Engineering, Mathematics, Physics, or related
  • 2+ years in Linux systems administration, SRE, DevOps, cloud operations, HPC, or infrastructure operations
  • Strong Linux troubleshooting skills
  • Scripting experience with Python or Bash
  • Experience with Slurm
  • Experience with GPU infrastructure
  • Experience with AWS, Azure, or GCP
  • Experience with Grafana, Prometheus, Datadog, or similar monitoring tools
  • Experience with containers and Kubernetes
  • Exposure to AI/ML infrastructure or research computing environments
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Essen
3,924 Employees
Year Founded: 1969

What We Do

First a passion, then an idea transformed into success – when it comes to pioneering automation and digitalisation technology, the ifm group is the ideal partner. Since its foundation in 1969, ifm has developed, produced and sold sensors, controllers, software and systems for industrial automation and for SAP-based solutions for supply chain management and shop floor integration worldwide. As one of the pioneers of Industry 4.0, ifm develops and implements consistent solutions to digitalise the entire value chain “from sensor to ERP”. Today, the second-generation family-run ifm group has more than 8,750 employees and is one of the worldwide market leaders. The group combines the internationality and innovative strength of a growing group of companies with the flexibility and close customer contact of a medium-sized company.

Similar Jobs

NVIDIA Logo NVIDIA

Support Engineer

Artificial Intelligence • Computer Vision • Hardware • Robotics • Metaverse
In-Office
5 Locations
21960 Employees
108K-207K Annually

Biohub Logo Biohub

Staff HPC Engineer

Artificial Intelligence • Healthtech • Machine Learning • Biotech
In-Office
San Francisco, CA, USA
468 Employees
214K-300K Annually
In-Office
Milpitas, CA, USA
10001 Employees
136K-232K Annually

Hewlett Packard Enterprise Logo Hewlett Packard Enterprise

Senior Software Engineer

Artificial Intelligence • Cloud • Information Technology • Consulting
In-Office
7 Locations
85422 Employees
120K-275K Annually

Similar Companies Hiring

Standard Template Labs Thumbnail
Artificial Intelligence • Information Technology • Software
New York, NY
25 Employees
Fortune Brands Innovations Thumbnail
Manufacturing
Deerfield, IL
10000 Employees
Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account