HPC Engineer

Reposted 29 Days Ago
Be an Early Applicant
Sunnyvale, CA, USA
In-Office
150K-300K Annually
Junior
Information Technology • Automation • Manufacturing
The Role
Provide overnight operational coverage for large-scale GPU clusters: monitor health and performance, triage incidents, support researchers, run recovery procedures, validate deployments and upgrades, track utilization, and build automation and monitoring tools.
Summary Generated by Built In
About MBZUAI
The Institute for Foundation Models (IFM) operates some of the world's largest AI supercomputing environments.
Position Summary
This role provides operational coverage during Abu Dhabi overnight hours and serves as a primary point of contact for infrastructure monitoring, incident triage, researcher support, and production operations.

Responsibilities

    • Monitor health, performance, and availability of large-scale GPU clusters.
    • Respond to incidents and perform first-level triage.
    • Support researchers and troubleshoot job failures.
    • Execute operational runbooks and recovery procedures.
    • Validate cluster deployments, upgrades, and maintenance activities.
    • Track infrastructure utilization and operational metrics.
    • Develop automation and monitoring tools.
    • Contribute to documentation and reporting.

Education

    Bachelor's degree in Computer Science, Computer Engineering, Software Engineering, Information Technology, Electrical Engineering, Mathematics, Physics, or related disciplines.

Experience

    • 2+ years in Linux systems administration, SRE, DevOps, cloud operations, HPC, or infrastructure operations.
    • Strong Linux troubleshooting skills.
    • Experience with scripting using Python or Bash.

Preferred Qualifications

    • Slurm.
    • GPU infrastructure.
    • AWS, Azure, or GCP.
    • Grafana, Prometheus, Datadog, or similar tools.
    • Containers and Kubernetes.
    • AI/ML infrastructure exposure.
    • Research computing environments.

Benefits Include
*Comprehensive medical, dental, and vision benefits 
 *Bonus
*401K Plan
*Generous paid time off, sick leave and holidays
*Paid Parental Leave
*Employee Assistance Program
*Life insurance and disability
 

Skills Required

  • Bachelor's degree in Computer Science, Computer Engineering, Software Engineering, IT, Electrical Engineering, Mathematics, Physics, or related
  • 2+ years in Linux systems administration, SRE, DevOps, cloud operations, HPC, or infrastructure operations
  • Strong Linux troubleshooting skills
  • Scripting experience with Python or Bash
  • Experience with Slurm
  • Experience with GPU infrastructure
  • Experience with AWS, Azure, or GCP
  • Experience with Grafana, Prometheus, Datadog, or similar monitoring tools
  • Experience with containers and Kubernetes
  • Exposure to AI/ML infrastructure or research computing environments
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Essen
3,924 Employees
Year Founded: 1969

What We Do

First a passion, then an idea transformed into success – when it comes to pioneering automation and digitalisation technology, the ifm group is the ideal partner. Since its foundation in 1969, ifm has developed, produced and sold sensors, controllers, software and systems for industrial automation and for SAP-based solutions for supply chain management and shop floor integration worldwide. As one of the pioneers of Industry 4.0, ifm develops and implements consistent solutions to digitalise the entire value chain “from sensor to ERP”. Today, the second-generation family-run ifm group has more than 8,750 employees and is one of the worldwide market leaders. The group combines the internationality and innovative strength of a growing group of companies with the flexibility and close customer contact of a medium-sized company.

Similar Jobs

Vast.ai Logo Vast.ai

Support Engineer

Artificial Intelligence • On-Demand • Software
In-Office
Los Angeles, CA, USA
41 Employees
90K-150K Annually

Vast.ai Logo Vast.ai

Systems Engineer

Artificial Intelligence • On-Demand • Software
In-Office
2 Locations
41 Employees
160K-320K Annually
In-Office
Milpitas, CA, USA
10001 Employees
136K-200K Annually
In-Office
Milpitas, CA, USA
10001 Employees
136K-232K Annually

Similar Companies Hiring

Rosendin Thumbnail
Other • Manufacturing
San Jose, CA
6219 Employees
Amalgamated Sugar Thumbnail
Food • Greentech • Agriculture • Industrial • Manufacturing
Boise, Idaho
768 Employees
Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account