HPC Systems Engineer

Posted 2 Days Ago
Be an Early Applicant
Fort Worth, TX, USA
Hybrid
Senior level
Artificial Intelligence • Cloud • Infrastructure as a Service (IaaS) • Renewable Energy
The Role
Designs, deploys, integrates, and optimizes high-performance computing clusters using Kubernetes, Slurm, and HPC management tools. Responsibilities include maintaining cluster availability, monitoring and troubleshooting systems, documenting architectures and procedures, adopting new technologies, collaborating on operational best practices, and providing technical leadership and training.
Summary Generated by Built In

Job Type: Full-time | Location: Dallas Fort-Worth, Texas | Department: Information Technologies (IT) | Reporting to: Manager, Systems Engineer | Work Location: #onsite #hybrid

IREN is a leading AI Cloud Service Provider, delivering large-scale GPU clusters for AI training and inference. IREN’s vertically integrated platform is underpinned by its expansive portfolio of grid-connected land and data centers in renewable-rich regions across the U.S. and Canada. 

With 100% renewable energy, we build, own and operate our data centers and take pride in being at the forefront of sustainable solutions for the ever-evolving applications of high-performance compute. We believe that human progress is invaluable, but it should be done in the right way – responsibly, sustainably and having a positive impact on the communities we operate in.    

The HPC Systems Engineer will spearhead the design, deployment, and optimization of our high-performance computing (HPC) systems with a focus on HPC clustering, Kubernetes, Slurm, and management tools. The ideal candidate will seamlessly blend traditional HPC solutions with modern container orchestration, ensuring a versatile and scalable computing environment that can be offered to our customers.
Job requirements
  • Minimum of 5 years of experience in HPC system architecture with proven expertise in designing, deploying, and managing HPC clusters.
  • Extensive knowledge of Kubernetes, with a focus on its integration within HPC environments.
  • Hands-on experience with the Slurm workload manager and its intricacies.
  • Familiarity with HPC management tools and software, ensuring efficient system monitoring and troubleshooting.
  • Proven track record of resolving complex system challenges and enhancing operational performance.
  • A Bachelor's or Master's degree in Computer Science, Engineering, or a related field.
  • Relevant certifications in Kubernetes, HPC technologies, or system architecture are advantageous.
  • Understanding of cloud platforms and their integration into HPC ecosystems.
  • Deep knowledge of network and storage solutions commonly used in HPC setups.

Key Attributes:

  • Analytical mindset, adept at envisioning and designing intricate systems.
  • Collaborative approach, with the ability to work effectively with diverse technical teams.
  • Excellent communication skills, translating complex technical concepts into understandable terms for varied audiences.
  • Detail-oriented focus, ensuring systems are both robust and efficient.
  • Continuous learner, keen to stay updated with rapid technological evolutions in the HPC and Kubernetes domains.

Other important requirements:  

  • Pre-employment screening, including background check and substance testing may be required according to company policies. 
  • Must provide own steel-toed work boots; other PPE will be supplied.
  • Must be able to reliably commute to the work site daily or have plans to relocate before starting work.

Job responsibilities
  • Lead the deployment and maintenance of HPC clusters, ensuring they operate effectively and maximise availability
  • Integrate and manage HPC software components such as Kubernetes, Slurm, cluster management software, and any infrastructure required to operate the HPC environment
  • Stay abreast of advancements in HPC, Kubernetes, and associated technologies, bringing innovations into our operations and product options.
  • Collaborate with technical teams to establish and implement best practices for system maintenance and optimization.
  • Draft comprehensive documentation, including system designs, operational procedures, and best practice guidelines.
  • Facilitate the selection and integration of relevant management tools to monitor, troubleshoot, and enhance HPC operations.
  • Provide technical leadership and training to other team members, fostering an environment of continuous learning and improvement.

Benefits

The IREN Package 

At IREN, we offer a highly competitive compensation package that includes base salary, annual performance incentives, and opportunities to build long-term wealth through equity programs. These offerings are part of our broader Total Rewards package, thoughtfully designed to support your health, well-being, and long-term success. 

Compensation 

  • Actual compensation will be determined based on factors such as experience, qualifications, and market data for the region.
  • Total Compensation package may be inclusive of annual incentive bonus, equity (long-term incentive).

Health & Wellness 

  • 100% company paid health insurance premiums(medical, dental, and vision)for employees, 75% company paid coverage for dependents.
  • Company-paid short-term and long-term disability insurance.
  • Voluntary life, critical illness, and accident coverage available.
  • Health Savings Accounts (HSA) – when combined with the High Deductible Health Plan.
  • Employee Assistance Program and wellness resources.

Retirement & Financial Wealth 

  • 401(k) retirement plan with company match.
  • Access to financial planning and legal services.

Time Off & Leave Programs 

  • Paid Time Off (PTO) and paid holidays.

Growth & Development 

  • Internal skills training and advancement pathways.
  • Professional development to support certifications, continuing education, or role related training.

Community & Culture 

  • Company events and team-building activities.

We value diverse perspectives and believe that skills can be developed. If you’re passionate about this role, we want to hear from you — whether you meet every criteria or not. Your unique experiences might be exactly what we need!   

IE US Operations Inc., the employing entity and proud member of the IREN group is an equal opportunity employer that is committed to creating an inclusive workplace. We are committed to evaluating qualified applicants and do not discriminate against protected characteristics under applicable legislation. 

By applying for this position and submitting your resume and application materials, you consent to the processing of your personal information in accordance with our Job Applicant Privacy Statement available on our website at www.iren.com.   

Skills Required

  • At least 5 years of experience in HPC system architecture, including designing, deploying, and managing HPC clusters
  • Extensive knowledge of Kubernetes and its integration with HPC environments
  • Hands-on experience with the Slurm workload manager
  • Familiarity with HPC management tools for system monitoring and troubleshooting
  • Experience resolving complex system challenges and improving operational performance
  • Bachelor's or Master's degree in Computer Science, Engineering, or a related field
  • Understanding of cloud platforms and their integration into HPC ecosystems
  • Deep knowledge of networking and storage solutions used in HPC environments
  • Relevant certifications in Kubernetes, HPC technologies, or system architecture
  • Ability to reliably commute to the work site daily or relocate before starting
  • Ability to provide personal steel-toed work boots
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
644 Employees
Year Founded: 2018

What We Do

IREN is a vertically integrated AI Cloud provider that develops, owns, and operates large-scale data centers and GPU clusters for AI training and inference. Its platform combines direct access to scalable NVIDIA GPUs with grid-connected land and renewable power across North America and other regions. The company emphasizes sustainable, high-performance computing infrastructure for AI builders and enterprise customers across diverse workloads.

Similar Jobs

NorthMark Strategies Logo NorthMark Strategies

Systems Engineer

Artificial Intelligence • Cloud • Machine Learning • Infrastructure as a Service (IaaS)
In-Office
Dallas, TX, USA
157 Employees

SpaceX Logo SpaceX

Systems Engineer

Aerospace • Other
In-Office
Star, TX, USA
8879 Employees

PwC Logo PwC

Pricing and Revenue Consulting Manager - Consumer Markets Sector

Artificial Intelligence • Professional Services • Business Intelligence • Consulting • Cybersecurity • Generative AI
Hybrid
58 Locations
370000 Employees
99K-232K Annually

PwC Logo PwC

Conversational AI and Agentic AI - Manager

Artificial Intelligence • Professional Services • Business Intelligence • Consulting • Cybersecurity • Generative AI
Hybrid
4 Locations
370000 Employees
99K-232K Annually

Similar Companies Hiring

Kepler  Thumbnail
Artificial Intelligence • Fintech • Software
New York, New York
9 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software
New York, New York
30 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account