HPC Infrastructure & Cluster Engineer

Posted 3 Days Ago
Be an Early Applicant
Springfield, VA, USA
In-Office
120K-162K Annually
Senior level
Aerospace • Information Technology • Professional Services • Security • Software
The Role
Administer and optimize a dedicated HPC compute cluster, including Linux systems, hardware, storage, networking, scheduling, patching, and upgrades. Configure Run:AI and SLURM for AI/ML workloads, manage InfiniBand infrastructure, provision OpenShift or Kubernetes environments, automate maintenance with Bash and Python, troubleshoot complex infrastructure issues, and maintain federal security compliance and accreditations.
Summary Generated by Built In

Type of Requisition:

Regular

Clearance Level Must Currently Possess:

Top Secret/SCI

Clearance Level Must Be Able to Obtain:

Top Secret SCI + Polygraph

Public Trust/Other Required:

None

Job Family:

IT Infrastructure and Operations

Job Qualifications:

Skills:

Cluster Administration, Infrastructure Optimization, Storage Area Network (SAN) Management

Certifications:

None

Experience:

5 + years of related experience

US Citizenship Required:

Yes

Job Description:

Position Summary

We are seeking an Infrastructure & Cluster Engineer to manage the administration, health, and performance of the foundational compute environment under the User Facing and Data Center Services (UDS) contract at GDIT. In this role, you will be responsible for the end-to-end administration of a dedicated customer compute cluster. Your primary mission is to ensure a highly available, secure, and optimized hardware foundation. By maintaining a robust infrastructure, you will directly contribute to the critical technology integration and performance engineering efforts, ensuring a highly reliable platform for integrating and executing complex customer workloads.

Key Responsibilities:

  • Cluster Administration: Manage the day-to-day operations of the customer compute cluster, including Linux operating system administration, hardware monitoring, patching, and system upgrades.
  • Resource and Job Management: Configure, maintain, and optimize workload management and orchestration platforms, utilizing the Run:AI job scheduler to ensure efficient distribution of intensive AI/ML workloads across the cluster.
  • Infrastructure Optimization: Tune cluster performance at the hardware, operating system, and network levels to maximize compute efficiency and data throughput for customer workloads.
  • Storage and Network Management: Administer storage solutions and high-speed networking fabrics. Support the transition to and ongoing management of an InfiniBand GPU-to-GPU network infrastructure to minimize latency for distributed operations. 
  • Environment Configuration: Partner with technology integration teams to provision specific environments, dependencies, and container platforms, specifically leveraging Red Hat OpenShift, required for seamless customer model deployment.
  • Security and Compliance: Ensure all infrastructure components remain compliant with federal security standards, implementing strict access controls and maintaining system accreditations.

Basic Qualifications:

  • Clearance: Active TS/SCI with the ability to obtain CI Poly.
  • Experience: 5+ years of experience in Linux systems administration and infrastructure management with a specific focus on high-performance computing environments.
  • Technical Skills:
    • Expertise in managing bare-metal servers, enterprise storage arrays, and advanced network configurations (Experience with InfiniBand).
    • Strong proficiency with workload managers, job schedulers, and AI orchestration tools (e.g., Run:AI, SLURM).
    • Hands-on experience with enterprise container orchestration platforms, specifically OpenShift or Kubernetes.
    • Experience writing automation and configuration scripts (e.g., Bash, Python) to streamline cluster maintenance.
  • Troubleshooting Focus: Proven ability to diagnose and resolve complex hardware, network, and OS-level issues.

Preferred Qualifications:

  • Familiarity with parallel file systems and high-throughput storage architectures.
  • Prior experience engineering or managing high-speed GPU-to-GPU communication topologies.

Location: Springfield, VA

US Citizenship Required 

GDIT IS YOUR PLACE:
● 401K with company match
● Comprehensive health and wellness packages
● Internal mobility team dedicated to helping you own your career
● Professional growth opportunities including paid education and certifications
● Cutting-edge technology you can learn from
● Rest and recharge with paid vacation and holidays

#RoverGSS

The likely salary range for this position is $119,850 - $162,150. This is not, however, a guarantee of compensation or salary. Rather, salary will be set based on experience, geographic location and possibly contractual requirements and could fall outside of this range.

Scheduled Weekly Hours:

40

Travel Required:

None

Telecommuting Options:

Onsite

Work Location:

USA VA Springfield

Additional Work Locations:

Total Rewards at GDIT:

Our benefits package for all US-based employees includes a variety of medical plan options, some with Health Savings Accounts, dental plan options, a vision plan, and a 401(k) plan offering the ability to contribute both pre and post-tax dollars up to the IRS annual limits and receive a company match. To encourage work/life balance, GDIT offers employees full flex work weeks where possible and a variety of paid time off plans, including vacation, sick and personal time, holidays, paid parental, military, bereavement and jury duty leave. To ensure our employees are able to protect their income, other offerings such as short and long-term disability benefits, life, accidental death and dismemberment, personal accident, critical illness and business travel and accident insurance are provided or available. We regularly review our Total Rewards package to ensure our offerings are competitive and reflect what our employees have told us they value most.

 



Our Identity Verification Process:

As part of the hiring process, we will ask you to complete an identity verification process that leverages advanced biometrics and artificial intelligence to ensure authenticity and protect against identity fraud. You are expected to be on camera during virtual interviews. We reserve the right to take your picture to verify your identity and prevent fraud. By proceeding, you authorize the collection, processing, and use of your biometric data for identity verification and security purposes.

About Our Work:

We are GDIT. A global technology and professional services company that delivers technology solutions and mission services to every major agency across the U.S. government, defense and intelligence community. Our 26,000 experts extract the power of technology to create immediate value and deliver solutions at the edge of innovation. We operate across 50+ countries worldwide, offering leading mission-ready capabilities in AI, cloud, cyber and software development.

Join our Talent Community to stay up to date on our career opportunities and events at

gdit.com/tc.

Equal Opportunity Employer / Individuals with Disabilities / Protected Veterans

Skills Required

  • 5+ years of related experience
  • 5+ years of Linux systems administration and infrastructure management experience focused on high-performance computing
  • Active Top Secret/SCI clearance
  • Ability to obtain CI Polygraph
  • US citizenship
  • Experience managing bare-metal servers, enterprise storage arrays, and advanced network configurations
  • Experience with InfiniBand
  • Proficiency with workload managers, job schedulers, and AI orchestration tools such as Run:AI or SLURM
  • Hands-on experience with OpenShift or Kubernetes
  • Experience writing Bash or Python automation and configuration scripts
  • Ability to diagnose and resolve complex hardware, network, and operating system issues
  • Familiarity with parallel file systems and high-throughput storage architectures
  • Experience engineering or managing high-speed GPU-to-GPU communication topologies

General Dynamics Information Technology Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about General Dynamics Information Technology and has not been reviewed or approved by General Dynamics Information Technology.

  • Affordable Benefits Pay and benefits are described as good or okay in multiple places, and the overall package is often portrayed as acceptable even when base pay is not viewed as top-tier.
  • Healthcare Strength Medical, dental, and vision plan options are presented as comprehensive, with additional protections like disability and life insurance contributing to a well-rounded health and protection offering.
  • Retirement Support A 401(k) plan with company match is consistently highlighted as part of the total rewards package, supporting longer-term financial planning.

General Dynamics Information Technology Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Falls Church, VA
21,625 Employees

What We Do

We are GDIT. The people supporting some of the most complex government, defense, and intelligence projects across the country. We deliver. Bringing the expertise needed to understand and advance critical missions. We transform. Shifting the ways clients invest in, integrate, and innovate technology solutions. We ensure today is safe and tomorrow is smarter. We are there. On the ground, beside our clients, in the lab, and everywhere in between. Offering the technology transformations, strategy, and mission services needed to get the job done.

Similar Jobs

INflow Federal Logo INflow Federal

HPC Infrastructure & Cluster Engineer

Cloud • Information Technology • Cybersecurity
In-Office
Springfield, VA, USA
44 Employees
140K-185K Annually

D2 Consulting Logo D2 Consulting

HPC Infrastructure & Cluster Engineer

Big Data • Cloud • Information Technology • Consulting • Cybersecurity
In-Office
Springfield, VA, USA
30 Employees
170K-180K Annually

Samsara Logo Samsara

Business Systems Analyst

Artificial Intelligence • Cloud • Computer Vision • Hardware • Internet of Things • Software
Easy Apply
Remote or Hybrid
USA
4000 Employees
111K-150K Annually

Samsara Logo Samsara

Security Engineer

Artificial Intelligence • Cloud • Computer Vision • Hardware • Internet of Things • Software
Easy Apply
Remote or Hybrid
United States
4000 Employees
165K-266K Annually

Similar Companies Hiring

Kepler  Thumbnail
Artificial Intelligence • Fintech • Software
New York, New York
9 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Revel.io Thumbnail
Aerospace • Hardware • Robotics • Software
US
50 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account