Senior HPC Hardware Engineer

Posted 5 Days Ago
Be an Early Applicant
Dallas, TX, USA
In-Office
Senior level
Artificial Intelligence • Cloud • Machine Learning • Infrastructure as a Service (IaaS)
The Role
Own the full lifecycle of a large-scale HPC and AI compute fleet, including bare-metal provisioning, firmware and BIOS management, hardware troubleshooting, performance tuning, capacity planning, security hardening, and infrastructure automation. Validate next-generation NVIDIA platforms, collaborate with software, networking, and vendor teams, resolve complex component issues, and mentor junior engineers.
Summary Generated by Built In

THE COMPANY

NorthMark Compute & Cloud (NMC²) is backed by dedicated leadership and investment, with a clear mission as it operates at the bleeding edge of technology. Its goal is to scale and enhance the high-performance computing (HPC) and cloud infrastructure that supports its clients’ research, production, and delivery, enabling breakthroughs that shape the industries of tomorrow. Its engineers build critical infrastructure to eliminate friction in scientific research, simulations, analysis, and decision-making, accelerating discovery and driving faster innovation.

THE POSITION

NMC² is seeking a Senior HPC Hardware Engineer to join the Compute Engineering team based at our Dallas, TX offices at Victory Commons. This is a hands-on role at the center of one of the most demanding and rapidly scaling HPC environments in operation, spanning a large fleet of GPU and CPU nodes built on the latest NVIDIA platforms including H200 and GB200/NVL72 architectures.

You will own the full hardware lifecycle for NMC²’s compute fleet — from bare-metal provisioning and firmware baseline management through production validation, troubleshooting, and capacity planning. You will be the go-to expert on server hardware architecture, driving standards and automation that ensure the fleet operates at peak performance and availability. Your work directly enables the research and delivery workloads that NMC²’s clients depend on.

This role requires close collaboration with Software Engineering, Networking, and Vendor teams, and involves mentoring junior engineers. The ideal candidate is a technically deep infrastructure leader who thrives in fast-paced environments, brings a strong automation mindset, and has a proven track record managing large-scale HPC or AI compute infrastructure.

RESPONSIBILITIES

  • Design, configure, and manage a high-performance compute fleet comprising large-scale GPU (NVIDIA V100/A100/H200/GB200) and CPU nodes across NMC²’s Dallas infrastructure.

  • Own the full firmware and BIOS lifecycle across the HPC/AI fleet — from establishing baselines and validation through rollout, compliance, and ongoing maintenance.

  • Lead troubleshooting of hardware components including CPUs, GPUs, DPUs, NVSwitches, NICs, memory, PSUs, and BMCs; drive component replacement and configuration remediation.

  • Automate health checks, onboarding workflows, and recurring hardware issue remediation to accelerate safe deployment and reduce recovery time.

  • Validate and operationalize next-generation AI platforms (e.g., NVL72 / Grace Blackwell) from day one, ensuring stability, performance readiness, and production fitness.

  • Collaborate with vendors on firmware and hardware issues, providing clear reproduction cases, diagnostic logs, and business impact to drive timely resolution.

  • Perform hardware performance analysis, tuning, and capacity planning to ensure reliable scale-out of the compute environment.

  • Define and implement security hardening best practices for hardware infrastructure, maintaining platform integrity across the fleet.

  • Leverage Infrastructure as Code (IaC) methodologies and scripting to drive efficient, repeatable, and scalable infrastructure management.

  • Mentor junior engineers, act as a subject matter expert for infrastructure-related escalations, and champion a culture of continuous improvement across the team.

REQUIREMENTS

  • Bachelor’s degree in Electrical Engineering, Computer Engineering, or a related field, or equivalent hands-on experience.

  • 5+ years of experience managing large-scale HPC or AI compute infrastructure in a production environment.

  • Deep knowledge of server hardware architecture, including processors, memory, storage, networking, power systems, and thermal management.

  • Hands-on experience with bare-metal provisioning, firmware and BIOS lifecycle management, and hardware automation tools such as Ansible, Puppet, or Chef.

  • Proficiency with Redfish API and BMC/IPMI tooling (iDRAC, iLO) for remote hardware management and diagnostics.

  • Demonstrated ability to troubleshoot and resolve complex hardware issues across GPU and CPU nodes, including NVIDIA-SMI and GPU diagnostics.

  • Experience with hardware monitoring platforms, performance tuning, and capacity planning at scale.

  • Familiarity with Linux-based environments and scripting proficiency in Python, Bash, or PowerShell for infrastructure automation.

  • Experience with OpenStack (particularly Ironic) or equivalent cloud/bare-metal provisioning platforms is strongly preferred.

  • Strong cross-functional communication skills and proven ability to collaborate effectively with software, networking, and vendor teams.

  • Prior technical leadership experience, including mentoring engineers and driving team-wide best practices.

It is impossible to list every requirement for, or responsibility of, any position.  Similarly, we cannot identify all the skills a position may require since job responsibilities and the Company’s needs may change over time.  Therefore, the above job description is not comprehensive or exhaustive.  The Company reserves the right to adjust, add to or eliminate any aspect of the above description.  The Company also retains the right to require all employees to undertake additional or different job responsibilities when necessary to meet business needs.

Must be legally authorized to work in the United States without the need for employer sponsorship, now or at any time in the future.

Benefits & Perks:

  • Company-Paid Lunch Stipend: Lunch is provided via GrubHub

  • Company-Paid Benefits: 100% Employer-Paid Medical in our High Deductible Health Plan, Dental and Vision benefits for employees and their families, 16 weeks of Paid Parental Leave, Employee Assistance Program, Life insurance, Short-Term Disability and Long-Term Disability

  • 401(k): Company will match 100% of your contributions up to 6%

  • Optional Employee-Paid Benefits: Medical insurance in our PPO plan and a variety of other benefits such as Health Savings Accounts (with Company Contribution!), Flexible Spending Accounts, Supplemental Life Insurance, Wellhub and more.

  • Time Off:  25 days of Paid Time Off plus 12 company holidays

EQUAL OPPORTUNITY EMPLOYER

NORTHMARK STRATEGIES LLC IS AN EQUAL EMPLOYMENT OPPORTUNITY EMPLOYER. THE COMPANY'S POLICY IS NOT TO DISCRIMINATE AGAINST ANY APPLICANT OR EMPLOYEE BASED ON RACE, COLOR, RELIGION, NATIONAL ORIGIN, GENDER, AGE, SEXUAL ORIENTATION, GENDER IDENTITY OR EXPRESSION, MARITAL STATUS, MENTAL OR PHYSICAL DISABILITY, AND GENETIC INFORMATION, OR ANY OTHER BASIS PROTECTED BY APPLICABLE LAW. THE FIRM ALSO PROHIBITS HARASSMENT OF APPLICANTS OR EMPLOYEES BASED ON ANY OF THESE PROTECTED CATEGORIES.

Skills Required

  • Bachelor’s degree in Electrical Engineering, Computer Engineering, a related field, or equivalent hands-on experience
  • 5+ years managing large-scale HPC or AI compute infrastructure in production
  • Deep knowledge of server hardware architecture, processors, memory, storage, networking, power systems, and thermal management
  • Hands-on experience with bare-metal provisioning and firmware and BIOS lifecycle management
  • Experience with hardware automation tools such as Ansible, Puppet, or Chef
  • Proficiency with Redfish API and BMC/IPMI tooling, including iDRAC or iLO
  • Ability to troubleshoot GPU and CPU nodes using NVIDIA-SMI and GPU diagnostics
  • Experience with hardware monitoring, performance tuning, and capacity planning at scale
  • Familiarity with Linux-based environments
  • Scripting proficiency in Python, Bash, or PowerShell
  • Experience with OpenStack, particularly Ironic, or equivalent cloud/bare-metal provisioning platforms
  • Strong cross-functional communication and collaboration skills
  • Technical leadership experience, including mentoring engineers and driving team-wide best practices
  • Legal authorization to work in the United States without employer sponsorship now or in the future
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
157 Employees

What We Do

NorthMark Strategies is a strategic capital firm that combines investment capital with engineering and technology to build enduring businesses. The firm operates a High-Performance Computing platform and supports simulation, AI/ML-enabled engineering and data-driven design to accelerate portfolio companies. NorthMark deploys capital, operates complex businesses, and builds infrastructure (including compute and cloud services) to drive long‑term innovation and operational outcomes.

Similar Jobs

Tapestry - Coach and Kate Spade Logo Tapestry - Coach and Kate Spade

Sales Support Associate III

eCommerce • Fashion • Retail • Sales • Wearables • Design
Hybrid
Cypress, TX, USA
16000 Employees
15-20 Hourly

Closinglock Logo Closinglock

Customer Support Representative

Fintech • Real Estate • Security • Software • Financial Services • Cybersecurity • PropTech
In-Office
Austin, TX, USA
110 Employees

SoFi Logo SoFi

Senior Manager, Secured Commercial Real Estate Lending

Fintech • Mobile • Software • Financial Services
Easy Apply
Hybrid
Frisco, TX, USA
4500 Employees
115K-198K Annually

Invenergy Logo Invenergy

O&M Manager

Greentech • Real Estate • Social Impact • Energy • Industrial • Solar • Renewable Energy
In-Office
Big Spring, TX, USA
2500 Employees
115K-155K Annually

Similar Companies Hiring

Legora Thumbnail
Artificial Intelligence • Legal Tech • Software
New York, New York
700 Employees
Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account