Senior HPC Support Engineer & Deployment Lead

Posted 5 Hours Ago
Be an Early Applicant
2 Locations
Remote or Hybrid
Senior level
Information Technology • Software
The Role
Lead customer-facing HPC deployments and serve as the escalation point for complex hardware, operating system, application, and compiler issues. Deploy and debug GPU-accelerated Kubernetes clusters across on-premises and cloud environments, using Python, C#/.NET, CUDA, and ROCm. Diagnose issues involving NVIDIA and AMD GPUs, high-speed interconnects, PCIe topology, and driver stacks. Create technical documentation, reproduce customer failures, and collaborate with engineering and R&D on bug reports and product improvements.
Summary Generated by Built In
Your Mission

You will act as the escalation point for our most challenging technical hurdles, ensuring that our compiler technology runs flawlessly on the world's most powerful hardware.

  • Think: Lead the architectural strategy for customer rollouts. You will analyze client infrastructure—evaluating power, high-speed interconnects (Infiniband/RoCE), and software environments—to plan successful cluster deployments. You will drive complex customer issues to resolution by diagnosing root causes that sit between hardware, the OS, and our application layer.

  • Implement:

    • Execute hands-on deployments of Kubernetes clusters (on-prem and cloud) tailored for GPU acceleration.

    • Dive deep into code and systems to detail, reproduce, and resolve issues. You will set up test environments using C#, CUDA, and ROCm to mimic customer failures.

    • Work directly with the latest silicon (NVIDIA H100, AMD MI300) and interconnects to ensure our software utilizes the hardware correctly.

  • Build:

    • The Knowledge Base: You will author detailed technical solutions, white papers, and "known issue" documentations. Your work will empower the rest of the team and our users to solve problems faster.

    • Feedback Loops: Collaborate closely with the Engineering and R&D teams. You will translate field data into clear bug reports and feature requests, helping to shape the future stability of the product.

What You Bring to the Table

You are a "System Doctor." You have the computer science fundamentals to understand code, but your expertise lies in making that code run reliably on physical systems.

  • Experience: You have a BS/MS in Computer Science, Electrical Engineering, or related field, with 8+ years of experience in system software development and hardware support. You have a proven track record in customer-facing roles.

  • HPC & Hardware Fluency: You have a deep understanding of GPU architectures and how they interact with the rest of the system. You are comfortable dealing with high-speed interconnects, PCIe topology, and driver stacks.

  • Software Ecosystem: You possess strong computer science fundamentals. You are an expert in Python and scripting for automation, but you are also comfortable navigating C#/.NET environments and the CUDA/ROCm ecosystems.

  • Containerization: You have practical experience deploying and debugging Kubernetes clusters in production environments.

  • Communication: Excellent interpersonal skills are non-negotiable. You can remain calm under pressure, communicate complex technical details to stakeholders, and manage customer expectations effectively.

Skills Required

  • Bachelor's or master's degree in Computer Science, Electrical Engineering, or a related field
  • 8+ years of experience in system software development and hardware support
  • Proven experience in customer-facing technical roles
  • Deep understanding of GPU architectures and their interaction with system hardware
  • Experience with high-speed interconnects, PCIe topology, and driver stacks
  • Expertise in Python and scripting for automation
  • Experience with C# and .NET environments
  • Experience with CUDA and ROCm ecosystems
  • Practical experience deploying and debugging Kubernetes clusters in production
  • Excellent interpersonal and stakeholder communication skills
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
1 Employee
Year Founded: 2008

What We Do

Hybridizer is a software platform for performance portability and GPU acceleration. Its compiler transforms C#/.NET and Java bytecode or high-level code into optimized source code for multicore CPUs and GPUs, allowing developers to use existing codebases without learning CUDA or rewriting applications. The technology supports debugging, profiling, cross-platform deployment, and demanding workloads such as quantitative finance, scientific simulation, and data processing.

Similar Jobs

CDW Logo CDW

Data Architect

Information Technology
Remote or Hybrid
US
15100 Employees
128K-193K Annually

CDW Logo CDW

Support Engineer

Information Technology
Remote or Hybrid
US
15100 Employees
78K-108K Annually

Liberty Mutual Insurance Logo Liberty Mutual Insurance

Inside Sales Representative

Artificial Intelligence • Fintech • Insurance • Marketing Tech • Software • Analytics
Remote or Hybrid
4 Locations
40000 Employees
44K-85K Annually

Liberty Mutual Insurance Logo Liberty Mutual Insurance

Inside Sales Representative

Artificial Intelligence • Fintech • Insurance • Marketing Tech • Software • Analytics
Remote or Hybrid
12 Locations
40000 Employees
45K-85K Annually

Similar Companies Hiring

Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Revel Thumbnail
Aerospace • Hardware • Robotics • Software
Marina Del Rey, California
60 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software
New York, New York
30 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account