LLM Pre-training & Distributed Engineer (AI Infrastructure)

Reposted 2 Months Ago
Be an Early Applicant
Seattle, WA, USA
In-Office
Senior level
Agency • Artificial Intelligence • Blockchain • Web3
The Role
Design and orchestrate large-scale LLM pre-training across 1,000+ GPUs using PyTorch, DeepSpeed, or Megatron-LM. Optimize InfiniBand/RDMA networking and memory to avoid OOM, automate checkpointing and failure recovery for month-long runs, and manage SLURM or Kubernetes GPU clusters. Implement systems-level improvements using C++, CUDA, and Python.
Summary Generated by Built In

We are seeking a highly skilled LLM Pre-training & Distributed Systems Engineer. This role is essential for orchestrating large-scale machine learning training runs and optimizing  distributed infrastructure. The ideal candidate will have a deep understanding of GPU clusters and extensive experience in system engineering to ensure efficient and reliable training processes.

Responsibilities:

  • Orchestrate distributed training runs across 1,000+ GPUs using PyTorch, DeepSpeed, or Megatron-LM.
  • Optimize networking (InfiniBand/RDMA) and memory management to prevent out-of-memory errors.
  • Automate checkpointing and failure recovery during month-long training runs.

Required Skills:

  • Deep expertise in 3D parallelism (Data, Tensor, Pipeline).
  • Experience managing SLURM or Kubernetes-based GPU clusters.
  • Strong systems engineering background (C++, CUDA, Python).

Skills Required

  • Deep expertise in 3D parallelism (Data, Tensor, Pipeline).
  • Orchestrate distributed training runs across 1,000+ GPUs using PyTorch, DeepSpeed, or Megatron-LM.
  • Optimize networking (InfiniBand/RDMA) and memory management to prevent out-of-memory errors.
  • Automate checkpointing and failure recovery during month-long training runs.
  • Experience managing SLURM or Kubernetes-based GPU clusters.
  • Strong systems engineering background (C++, CUDA, Python).
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
7 Employees
Year Founded: 2024

What We Do

Hyphen Connect is a Web3 and AI talent agency and crypto-integrated software solutions provider that connects blockchain, DeFi, NFT, and AI companies with specialized technical and go-to-market talent globally and remotely. They deliver headhunting, data-driven research, and recruitment services across infrastructure, exchanges, gaming, and DeFi projects, plus industry analysis and hiring insights to help clients build engineering and product teams.

Similar Jobs

Snap Inc. Logo Snap Inc.

Technical Program Manager

Artificial Intelligence • Cloud • Machine Learning • Mobile • Software • Virtual Reality • App development
Hybrid
3 Locations
5000 Employees
162K-284K Annually

Applied Systems Logo Applied Systems

Portfolio Strategy Director

Artificial Intelligence • Cloud • Payments • Software • Business Intelligence • Generative AI • Automation
Remote or Hybrid
United States
3116 Employees
159K-190K Annually

CoreWeave Logo CoreWeave

Product Manager

Cloud • Information Technology • Machine Learning
In-Office
5 Locations
1450 Employees
153K-204K Annually

Tapestry - Coach and Kate Spade Logo Tapestry - Coach and Kate Spade

Temporary Sales Associate

eCommerce • Fashion • Retail • Sales • Wearables • Design
Hybrid
Bellevue, WA, USA
16000 Employees
15-24 Hourly

Similar Companies Hiring

Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software
New York, New York
30 Employees
Vega Thumbnail
Artificial Intelligence • Automotive • Insurance • Transportation
US
43 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account