LLM Pre-training & Distributed Engineer (AI Infrastructure)

Reposted 19 Days Ago
Hiring Remotely in Oregon, USA
Remote
Senior level
Agency • Artificial Intelligence • Blockchain • Web3
The Role
Operate and optimize large-scale LLM pre-training on 1,000+ GPU clusters using PyTorch, DeepSpeed, or Megatron-LM. Improve networking (InfiniBand/RDMA), memory management, checkpointing, and failure recovery. Manage SLURM/Kubernetes GPU clusters and apply systems engineering (C++, CUDA, Python) and 3D parallelism techniques.
Summary Generated by Built In

We are seeking a highly skilled LLM Pre-training & Distributed Systems Engineer. This role is essential for orchestrating large-scale machine learning training runs and optimizing  distributed infrastructure. The ideal candidate will have a deep understanding of GPU clusters and extensive experience in system engineering to ensure efficient and reliable training processes.

Responsibilities:

  • Orchestrate distributed training runs across 1,000+ GPUs using PyTorch, DeepSpeed, or Megatron-LM.
  • Optimize networking (InfiniBand/RDMA) and memory management to prevent out-of-memory errors.
  • Automate checkpointing and failure recovery during month-long training runs.

Required Skills:

  • Deep expertise in 3D parallelism (Data, Tensor, Pipeline).
  • Experience managing SLURM or Kubernetes-based GPU clusters.
  • Strong systems engineering background (C++, CUDA, Python).

Skills Required

  • Experience orchestrating distributed training runs across 1,000+ GPUs using PyTorch, DeepSpeed, or Megatron-LM
  • Deep expertise in 3D parallelism (Data, Tensor, Pipeline)
  • Experience optimizing networking (InfiniBand/RDMA) and memory management for large-scale training
  • Experience automating checkpointing and failure recovery for long-running training jobs
  • Experience managing SLURM or Kubernetes-based GPU clusters
  • Strong systems engineering background with C++, CUDA, and Python
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
7 Employees
Year Founded: 2024

What We Do

Hyphen Connect is a Web3 and AI talent agency and crypto-integrated software solutions provider that connects blockchain, DeFi, NFT, and AI companies with specialized technical and go-to-market talent globally and remotely. They deliver headhunting, data-driven research, and recruitment services across infrastructure, exchanges, gaming, and DeFi projects, plus industry analysis and hiring insights to help clients build engineering and product teams.

Similar Jobs

Elevate Leadership Logo Elevate Leadership

Sales Development Representative

HR Tech • Professional Services • Sales • Consulting
Remote
United States
14 Employees

LogicGate Logo LogicGate

Technical Account Manager

Cloud • Information Technology • Security • Software
Easy Apply
Remote
United States
202 Employees
100K-125K Annually

Zapier Logo Zapier

Sr. Manager, Strategic Finance, Self Serve & New Products

Artificial Intelligence • Productivity • Software • Automation
Remote
2 Locations
800 Employees
192K-287K Annually

Alloy Logo Alloy

Account Executive

Fintech • Information Technology • Software • Financial Services
Easy Apply
Remote or Hybrid
USA
315 Employees
300K-350K Annually

Similar Companies Hiring

Legora Thumbnail
Artificial Intelligence • Legal Tech • Software
New York, New York
700 Employees
Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account