Senior Machine Learning Infrastructure Engineer (Precision, Diagnostics & Hardware)

Posted 4 Hours Ago
3 Locations
Remote or Hybrid
Senior level
Artificial Intelligence • Computer Vision • Machine Learning • Robotics
The Role
Design, optimize, and maintain large-scale distributed multi-GPU training systems. Diagnose numerical precision and stability issues (FP16/BF16/FP8), detect and recover from hardware/software faults, profile performance bottlenecks, and build developer tooling for resilient checkpointing and rapid fault recovery to maximize researcher productivity and hardware utilization.
Summary Generated by Built In
About Us

Veeda AI is building the next generation of multimodal foundation world models for Physical AI. We're a small, fast-moving team of engineers and researchers from leading AI labs, tackling some of the most challenging problems at the intersection of AI, robotics, and embodied intelligence. If you're excited about pushing the boundaries of what's possible with Physical AI, you'll have the opportunity to make an outsized impact from day one.

Responsibilities
  • Distributed Training Systems & Scalability: Design, optimize, and maintain high-throughput distributed training systems across large-scale GPU clusters for multi-modal foundation models.

  • Precision & Numerical Stability: Debug, diagnose, and resolve subtle numerical instability issues (underflow/overflow, loss spikes, gradient explosion, and mixed-precision divergence) in FP16, BF16, FP8, and custom quantization schemes.

  • Fault Diagnostics & Recovery: Build advanced fault-detection mechanisms and automated diagnostics to rapidly pinpoint and isolate silent data corruption (SDC), hardware hang/deadlock, memory leaks, and "card-freeze" issues during large training runs.

  • Performance Profiling & Optimization: Profile distributed communication bottlenecks, memory usage, and kernel execution to improve overall FLOPS utilization across multi-node, multi-GPU training jobs.

  • Developer Tooling & Infrastructure: Develop resilient checkpointing systems, rapid fault-recovery pipelines, and execution telemetry to keep researcher productivity high and hardware downtime minimal.

Requirements
  • You have a Bachelor's degree or equivalent hands-on experience in Computer Science, Computer Engineering, or a related technical field.

  • You have deep hands-on experience with deep learning training frameworks (e.g., PyTorch) and distributed training paradigms (FSDP, Megatron-LM, DeepSpeed, Tensor Parallelism, Pipeline Parallelism).

  • You have proven experience in numerical precision analysis, low-precision training (BF16/FP8), and debugging complex loss divergence/stability issues in massive training runs.

  • You have strong root-cause analysis skills for hardware/software interaction bugs, including stuck CUDA kernels, NCCL timeouts, GPU hardware faults, and silent training corruptions.

  • You have strong programming skills in Python and C++/CUDA, with a deep understanding of low-level GPU architectures and memory hierarchies.

Nice to Have
  • You have experience running or porting large-scale training workloads on AMD GPUs (ROCm platform) or Google TPUs (JAX/XLA stack).

  • You have contributed to low-level training infrastructure, custom CUDA/Triton kernels, or distributed training open-source projects.

  • You have built resilient fault-tolerant training frameworks with dynamic node re-queueing and rapid checkpointing/saving mechanisms.

Skills Required

  • Bachelor's degree or equivalent hands-on experience in Computer Science, Computer Engineering, or related technical field
  • Deep hands-on experience with deep learning training frameworks (PyTorch) and distributed training paradigms (FSDP, Megatron-LM, DeepSpeed, tensor and pipeline parallelism)
  • Proven experience in numerical precision analysis and low-precision training (BF16/FP8) and debugging loss divergence/stability issues
  • Strong root-cause analysis skills for hardware/software interaction bugs including stuck CUDA kernels, NCCL timeouts, GPU hardware faults, and silent data corruption
  • Strong programming skills in Python and C++/CUDA with deep understanding of low-level GPU architectures and memory hierarchies
  • Experience running or porting large-scale training workloads on AMD GPUs (ROCm) or Google TPUs (JAX/XLA stack)
  • Contributions to low-level training infrastructure, custom CUDA/Triton kernels, or distributed training open-source projects
  • Built resilient fault-tolerant training frameworks with dynamic node re-queueing and rapid checkpointing/saving mechanisms
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company

What We Do

Veeda AI is a small, fast-moving team of engineers and researchers building the next generation of multimodal foundation world models for Physical AI. Its work sits at the intersection of artificial intelligence, robotics, and embodied intelligence, with engineering roles involving high-throughput image and video data pipelines. The company aims to advance intelligent systems capable of operating in and understanding the physical world.

Similar Jobs

Tulip Logo Tulip

Customer Success Manager

Enterprise Web • Hardware • Internet of Things • Software
Easy Apply
Remote or Hybrid
27 Locations
310 Employees

The Aerospace Corporation Logo The Aerospace Corporation

Director Sales

Aerospace • Artificial Intelligence • Cloud • Machine Learning • Software • Cybersecurity • Defense
Remote or Hybrid
4 Locations
4600 Employees

Zapier Logo Zapier

Back-end Engineer

Artificial Intelligence • Productivity • Software • Automation
Remote
32 Locations
800 Employees
211K-316K Annually

Zapier Logo Zapier

Systems Engineer

Artificial Intelligence • Productivity • Software • Automation
Remote
27 Locations
800 Employees

Similar Companies Hiring

Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
LTX Thumbnail
Robotics • Conversational AI • Generative AI
Jerusalem, Israel
200 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account