ML Systems Performance Engineer (MFU)

Posted An Hour Ago
Be an Early Applicant
Almaty, KAZ
In-Office
Entry level
Artificial Intelligence • Marketing Tech • Software • Generative AI
The Role
Profile and optimize distributed AI training across compute, memory, communication, storage, and orchestration. Improve MFU, throughput, scaling efficiency, goodput, GPU uptime, parallelism, collective communication, data loading, checkpointing, fault tolerance, and reliability. Develop or integrate CUDA and Triton kernels, diagnose distributed hangs and performance regressions, and optimize multi-GPU and multi-node training systems.
Summary Generated by Built In

Why work at Higgsfield AI?

Higgsfield AI is the fastest-scaling generative AI company in history, hitting $500M in annual revenue run rate, 25M+ users worldwide, 6M+ generations per day, and powering 390 of Fortune 500 brands. We're building at the absolute frontier of AI-powered video creation and next-generation creative tools. Joining Higgsfield means becoming part of a high-impact team shaping the future of AI-native experiences, at a company that isn't just moving fast, but rewriting what fast looks like.

What you will do

• Profile end-to-end training runs and identify bottlenecks across compute, memory, communication, storage, and orchestration.

• Define, measure, and improve MFU, tokens/sec/GPU, scaling efficiency, training goodput, and GPU uptime.

• Optimize distributed training and model-sharding strategies, including data, tensor, pipeline, context, and expert parallelism.

• Improve collective communication through topology-aware placement and compute/communication overlap.

• Develop or integrate optimized CUDA and Triton kernels

• Optimize data loading, preprocessing, sequence packing, and checkpointing so that I/O does not leave accelerators idle.

• Diagnose distributed hangs фтв performance regressions.

• Improve fault tolerance for long-running training jobs.

 

What we are looking for

• Strong experience running and optimizing multi-GPU or multi-node training.

• Experience with PyTorch Distributed or an equivalent training framework.

• Understanding of GPU architecture, including memory hierarchy, Tensor Cores

• Understanding of collective communication, cluster topology, and distributed-training bottlenecks.

• Experience with distributed parallelism technologies such as FSDP, DeepSpeed, Megatron-LM, TorchTitan, or similar.

• Ability to debug complex performance and reliability problems across multiple layers of the training stack.

 

Nice to have

• CUDA, Triton or GPU-kernel development experience.

• Experience with NCCL, MPI, UCX, RDMA, InfiniBand, RoCE, GPUDirect, NVLink, or NVSwitch.

• Experience training Mixture-of-Experts, multimodal, or reinforcement-learning models.

• Knowledge of PyTorch internals, torch.compile, XLA, ML compilers, or custom operators.

• Experience with mixed-precision training, including BF16, FP8, or FP4.

What We Offer

  • Competitive base salary in USD, based on your experience, skills, and the scope of the role.

  • Equity participation through the company’s stock option program, giving you the opportunity to share in Higgsfield’s long-term growth.

  • Relocation support to Almaty for candidates moving from another city or country.

  • A highly collaborative, fast-paced environment where you can work directly with experienced leaders and have a meaningful impact on the product and company.

  • Opportunities for professional growth, ownership, and career development as the company scales.

  • Company-provided equipment, meals, transportation, or other office benefits.

This is a fully on-site role based in our Almaty office. Our team works from the office five days per week for the full working day. We believe in-person collaboration is an important part of how we move quickly, solve complex problems, and build strong teams.

Skills Required

  • Strong experience running and optimizing multi-GPU or multi-node training
  • Experience with PyTorch Distributed or an equivalent training framework
  • Understanding of GPU architecture, including memory hierarchy and Tensor Cores
  • Understanding of collective communication, cluster topology, and distributed-training bottlenecks
  • Experience with distributed parallelism technologies such as FSDP, DeepSpeed, Megatron-LM, TorchTitan, or similar
  • Ability to debug complex performance and reliability problems across multiple layers of the training stack
  • CUDA, Triton, or GPU-kernel development experience
  • Experience with NCCL, MPI, UCX, RDMA, InfiniBand, RoCE, GPUDirect, NVLink, or NVSwitch
  • Experience training Mixture-of-Experts, multimodal, or reinforcement-learning models
  • Knowledge of PyTorch internals, torch.compile, XLA, ML compilers, or custom operators
  • Experience with mixed-precision training, including BF16, FP8, or FP4
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
70 Employees
Year Founded: 2023

What We Do

Higgsfield is an AI-native generative video platform that provides infrastructure for AI video and image generation, enabling creators, brands, and marketing teams to produce cinematic-quality visuals at scale by integrating various AI models into a unified workflow.

Similar Jobs

Mondelēz International Logo Mondelēz International

Statutory Reporting Analyst

Big Data • Food • Hardware • Machine Learning • Retail • Automation • Manufacturing
Hybrid
Almaty, KAZ
90000 Employees

Mondelēz International Logo Mondelēz International

Analyst, Accounting & External Reporting, Indirect Tax

Big Data • Food • Hardware • Machine Learning • Retail • Automation • Manufacturing
Hybrid
Almaty, KAZ
90000 Employees

Academy of Digital Industries Logo Academy of Digital Industries

Expansion Manager KZ

Digital Media • Edtech • Design
In-Office or Remote
2 Locations
150 Employees

AVIAREPS Group Logo AVIAREPS Group

Sales Manager

Agency • Professional Services • Travel • Hospitality
In-Office
Almaty, KAZ
800 Employees

Similar Companies Hiring

Kepler  Thumbnail
Fintech • Software
New York, New York
6 Employees
LTX Thumbnail
Robotics • Conversational AI • Generative AI
Jerusalem, Israel
200 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account