What You Will Work On
Design and operate distributed training systems for large neural networks (autoregressive, diffusion, State Space Models etc.) across GPU clusters
Optimise multi‑node, multi‑GPU execution to maximize throughput and utilization
Diagnose & resolve bottlenecks across compute, memory, and network
Improve training stability and fault tolerance at scale
Partner with research and applied ML teams to productionize large‑model training pipelines
Core Responsibilities
Distributed Training Infrastructure
Build and optimize GPU cluster orchestration using:
Slurm
Kubernetes
Ray
RunAI
Ensure efficient scheduling, isolation, and fairness across training workloads
Communication & Networking
Optimize and debug distributed communication using:
NCCL
RDMA
InfiniBand
NVLink
Minimize networking bottlenecks that dominate end‑to‑end training time
Training Frameworks
Scale large-model training using:
PyTorch Distributed
Megatron‑LM
DeepSpeed
Own multi‑node launch configurations, failure recovery, and performance tuning
Memory & Performance Optimization
Apply advanced memory optimization techniques:
Activation checkpointing
ZeRO (Stage 1–3) and offload strategies
Balance compute, memory, and communication to push model size and batch scale
What Success Looks Like
GPU utilization consistently stays high (>80–90%)
Training scales cleanly from single node to dozens or hundreds of GPUs
Communication overhead is minimized and predictable
Large training jobs run stably for days or weeks without failure
New models can be trained faster, larger, and more reliably than before
Required Experience & Skills
Strongly Required
Deep hands‑on experience with distributed systems or ML systems
Experience running large‑scale workloads on GPU clusters
Production experience with PyTorch distributed training
Strong understanding of parallelism strategies (data, tensor, pipeline parallelism)
Low‑level understanding of GPU communication and networking
Critical Technical Skills
GPU orchestration: Slurm, Kubernetes, Ray, RunAI
Communication libraries: NCCL, RDMA, InfiniBand, NVLink
Training frameworks: PyTorch Distributed, Megatron‑LM, DeepSpeed
Memory optimisation: activation checkpointing, ZeRO offload techniques
Common Problems You’ll Be Solving
Many teams fail at scale because:
GPU utilization is low despite large clusters
Networking and communication dominate training time
Training jobs crash or become unstable at large scale
You will be explicitly focused on eliminating these failure modes.
Ideal Background
This role is a strong fit for individuals who have worked as:
ML Systems Engineer
Distributed Systems Engineer
AI Infrastructure Engineer
HPC Engineer transitioning into ML
Experience working with large language models or foundation models is a strong plus, but deep systems expertise is valued over pure model architecture experience.
Why This Role Matters
Without robust distributed training infrastructure, progress on large models stalls. This role directly enables:
Larger models
Faster iteration cycles
More reliable research-to-production pipelines
You will be building the foundation that makes large‑scale AI possible.
Cerence Inc. (Nasdaq: CRNC and www.cerence.com) is the global industry leader in creating unique, moving experiences for the automotive world. Spun out from Nuance in October 2019, Cerence is a new, independent company that has quickly gained traction as a leader in the automotive voice assistant space, working with all of the world’s leading automakers – from Ford and Fiat Chrysler to Daimler, Audi and BMW to Geely and SAIC – to transform how a car feels, responds and learns. Its track record is built on more than 20 years of industry experience and leadership and more than 500 million cars on the road today across more than 70 languages.
As Cerence looks to the future and continues an ambitious growth agenda, we need someone to join the team and help build the future of voice and AI in cars. This is an exciting opportunity to join Cerence’s passionate, dedicated, global team and be a part of meaningful innovation in a rapidly growing industry.
EQUAL OPPORTUNITY EMPLOYERCerence is firmly committed to Equal Employment Opportunity (EEO) and to compliance with all federal, state and local laws that prohibit employment discrimination on the basis of age, race, color, gender, gender identity, gender expression, sex, sex stereotyping, pregnancy, national origin, ancestry, religion, physical or mental disability, medical condition, marital status, citizenship status, sexual orientation, protected military or veteran status, genetic information and other protected classifications. Cerence Equal Employment Opportunity Policy Statement.
All prospective and current Employees need to remain vigilant when it comes to executing security policies in the workplace. This includes:
- Following workplace security protocols and training programs to familiarize with the ways to maintain a safe workplace.
- Following security procedures to report any suspicious activity.
- Having respect for corporate security procedures to allow those procedures to be effective.
- Adhering to company's compliance and regulations.
- Encouraging to follow a zero tolerance for workplace violence.
- Basic knowledge of information security and data privacy requirements (e.g., how to protect data & how to be handling this data).
- Demonstrative knowledge of information security through internal training programs.
Skills Required
- Deep hands-on experience with distributed systems or ML systems
- Experience running large-scale workloads on GPU clusters
- Production experience with PyTorch distributed training
- Strong understanding of parallelism strategies (data, tensor, pipeline parallelism)
- Low-level understanding of GPU communication and networking
- GPU orchestration: Slurm, Kubernetes, Ray, RunAI
- Communication libraries: NCCL, RDMA, InfiniBand, NVLink
- Training frameworks: PyTorch Distributed, Megatron-LM, DeepSpeed
- Memory optimisation: activation checkpointing, ZeRO offload techniques
- Experience working with large language models or foundation models
- Background as ML Systems Engineer, Distributed Systems Engineer, AI Infrastructure Engineer, or HPC Engineer
Cerence Inc. Compensation & Benefits Highlights
The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Cerence Inc. and has not been reviewed or approved by Cerence Inc..
-
Healthcare Strength — Healthcare is described as strong and affordable, with mentions of a great health care plan and wellness support. Feedback suggests these offerings contribute to work-life balance and peace of mind.
-
Wellbeing & Lifestyle Benefits — Lifestyle perks such as gym membership reimbursement and transit subsidies are highlighted as meaningful add-ons. Feedback suggests these perks enhance the overall rewards package.
-
Strong & Reliable Incentives — Bonuses and equity incentives, including spot awards, short-term incentive targets, and ESPP, are cited as part of total rewards. Feedback suggests these programs are a notable component beyond base pay.
Cerence Inc. Insights
What We Do
Cerence (NASDAQ: CRNC) is the global industry leader in creating unique, moving experiences for the mobility world. As an innovation partner to the world’s leading automakers and mobility OEMs, it is helping advance the future of connected mobility through intuitive, powerful interaction between humans and their cars, two-wheelers, and even elevators, connecting consumers’ digital lives to their daily journeys no matter where they are. Cerence’s track record is built on more than 20 years of knowledge and more than 400 million cars shipped with Cerence technology. Whether it’s connected cars, autonomous driving, e-vehicles, or buildings, Cerence is mapping the road ahead.







