Research MLE (Training Optimization)大模型训练优化工程师

Reposted One Month Ago
Be an Early Applicant
Beijing, CHN
Hybrid
Senior level
Digital Media • Information Technology • Software • Design
The Role
Design, implement, and optimize large-scale distributed training systems for foundation and multimodal models. Improve GPU utilization, communication, and memory efficiency; write custom CUDA/Triton kernels; profile and debug training workflows; and collaborate with research teams to scale model training.
Summary Generated by Built In
Company Description

At Canva, we're building a future powered by AI that's as magical as it is impactful. As a Research Scientist at Canva, you'll be responsible for advancing the future of AI by experimenting with cutting-edge techniques, as well as improving models for real-world quality and performance.

Job Description

About the Group/Team

We're the CORE team within the Generative AI supergroup. Our mission is to invent foundational technologies that will power the future of AI-assisted design. From large-scale models to groundbreaking research, our team builds the technical core of Canva’s creative intelligence engine. We collaborate globally to ship research that makes a real impact—from smart editing to AI video tools—at massive scale.

 

About the Role/Specialty

As a Machine Learning Engineer, you’ll lead efforts to scale and optimize the training system for our large-scale multimodal and foundation models. You’ll design distributed training systems using Megatron-LM, NVIDIA NeMo, FSDP, and Triton—pushing the limits of performance across compute, memory, and communication layers. You'll sit at the intersection of systems and AI research, directly shaping how we train the models that will power Canva’s next generation of products.

 

What you’ll do (responsibilities)

  • You’ll design, implement, and optimize large-scale machine learning systems for training
  • You’ll improve all aspects of performance, including GPU utilization, communication overhead, and memory efficiency.
  • You’ll partner with research and modeling teams to align systems with algorithmic needs.
  • You’ll evaluate and apply best practices for distributed training using industry-leading frameworks.
  • You’ll dive deep into low-level optimization, including custom CUDA or Triton kernels.

• • You’ll debug, profile, and fine-tune training workflows to unlock new levels of scalability.

Qualifications

What we're looking for

We’re looking for a systems-first engineer who thrives in fast-paced, high-impact environments. You’re deeply familiar with distributed model training at scale and understand the nuances of optimizing compute at every level of the stack. You're excited by challenges that stretch current boundaries, and you’re a strong collaborator who communicates clearly across domains.

  • Strong background in LLMs, multimodal AI, or diffusion models.
  • Proficiency in Python. Familiarity with a system programming language (e.g. C++ or Rust) is a plus.
  • Deep knowledge of PyTorch or JAX as well as libraries such as Megatron-LM, NeMo, or DeepSpeed.
  • Familiarity with common optimization techniques such as FSDP/ZeRO, gradient checkpointing, or low-precision data types.
  • Hands-on experience writing custom GPU kernels in CUDA or Triton.
  • Excellent communication and problem-solving skills, incl. full proficiency in English.

Additional Information

大模型训练优化工程师(多模态/图像生成),技术要求:算子优化/分布式训练/GPU集群/训练框架。该岗位面向所有经验阶段的候选人开放,包括社会招聘、应届毕业生,同时开放实习生岗位。

Skills Required

  • Strong background in LLMs, multimodal AI, or diffusion models.
  • Proficiency in Python.
  • Familiarity with a systems programming language (C++ or Rust).
  • Deep knowledge of PyTorch or JAX.
  • Familiarity with Megatron-LM, NVIDIA NeMo, or DeepSpeed.
  • Familiarity with optimization techniques such as FSDP/ZeRO, gradient checkpointing, and low-precision data types.
  • Hands-on experience writing custom GPU kernels in CUDA or Triton.
  • Experience designing, implementing, and optimizing large-scale distributed training systems and GPU cluster workflows.
  • Ability to debug, profile, and fine-tune training workflows.
  • Excellent communication and problem-solving skills; full proficiency in English.

Canva Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Canva and has not been reviewed or approved by Canva.

  • Parental & Family Support Permanent employees are eligible for 18 weeks of paid parental leave from day one, inclusive of all parents. Additional supports include flexibility in how leave is taken and resources around pregnancy loss and caregiving.
  • Leave & Time Off Breadth Extra paid time off categories expand beyond standard PTO. Examples include five Flex Leave days, three days of paid volunteering, and milestone Epic Experience time off at 5 and 10 years.
  • Equity Value & Accessibility Equity participation is a core element of total rewards. Company materials and industry reporting note periodic opportunities to realize value through secondary transactions, enhancing perceived upside.

Canva Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Sydney
5,500 Employees
Year Founded: 2013

What We Do

Canva is an online graphic design platform with a mission to empower everyone to design anything and publish anywhere, offering a free-to-use tool for creating social media posts, presentations, posters, videos, logos, and more.

Similar Jobs

Ericsson Logo Ericsson

New Grad-Power Amplifier Developer-BJ

Cloud • Information Technology • Internet of Things • Machine Learning • Software • Cybersecurity • Infrastructure as a Service (IaaS)
In-Office
Beijing, CHN
88000 Employees

Ericsson Logo Ericsson

Hardware Engineer

Cloud • Information Technology • Internet of Things • Machine Learning • Software • Cybersecurity • Infrastructure as a Service (IaaS)
In-Office
Beijing, CHN
88000 Employees

Ericsson Logo Ericsson

New Grad-Radio Developer TRX-BJ

Cloud • Information Technology • Internet of Things • Machine Learning • Software • Cybersecurity • Infrastructure as a Service (IaaS)
In-Office
Beijing, CHN
88000 Employees

Ericsson Logo Ericsson

Network Engineer

Cloud • Information Technology • Internet of Things • Machine Learning • Software • Cybersecurity • Infrastructure as a Service (IaaS)
In-Office
2 Locations
88000 Employees

Similar Companies Hiring

Kepler  Thumbnail
Artificial Intelligence • Fintech • Software
New York, New York
9 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Revel.io Thumbnail
Aerospace • Hardware • Robotics • Software
US
50 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account