Senior Distributed Systems Engineer

Posted 11 Days Ago
Be an Early Applicant
Palo Alto, CA
180K-250K Annually
Senior level
Digital Media
The Role
The Senior Distributed Systems Engineer will collaborate with researchers to develop systems that efficiently train large models on massive GPU clusters. Responsibilities include optimizing code for hardware efficiency, distributing work across clusters, and handling hardware failures during model training.
Summary Generated by Built In

We are looking for people with strong ML & Distributed systems backgrounds. This role will work within our Research team, closely collaborating with researchers to build the platforms for training our next generation of foundation models.

Responsibilities

  • Work with researchers to scale up the systems required for our next generation of models trained on multi-thousand GPU clusters.
  • Profile and optimize our model training code-base to achieve best in class hardware efficiency.
  • Build systems to distribute work across massive GPU clusters efficiently.
  • Design and implement methods to robustly train models in the presence of hardware failures.
  • Build tooling to help us better understand problems in our largest training jobs.

Experience

  • 5+ years of work experience.
  • Experience working with multi-modal ML pipelines, high performance computing and/or low level systems.
  • Passion for diving deep into systems implementations and understanding their fundamentals in order to improve their performance and maintainability.
  • Experience building stable and highly efficient distributed systems.
  • Strong generalist Python and Software skills including significant experience with Pytorch.
  • Good to have experience working with high performance C++ or CUDA.
  • Please note this role is not meant for recent grads.

Compensation

  • The pay range for this position in California is $180,000 - $250,000yr; however, base pay offered may vary depending on job-related knowledge, skills, candidate location, and experience. We also offer competitive equity packages in the form of stock options and a comprehensive benefits plan. 

Your application is reviewed by real people.

Top Skills

C++
Cuda
Python
PyTorch
The Company
Minneapolis, MN
0 Employees
On-site Workplace

What We Do

Luma is a multimedia platform that delivers personalized movie and TV program selections from a range of sources to its viewers.

Similar Jobs

Snowflake Logo Snowflake

Senior Distributed Systems Engineer

Artificial Intelligence • Big Data • Cloud • Machine Learning • Software • Database • Analytics
San Mateo, CA, USA
7630 Employees
Easy Apply
Hybrid
Hollywood, Los Angeles, CA, USA
250 Employees
93K-120K Annually

Crunchyroll Logo Crunchyroll

Senior Software Engineer - Web Video Players

Digital Media • eCommerce • Gaming • Mobile • News + Entertainment
San Francisco, CA, USA
1200 Employees
185K-232K Annually

Atlassian Logo Atlassian

Principal Application Engineer, Marketing Technologies

Cloud • Information Technology • Productivity • Security • Software • App development • Automation
Remote
San Francisco, CA, USA
11000 Employees
171K-274K Annually

Similar Companies Hiring

Effectv Thumbnail
Marketing Tech • Digital Media • AdTech
New York, NY
2157 Employees
Artlist Thumbnail
Social Media • Other • Music • Digital Media
Tel Aviv, IL
450 Employees
bet365 Thumbnail
Software • Gaming • eSports • Digital Media • Automation
Denver, Colorado
6100 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account