ML Engineer L4, Consumer Inference

Posted 4 Days Ago
Los Gatos, CA, USA
In-Office
100K-464K Annually
Mid level
News + Entertainment
The Role
Develop and maintain customer-facing libraries and online inference services for low-latency, reliable real-time predictions. Optimize and deploy LLMs for efficient GPU inference, maintain a model registry, and improve ML Platform CI/CD, observability, and incident management. Collaborate cross-functionally to productize models and accelerate research-to-production velocity for Netflix ML practitioners.
Summary Generated by Built In

Netflix is one of the world’s leading entertainment services with 278 million paid memberships in over 190 countries enjoying TV series, films and games across a wide variety of genres and languages. Members can play, pause and resume watching as much as they want, anytime, anywhere, and can change their plans at any time.

The Role

With more than 230 million members in over 190 countries, Netflix continues to shape the future of entertainment around the world. Machine Learning/Artificial Intelligence is powering innovation in all areas of the business, from helping members choose the right title for them through personalization, to better understanding our audience and our content slate, to optimizing our payment processing and other revenue-focused initiatives. The Machine Learning Platform (MLP) provides the foundation for all of this innovation. It offers ML/AI practitioners across Netflix the means to achieve the highest possible impact with their work by making it easy to develop, deploy and improve their machine learning models.  As part of our mission to support the infrastructure for machine learning across the company, we are hiring for a Machine Learning Engineer to join our team to contribute to the team's mission of bridging the gap between ML research and productization. In this role, you will: -Develop customer facing libraries and services to productize machine learning models for efficient and scalable inference.  -Develop and maintain online inference services that provide real-time predictions with low latency and high reliability. -Optimize and deploy large language models (LLMs) for efficient, scalable inference, ensuring high performance and low latency in production environments. -Maintain and improve a model registry to facilitate the discovery, versioning, and governance of machine learning models. -Participate in and improve ML Platform incident management and support workflows. What we offer Opportunity for impact. You will work on cutting edge ML infrastructure use cases and technologies. This role will develop into strategic ownership opportunities for defining MLP’s path from research to production, specifically focusing on building services and tools to accelerate research to production velocity for Netflix ML practitioners.. Responsibility. Netflix offers true transparency and autonomy. Our culture is unique and is key to how we innovate. From day one, your expertise and opinion will be respected and valued by the team and you’ll be given autonomy in deciding the best direction to set for optimizing research to the production path for ML practitioners at Netflix. Learning. You will be developing libraries, tools and services to ensure an efficient and reliable journey of productizing ML models. You will have the opportunity to work with stunning colleagues who value collaboration and have a wealth of experience you can tap into. A work environment where you can grow your career. ML Platform offers a wide variety of projects that can help find the areas you are passionate about. Who will be successful in this role?
  • You are highly customer-driven / developer-driven and empathic. You strive to always focus on delivering customer / user value with an excellent customer service mentality.
  • You have a strong understanding of building scalable and efficient model serving solutions to support large-scale inference for generative models and large language models (LLMs). You create solutions that your stakeholders love and you drive development success from planning to implementation to delivery.
  • You can successfully execute changes within a team's systems, including developing, testing, deploying, and revising solutions.
  • You can communicate and collaborate effectively  (e.g. project meetings, team meetings, code reviews) with immediate team peers and cross-functional project teams.
  • You are eager to both go deep and wide on ML-facing projects. When a project needs deep technical expertise in a domain area you are able to get up to speed quickly. When projects require breadth of focus you are eager to do what’s needed to deliver value even if it means going outside of your comfort zone.
Skills:
  • Strong programming skills, particularly in languages such as Python and Java, and familiarity with ML libraries and frameworks like TensorFlow, PyTorch.
  • Familiarity tools and techniques for deploying machine learning models into production environments, with a particular emphasis on GPU inference optimization (e.g., Triton Inference Server, TensorRT), as well as containerization (e.g., Docker) and orchestration (e.g., Kubernetes).
  • Experience designing with data handling, preprocessing, and transformation techniques to prepare data for model inference.
  • Demonstrated industry-leading experience in large-scale build, release, CI/CD and observability techniques, with particular emphasis on multi-language environments including Scala, Java, and Python.
  • Adopt and promote best practices in operations, including observability, logging, reporting, and on-call processes to ensure engineering excellence.
Our compensation structure consists solely of an annual salary; we do not have bonuses. You choose each year how much of your compensation you want in salary versus stock options. To determine your personal top of market compensation, we rely on market indicators and consider your specific job family, background, skills, and experience to determine your compensation in the market range. The range for this role is $100,000 - $464,000. Netflix provides comprehensive benefits including Health Plans, Mental Health support, a 401(k) Retirement Plan with employer match, Stock Option Program, Disability Programs, Health Savings and Flexible Spending Accounts, Family-forming benefits, and Life and Serious Injury Benefits. We also offer paid leave of absence programs.  Full-time hourly employees accrue 35 days annually for paid time off to be used for vacation, holidays, and sick paid time off. Full-time salaried employees are immediately entitled to flexible time off. See more detail about our Benefits here. Netflix is a unique culture and environment.  Learn more here.

We are an equal-opportunity employer and celebrate diversity, recognizing that diversity of thought and background builds stronger teams. We approach diversity and inclusion seriously and thoughtfully. We do not discriminate on the basis of race, religion, color, ancestry, national origin, caste, sex, sexual orientation, gender, gender identity or expression, age, disability, medical condition, pregnancy, genetic makeup, marital status, or military service.

Skills Required

  • Strong programming skills in Python and Java
  • Experience with ML frameworks such as TensorFlow and PyTorch
  • Experience optimizing and deploying LLMs and large-scale models for GPU inference
  • Experience with Triton Inference Server and TensorRT (GPU inference optimization)
  • Experience with containerization and orchestration (Docker, Kubernetes)
  • Experience designing data handling, preprocessing, and transformation for model inference
  • Industry experience with build, release, CI/CD and observability in multi-language environments (Scala, Java, Python)
  • Adopt and promote operational best practices including observability, logging, reporting, and on-call processes
  • Strong collaboration and communication skills; customer/developer-driven mindset

Netflix Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Netflix and has not been reviewed or approved by Netflix.

  • Fair & Transparent Compensation Compensation is positioned as “personal top of market” with regular recalibration and broad posted ranges for senior roles that signal the philosophy. The cash‑forward structure and clearly described pay‑mix choices help set expectations on how pay is determined.
  • Equity Value & Accessibility Employees can choose the mix of cash versus fully vested 10‑year stock options, with grants structured to be retained even after departure. This employee‑directed design increases accessibility and control over equity participation.
  • Healthcare Strength Health coverage is described as comprehensive across medical, dental, vision, and mental health, with employer funding designed to offset premiums. Additional resources like counseling/coaching and wellness support reinforce breadth in care access.

Netflix Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Los Gatos, CA
13,212 Employees
Year Founded: 1997

What We Do

Netflix is the world's leading streaming entertainment service with 209 million paid memberships in over 190 countries enjoying TV series, documentaries and feature films across a wide variety of genres and languages. Members can watch as much as they want, anytime, anywhere, on any internet-connected screen. Members can play, pause and resume watching, all without commercials or commitments.

Similar Jobs

ServiceNow Logo ServiceNow

Senior Information Security Analyst

Artificial Intelligence • Cloud • HR Tech • Information Technology • Productivity • Software • Automation
Remote or Hybrid
Santa Clara, CA, USA
29000 Employees
128K-217K Annually

ServiceNow Logo ServiceNow

Product Manager

Artificial Intelligence • Cloud • HR Tech • Information Technology • Productivity • Software • Automation
Hybrid
Santa Clara, CA, USA
29000 Employees
191K-334K Annually

Crunchyroll Logo Crunchyroll

Senior Software Engineer

Digital Media • eCommerce • Gaming • Mobile • News + Entertainment
Hybrid
San Francisco, CA, USA
1300 Employees
203K-235K Annually

Nasuni Logo Nasuni

Senior Manager, Sales

Artificial Intelligence • Big Data • Cloud • Security • Software • Cybersecurity • Infrastructure as a Service (IaaS)
Easy Apply
Hybrid
8 Locations
550 Employees

Similar Companies Hiring

TIDAL Thumbnail
Software • News + Entertainment • Mobile • Information Technology • Music • Consumer Web
New York, NY
450 Employees
Sandbox VR Thumbnail
Events • Gaming • News + Entertainment • Retail • Virtual Reality
Tsim Sha Tsui East, Kowloon
650 Employees
Hedra Thumbnail
Software • News + Entertainment • Marketing Tech • Generative AI • Enterprise Web • Digital Media • Consumer Web
San Francisco, CA
14 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account