Machine Learning Engineer, Infra (MLOps) - CN

Posted 17 Days Ago
Be an Early Applicant
Beijing, CHN
In-Office
Senior level
AdTech • Artificial Intelligence • Machine Learning • Marketing Tech • Software
The Role
Design, build, and operate scalable, reliable model training and deployment infrastructure. Build automated orchestration pipelines, ensure reproducibility and versioning, monitor data and model quality, optimize distributed training and online inference, apply MLOps/DevOps practices, and collaborate with algorithm engineers to support high-QPS advertising recommendation systems.
Summary Generated by Built In
Who Are We?

RZR is an AI-native advertising platform built for the next era of performance marketing. We operate at the intersection of machine learning, programmatic media, and full-funnel mobile growth, powering campaigns for some of the world's most ambitious advertisers. Our platform is purpose-built to deliver outcomes at scale, not just impressions.

We are a team of builders, operators, and technologists who believe the advertising industry is overdue for a fundamental rethink. We move fast, operate with a high degree of ownership, and hold ourselves to an exceptionally high standard of craft.

RZR is scaling aggressively with an active M&A pipeline and a platform vision that puts us on a path to becoming an industry leader. This is a rare opportunity to join a company at an inflection point and help shape what it becomes.

Role Overview

As Machine Learning Engineer (Infra / MLOps) at RZR, you will design, build, and operate the model training and deployment infrastructure that powers our Demand-Side Platform (DSP). This role focuses on building scalable, flexible, and reliable systems for training models on billions of records across bidding, ranking, pacing, and fraud use cases.

You will work at the intersection of machine learning, data platforms, and infrastructure — with a strong focus on automation, reproducibility, and reliability. This is a P0 priority hire directly tied to accelerating RZR's migration from legacy model training systems to Prefect-based DNN pipelines, enabling 100% UA on DNN.

The right person for this role combines production-grade ML systems experience with a strong bias to automate, document, and build for reliability — someone who takes end-to-end ownership from data to serving, and is energized by the complexity of high-QPS real-time bidding infrastructure.

Key Responsibilities
  • Own the development and evolution of infrastructure that enables faster, more reliable, and more cost-efficient model training
  • Design, build, and maintain automated model training and orchestration pipelines that scale across large datasets and support rapid recovery from failures
  • Develop standardized training workflows that support experimentation, reproducibility, versioning, and traceability
  • Build and operate observability and monitoring systems to detect data quality issues, training instabilities, model anomalies, and performance regressions
  • Improve the efficiency, scalability, and maintainability of the model training codebase, defining and enforcing best practices across the ML organization
  • Apply DevOps and MLOps best practices to machine learning training workflows, including CI/CD and automated testing
  • Design, develop, and continuously optimize ML infrastructure for advertising recommendation systems, covering model training, online inference, model serving, and feature pipelines
  • Build a high-performance, highly scalable ML platform to support rapid iteration and stable deployment of advertising recommendation models
  • Optimize distributed training, online inference, and resource scheduling to continuously improve system performance, stability, and resource utilization
  • Collaborate closely with algorithm engineers to drive efficient implementation of recommendation, ranking, and ad-serving models
  • Stay current with advancements in ML infrastructure and AI technologies, including the application of LLMs in recommendation and advertising scenarios
Required Skills and ExperienceMust-Have
  • Strong proficiency in Python and Spark for ML training and deployment workflows
  • Experience building and operating machine learning pipelines in production environments
  • Hands-on experience with DevOps practices including CI/CD, infrastructure as code, and automated testing
  • Experience with workflow orchestration tools such as Airflow or Prefect for ML pipelines
  • Solid understanding of ML experimentation, reproducibility, model versioning, and dataset management
  • Experience with large-scale data pipelines, feature generation, and offline/online data consistency
  • Experience developing recommendation systems, advertising systems, search systems, or machine learning platforms
  • Familiarity with mainstream ML frameworks such as PyTorch and TensorFlow
  • Experience with ML infrastructure, model training, online inference, or model serving
Nice-to-Have
  • Familiarity with system programming languages including C++ and Rust
  • Strong grasp of probability, statistics, and data analysis principles
  • Exposure to online inference systems, gRPC/REST model endpoints, or streaming features via Kafka or Flink
  • Ad-tech familiarity: auction dynamics, pacing, fraud signals, creative personalization
  • Experience with large-scale distributed training, high-performance computing (HPC), or GPU optimization
  • Familiarity with distributed computing frameworks such as Kubernetes, Ray, Spark, and Flink
  • Interest in or practical experience with LLMs and their application in recommendation and advertising scenarios
  • Experience with on-prem deployments of open source tools including Spark, ClickHouse, and Redash
  • Strong English reading and writing skills for collaboration with global teams
Why Join RZR?
  • End-to-end ownership across the full ML stack — data, features, training, evaluation, serving, A/B testing, and monitoring. You will not be working on one slice of the pipeline; you will shape all of it.
  • Real-time bidding and training pipelines at genuine scale — high QPS with tight latency SLOs. The infrastructure challenges here are not academic.
  • Shape the MLOps platform from the ground up — you will drive observability, data and model quality systems, and the MLflow-first platform, with direct influence on how the ML organization operates.
  • Mentorship and structured growth — paired with a senior ML engineer, with structured growth goals and a strong code review culture.
  • Immediate, measurable impact — your work will directly accelerate model iteration speed, improve feature quality, and improve offline/online metric alignment for RZR's core bidding and ranking systems.
  • Exposure to emerging AI technologies — RZR is actively exploring LLM applications in recommendation and advertising, and this role sits at the center of that work.
RZR Behaviors

RZR operates by eight core behaviors: Extreme Ownership · Move Fast · Drive for Excellence · Proactive Communication · Courage · Curiosity · Deliver Results · Manage Ambiguity

Skills Required

  • Strong proficiency in Python and Spark for ML training and deployment workflows
  • Experience building and operating machine learning pipelines in production environments
  • Hands-on experience with DevOps practices including CI/CD, infrastructure as code, and automated testing
  • Experience with workflow orchestration tools such as Airflow or Prefect for ML pipelines
  • Solid understanding of ML experimentation, reproducibility, model versioning, and dataset management
  • Experience with large-scale data pipelines, feature generation, and offline/online data consistency
  • Experience developing recommendation systems, advertising systems, search systems, or machine learning platforms
  • Familiarity with mainstream ML frameworks such as PyTorch and TensorFlow
  • Experience with ML infrastructure, model training, online inference, or model serving
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: San Francisco, California
65 Employees

What We Do

RZR powers performance for the world's most ambitious brands through proprietary neural architecture that optimizes across user acquisition, retargeting, CTV, and influencer campaigns as one connected performance system. With four owned-and-operated data centers that process 6M+ queries per second, RZR turns signals into strategy and impressions into impact. Trusted by brands across gaming, consumer, food and beverage, retail, and entertainment. Built on over a decade of performance data and backed by AI and ML experts and industry veterans across offices in San Francisco, New York, London, Bangalore, Beijing, Manila, and Seoul. RZR delivers retention-led growth intelligence: faster, sharper, and built for performance at scale.

Similar Jobs

CSC Logo CSC

Senior Manager Payroll & Corporate Secretarial Services

Fintech • Legal Tech • Software • Financial Services • Cybersecurity • Data Privacy
Hybrid
Beijing, CHN
8500 Employees

Snap Inc. Logo Snap Inc.

Senior Client Partner, Business Development

Artificial Intelligence • Cloud • Machine Learning • Mobile • Software • Virtual Reality • App development
Hybrid
2 Locations
5000 Employees

Taboola Logo Taboola

Account Manager

AdTech • Big Data • Digital Media • Marketing Tech
Hybrid
Beijing, CHN
1900 Employees

Taboola Logo Taboola

Marketing Manager

AdTech • Big Data • Digital Media • Marketing Tech
Hybrid
Beijing, CHN
1900 Employees

Similar Companies Hiring

Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees
Kepler  Thumbnail
Fintech • Software
New York, New York
6 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account