Reinforcement Learning Infrastructure Engineer

Posted 6 Days Ago
Be an Early Applicant
Hiring Remotely in Office, Machaze, Manica, MOZ
Remote or Hybrid
275K-475K Annually
Mid level
Artificial Intelligence • Computer Vision • Machine Learning • Generative AI
The Role
Design and build end-to-end reinforcement learning training infrastructure: rollout and reward pipelines, actor-learner architectures, distributed orchestration, monitoring, and GPU utilization improvements, partnering with researchers to productionize RL training for multimodal models.
Summary Generated by Built In
About Us

We are a well-funded, early-stage AI lab focused on building the next generation of frontier multimodal AI models. Founded by former DeepMind researchers, including Andrew Dai, who was previously a leader on Gemini. Our team currently consists of 20 world-class scientists and engineers. We recently raised $55M in seed funding from Striker Ventures, Menlo Ventures, Altimeter Capital, and NVIDIA. We are tackling some of the hardest problems in artificial intelligence, and we are growing fast.

The Role

We're looking for an infrastructure engineer to design and build the core systems behind how we train our models with reinforcement learning (RL).

You'll own the training infrastructure end to end, from rollout and reward pipelines to orchestration, reliability, and observability. The work spans both the algorithmic side of RL and the systems reality of running distributed training at scale, and you'll partner closely with our research team to keep RL training fast, stable, and dependable for the multimodal, visual reasoning models at the center of our work.


What You Will Do
  • Design, build, and optimize the infrastructure that powers our large-scale RL and post-training workloads

  • Improve the reliability, scalability, and throughput of distributed RL training pipelines

  • Build actor-learner architectures and orchestrate environment rollouts at scale

  • Develop monitoring and observability tools that ensure high uptime, debuggability, and reproducibility across RL systems

  • Collaborate with researchers to translate algorithmic ideas into production-grade training pipelines

  • Improve GPU utilization and training throughput across the cluster

What We're Looking For

Minimum qualifications:

  • 3+ years of distributed systems experience, including building or optimizing large-scale RL training pipelines (PPO, GRPO, or similar on-policy methods)

  • Experience with actor-learner architectures and environment rollout orchestration at scale

  • Strong Python skills, plus PyTorch or JAX

  • Experience with async training infrastructure, replay buffers, or simulation-based environment frameworks

  • Multi-node GPU orchestration experience (Ray, SLURM, or Kubernetes)

  • A track record of improving training throughput and GPU utilization at scale

  • Strong engineering skills; ability to contribute performant, maintainable code and debug in complex codebases

Preferred qualifications (strong candidates may have some, not all):

  • Experience with multimodal or agentic RL environments

  • Experience with RLHF or reward modeling pipelines

  • A self-directed builder who moves quickly and works across teams in an early-stage setting

Logistics

Location: This role is based on-site in Palo Alto, California.

Compensation: Depending on background, skills, and experience, the expected annual base salary range for this position is $200,000 - $400,000 USD, plus equity and benefits.

Visa sponsorship: We sponsor work visas. We can't promise every case will succeed, but for the right person we'll work through the process with you.

Benefits: We offer health, dental, and vision benefits, unlimited PTO, paid parental leave, and relocation support as needed.

Elorian AI is an equal opportunity employer. We are committed to building a diverse team and inclusive environment.

Skills Required

  • 3+ years of distributed systems experience, including building or optimizing large-scale RL training pipelines (PPO, GRPO, or similar on-policy methods)
  • Experience with actor-learner architectures and environment rollout orchestration at scale
  • Strong Python skills, plus PyTorch or JAX
  • Experience with async training infrastructure, replay buffers, or simulation-based environment frameworks
  • Multi-node GPU orchestration experience (Ray, SLURM, or Kubernetes)
  • A track record of improving training throughput and GPU utilization at scale
  • Strong engineering skills; ability to contribute performant, maintainable code and debug in complex codebases
  • Experience with multimodal or agentic RL environments
  • Experience with RLHF or reward modeling pipelines
  • Self-directed builder who moves quickly and works across teams in an early-stage setting
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
18 Employees
Year Founded: 2025

What We Do

Elorian is an early-stage AI research lab founded by former DeepMind and Apple researchers. They are focused on building the next generation of frontier multimodal AI models with a specific emphasis on advanced visual reasoning. By moving beyond text-centric LLMs, they aim to enable AI to understand the spatial and structural complexity of the physical world, pushing the boundaries toward Artificial General Intelligence (AGI) for applications in robotics, science, and industry.

Similar Jobs

Treeswift Logo Treeswift

Field Operations Technician

Machine Learning • Robotics • Software
Remote or Hybrid
Office, Machaze, Manica, MOZ
27 Employees
90K-110K Annually

Auror Logo Auror

Director Of Product Management

Artificial Intelligence • Big Data • Retail • Security • Social Impact • Software • Business Intelligence
Remote or Hybrid
Office, Machaze, Manica, MOZ
212 Employees
175K-245K Annually

Auror Logo Auror

Enterprise Account Manager

Artificial Intelligence • Big Data • Retail • Security • Social Impact • Software • Business Intelligence
Remote or Hybrid
Office, Machaze, Manica, MOZ
212 Employees
68K-83K Annually

Auror Logo Auror

Finance Business Partner, UK

Artificial Intelligence • Big Data • Retail • Security • Social Impact • Software • Business Intelligence
Remote or Hybrid
Office, Machaze, Manica, MOZ
212 Employees
70K-80K Annually

Similar Companies Hiring

Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
LTX Thumbnail
Robotics • Conversational AI • Generative AI
Jerusalem, Israel
300 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account