RL Environments Architect

Posted Yesterday
Be an Early Applicant
Hiring Remotely in United States
Remote
Entry level
Artificial Intelligence • Big Data • Machine Learning • Software
The Role
Design and govern reinforcement learning environments, including modular APIs, curricula, reward and termination schemas, quality standards, telemetry, and evaluation suites. Build robust simulations of real-world tasks, synthetic data generators, and multi-agent ecosystems while mitigating reward hacking, mode collapse, and exploitable loopholes. Partner with research and engineering teams to create deterministic, observable, scalable environments and maintain reproducible, high-quality training signals.
Summary Generated by Built In
About Us

Our mission is to raise AGI with the richness of human intelligence — curious, witty, imaginative, and full of unexpected brilliance.

Surge was founded by engineers and researchers who dreamed of building the next generation AI. We're building a platform that powers the most powerful models in the world in partnership with companies like Anthropic, Google, Microsoft, and Meta.

At Surge, we believe the path to AGI isn't just about scaling compute—it's about embracing the unlimited ceiling of human intelligence and creativity in the data that shapes these systems. Our platform combines elite human expertise with cutting-edge tools for scalable oversight, from building rich RL environments to conducting rigorous evaluations that go beyond benchmarks. We've run a profitable business from day one without raising venture funding.

The Role

As an RL Environments Architect, you’ll design, instrument, and govern the simulated worlds where agents learn — from compact task microcosms to multi-agent, tool-using ecosystems. You’ll define the primitives, reward structures, interfaces, and telemetry that let us stress-test emerging capabilities while keeping training signals faithful, stable, and scalable.

Not only will you build environments, you’ll craft standards for data quality and reproducibility across large-scale agent gyms. This is a role for someone who sweats the details of simulation fidelity, thinks in terms of coverage and failure surfaces, and loves turning messy real-world phenomena into learnable curricula. Your work will form the backbone for safe, rapid progress in agentic systems.

What You'll Do
  • Architect a modular environment framework with clear APIs, curriculum scaffolds, and configurable reward/termination schemas

  • Establish quality bars: coverage metrics, invariance checks, and trace audits for environment outputs and agent experience buffers

  • Instrument rich telemetry for episode rollouts; mitigating reward hacking, mode collapse, and exploitable loopholes

  • Partner with researchers to translate real-world tasks into robust simulations, including synthetic data generators and evaluation suites

What We’re Looking for
  • Simulation & Systems Depth – Experience building RL environments or simulators (e.g., custom physics, multi-agent, tool APIs) with an eye for determinism, performance, and observability

  • Data Quality Leadership – Strong instincts for designing reward functions, scenario taxonomies, and QA pipelines that keep signals aligned and drift-free

  • Builder’s Mindset – Comfort collaborating across research and engineering to ship pragmatic, testable environments that evolve with model capabilities

Skills Required

  • Experience building reinforcement learning environments or simulators, including custom physics, multi-agent systems, or tool APIs, with focus on determinism, performance, and observability.
  • Ability to design reward functions, scenario taxonomies, coverage metrics, invariance checks, trace audits, and QA pipelines for data quality and signal alignment.
  • Comfort collaborating across research and engineering teams to ship pragmatic, testable environments that evolve with model capabilities.
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
485 Employees
Year Founded: 2020

What We Do

Surge AI is an artificial intelligence company focused on providing high-skill data-labeling services and technology for AI development. It combines a data-labeling platform with a specialized workforce supporting natural-language processing, code generation, search evaluation, adversarial training, and related machine-learning applications. The company’s broader mission is to help raise artificial general intelligence with the richness, curiosity, and creativity of human intelligence.

Similar Jobs

SentiLink Logo SentiLink

Product Leader, Prefill

Fintech • Information Technology • Software
Remote
United States
170 Employees
180K-230K Annually

Zapier Logo Zapier

Growth Marketing Design Lead

Artificial Intelligence • Productivity • Software • Automation
Remote
2 Locations
800 Employees
192K-287K Annually

Applied Systems Logo Applied Systems

Director, Product Management

Artificial Intelligence • Cloud • Payments • Software • Business Intelligence • Generative AI • Automation
Remote or Hybrid
United States
3116 Employees
150K-220K Annually

Applied Systems Logo Applied Systems

User Experience Designer

Artificial Intelligence • Cloud • Payments • Software • Business Intelligence • Generative AI • Automation
Remote or Hybrid
2 Locations
3116 Employees
120K-175K Annually

Similar Companies Hiring

Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Revel.io Thumbnail
Aerospace • Hardware • Robotics • Software
US
50 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software
New York, New York
30 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account