Founding Research Engineer

Posted 17 Days Ago
Be an Early Applicant
2 Locations
In-Office
215K-330K Annually
Mid level
Artificial Intelligence • Productivity • Software • Automation
The Role
Build a scalable replayable-environment engine from enterprise historical data to train, evaluate, and improve agents. Own environment factory, reward/grade design, task mining, eval set creation, agent training/evaluation, replay/observability, and production-scale infrastructure.
Summary Generated by Built In
What we do

Ambral helps enterprises own the intelligence behind their most important workflows.

Every company has years of historical evidence showing how work gets done: the context people had, the decisions they made, the actions they took, and the outcomes that followed. Today, most of that history is inert. It isn’t structured in a way that companies can use to evaluate models and improve agent behavior.

Ambral turns this history into replayable environments and eval sets grounded in real workflows and observed outcomes. We use those environments to improve model performance through reinforcement learning and other post-training techniques, alongside context engineering, harness design, and agent engineering.

The result is better, more cost-efficient AI for each enterprise’s specific work, powered by open-weight models that the company owns and controls. This allows each company to retain ownership of its core workflow intelligence instead of outsourcing it to a model provider.
We graduated from YC S2025, raised millions in funding, and are already deployed within multi-billion dollar enterprises. Now we're growing the founding team.

What you’ll do

We’re building a replayable environment engine over real enterprise history.

The system reconstructs a company’s context as it existed at any past time, then exposes that state through the same tools an agent would use in production. This lets us place new policies and agent configurations inside real historical environments, observe how they reason and act, and grade their performance against real outcomes.

You’ll own the research and infrastructure required to turn this into a scalable model-improvement system. The core problems include:

  • Building an environment factory that converts recorded enterprise data and task definitions into runnable environments

  • Designing graders that turn ambiguous business objectives into verifiable rewards

  • Developing methods for mining useful tasks, trajectories, and evaluation cases from historical workflows

  • Creating eval sets that are representative, reproducible, and resistant to overfitting

  • Finding the right combinations of models, tools, context, and policies to maximize performance while reducing inference cost

  • Training and evaluating agents that operate over long horizons, incomplete information, and large tool spaces

  • Building replay and observability systems that make agent behavior explainable and measurable

  • Scaling from individual environments to thousands of concurrent training and evaluation runs

These problems are wide open. You’ll have significant ownership over both the research direction and the production systems that make it real.

You’ll work directly with the CTO, deploy into real enterprise workflows, and see your research tested against consequential problems and observable outcomes.

Who you are
  • You have 4+ years of experience building production software or machine-learning systems, including at least 2 years working on reinforcement-learning environments, LLM post-training, evaluation infrastructure, agent harnesses, or closely related systems

  • You understand how environment design, reward design, context, tooling, and policy behavior interact

  • You’re comfortable turning fuzzy business objectives into tasks and signals that can be evaluated reliably

  • You can diagnose whether a model’s limitations come from the model itself, its context, its tools, its harness, or its training

  • You can move between research questions and production implementation without treating them as separate jobs

  • You write strong software and can build systems that process large, messy datasets at scale

  • You care about reproducibility, observability, and understanding why a model behaves the way it does

  • You’re looking to do the best work of your life and build something you’ll be proud of for decades

We care much more about what you’ve built and how you think than credentials or conventional career paths.

Benefits
  • Significant equity and ownership

  • Equinox membership

  • Free meals, coffee, and snacks

  • Health insurance

  • Unlimited PTO

Skills Required

  • 4+ years building production software or machine-learning systems
  • At least 2 years working on reinforcement-learning environments, LLM post-training, evaluation infrastructure, or agent harnesses
  • Experience designing environments, reward design, context, tooling, and policy interactions
  • Ability to convert fuzzy business objectives into measurable tasks and evaluation signals
  • Skill diagnosing model limitations across model, context, tools, harness, and training
  • Ability to move between research questions and production implementation
  • Strong software engineering skills and experience building systems that process large, messy datasets at scale
  • Focus on reproducibility, observability, and explainability of model/agent behavior
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
2 Employees
Year Founded: 2025

What We Do

Ambral provides AI agents for enterprise account management and customer success, aiming to turn under-managed accounts into growth opportunities.

Similar Jobs

CrowdStrike Logo CrowdStrike

Security Engineer

Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity
Remote or Hybrid
USA
11000 Employees
120K-180K Annually

CrowdStrike Logo CrowdStrike

Senior Consultant

Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity
Remote or Hybrid
USA
11000 Employees
115K-160K Annually

CrowdStrike Logo CrowdStrike

Sr. Cloud Threat Hunter - AWS (Remote)

Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity
Remote or Hybrid
USA
11000 Employees
115K-160K Annually

CrowdStrike Logo CrowdStrike

Consultant

Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity
Remote or Hybrid
USA
11000 Employees
115K-160K Annually

Similar Companies Hiring

Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
Kepler  Thumbnail
Fintech • Software
New York, New York
6 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account