Research Engineer - Midtraining

Posted 2 Days Ago
Be an Early Applicant
Menlo Park, CA, USA
In-Office
250K-350K Annually
Mid level
Artificial Intelligence • Hardware • Information Technology • Robotics
From bits to atoms.
The Role
Improve scientific reasoning of foundation models by curating and generating scientific data, building evaluations, and running large-scale mid-training experiments. Apply techniques like self- and on-policy distillation, calculate scaling laws and compute-optimal hyperparameters, and collaborate with RL researchers, physicists, chemists, and supercompute engineers to scale training across thousands of GPUs.
Summary Generated by Built In

We're an AI and physical sciences company building state-of-the-art models to accelerate breakthroughs across materials, energy, and beyond. Backed by world-class investors and growing rapidly, we operate at the pace the frontier requires. Our team brings deep expertise, genuine ownership, and a drive to push the boundaries of what's scientifically possible.

About the Role

We're training frontier models to develop deep scientific knowledge and reasoning for scientific discovery. As a Midtraining Research Engineer, you'll take base models and improve their scientific reasoning: curating and generating data, building evals, and running large-scale training experiments. Your work will also lay the groundwork for our pre-training efforts down the line.

What You'll Do
  • Identify, process, and curate novel sources of scientific data for large-scale model training.

  • Generate high-quality synthetic data to fill gaps in scientific knowledge and reasoning.

  • Build evaluations that correlate with downstream scientific task performance, working closely with RL researchers, physicists, and chemists.

  • Develop and apply techniques such as self-distillation and on-policy distillation to improve model capability.

  • Design and run large-scale training experiments, partnering with supercompute engineers to scale efficiently across thousands of GPUs.

  • Build tools for yourself and the team to investigate how data choices shape model intelligence.

You Will Thrive in This Role If You Have
  • Experience training LLMs on curated mixes of trillions of tokens.

  • Experience on a dedicated evals team supporting a large production training run.

  • Hands-on use of self-distillation, on-policy distillation, or similar methods in a real training pipeline.

  • Experience with scaling laws and compute-optimal hyperparameters.

  • Comfort working across data, evals, and training infrastructure.

Especially Strong Candidates May Also Have
  • Experience optimizing throughput and reliability for large-scale distributed training runs.

  • A background in AI for science or training on specialized domain data (e.g., protein, materials, or other scientific datasets).

  • Experience creating evals or synthetic data for non verifiable tasks and tracking performance over live runs.

Mechanics
  • Minimum education: Bachelor's degree or similar experience

  • Location: Menlo Park, CA (Soon: San Francisco, too)

  • Compensation: $250,000–$350,000 + equity

  • Visa sponsorship: Yes, we sponsor visas and will do everything we can to assist in this process.

Skills Required

  • Bachelor's degree or equivalent experience
  • Experience training LLMs on curated mixes of trillions of tokens
  • Experience with mid-training or pre-training at scale (big-lab experience a plus)
  • Experience on a dedicated evals team supporting a large production training run
  • Hands-on use of self-distillation, on-policy distillation, or similar methods in a real training pipeline
  • Ability to calculate scaling laws and compute-optimal hyperparameters
  • Comfort working across data, evals, and training infrastructure
  • Experience optimizing throughput and reliability for large-scale distributed training runs
  • Background in AI for science or training on specialized domain data (protein, materials, scientific datasets)
  • Experience tracking evals and driving interventions during live big training runs
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
32 Employees
Year Founded: 2025

What We Do

We're building AI scientists and the autonomous laboratories for them to operate.

Similar Jobs

Ericsson Logo Ericsson

Senior Product Manager

Cloud • Information Technology • Internet of Things • Machine Learning • Software • Cybersecurity • Infrastructure as a Service (IaaS)
In-Office
2 Locations
88000 Employees
177K-221K Annually

UL Solutions Logo UL Solutions

Software Engineer

Automotive • Professional Services • Software • Consulting • Energy • Chemical • Renewable Energy
Hybrid
Fremont, CA, USA
15000 Employees
84K-93K Annually

Square Logo Square

Mid-market Account Executive

eCommerce • Fintech • Hardware • Payments • Software • Financial Services
Hybrid
Los Angeles, CA, USA
12000 Employees
130K-234K Annually

Square Logo Square

Manager, Mid-Market Sales

eCommerce • Fintech • Hardware • Payments • Software • Financial Services
Hybrid
Los Angeles, CA, USA
12000 Employees
214K-377K Annually

Similar Companies Hiring

Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees
LTX Thumbnail
Robotics • Conversational AI • Generative AI
Jerusalem, Israel
200 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account