Research Scientist, Data

Posted Yesterday
Be an Early Applicant
Menlo Park, CA, USA
In-Office
250K-350K Annually
Mid level
Artificial Intelligence • Hardware • Information Technology • Robotics
From bits to atoms.
The Role
Lead evaluation and data strategy for scientific AI: design benchmarks and RL environments, source and integrate external and experimental datasets, build pipelines and tooling, and collaborate with domain and ML researchers to iterate model training and evaluation.
Summary Generated by Built In
About Periodic Labs

The most important scientific discoveries of our time won’t happen in a traditional lab. We’re an AI and physical sciences company building state-of-the-art models to accelerate breakthroughs across materials, energy, and beyond. Backed by world-class investors and growing rapidly, we operate at the pace the frontier requires. Our team brings deep expertise, genuine ownership, and an insatiable drive to push the boundaries of what’s scientifically possible.

About the Role

You will work on the most important aspect of Scientific AI creation: evaluations and data. This means constructing cutting-edge evaluations based on advanced scientific use cases, sourcing and procuring external datasets, integrating internally generated experimental data into the training stack, constructing training environments for RL. You’ll ensure that the team always has the right assets, in the right shape, to evaluate and improve AI models.

You will work with computational and experimental scientists to translate complex scientific workflows into rigorous evaluations and agentic benchmarks, and partner with pretraining, midtraining, and reinforcement learning researchers to identify the data models needed, then build the datasets, environments, and pipelines to deliver it. Your goal will be to create a tight feedback loop between scientific use cases, model evaluation, and training data.

 
What You’ll Do
  • Own the evaluation and data strategy across the training stack, identifying capability gaps and shaping the roadmap with leads of physical science and AI research

  • Work with domain experts to translate advanced scientific workflows into rigorous evals, benchmarks, and RL environments

  • Source, evaluate, and procure external datasets across chemistry, physics, materials science, mathematics, simulations, and laboratory instrumentation

  • Build robust pipelines to ingest, clean, and transform for training large-scale datasets from heterogeneous sources

  • Build tooling and analysis workflows that help researchers inspect data, understand model failures, and determine which evaluations or datasets to develop next

You Will Thrive in This Role If You Have
  • Designed evaluations, benchmarks, or RL environments for language models, agents, or scientific AI systems

  • Built large-scale data pipelines for LLM pretraining, midtraining, post-training, or evaluation

  • Strong judgment about dataset and evaluation quality, including scientific relevance, coverage, provenance, licensing, and contamination risks

  • Strong software and data engineering skills, including familiarity with data processing at scale, dataset versioning, lineage tracking

  • A research-oriented mindset: you form hypotheses about data, run controlled experiments, measure model outcomes, and iterate with rigor

Mechanics

Minimum education: Bachelor’s degree or similar experience

Location: Menlo Park, CA or Montreal, Canada. (Soon: San Francisco, too)

Compensation: $250,000-350,000 + equity

Visa sponsorship: Yes, we sponsor visas.

Skills Required

  • Bachelor's degree or equivalent experience
  • Designed evaluations, benchmarks, or RL environments for language models, agents, or scientific AI systems
  • Built large-scale data pipelines for LLM pretraining, midtraining, post-training, or evaluation
  • Strong judgment about dataset and evaluation quality, including provenance, licensing, contamination risks
  • Strong software and data engineering skills, including data processing at scale, dataset versioning, lineage tracking
  • Research-oriented mindset: form hypotheses, run controlled experiments, measure outcomes, iterate rigorously
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
32 Employees
Year Founded: 2025

What We Do

We're building AI scientists and the autonomous laboratories for them to operate.

Similar Jobs

HealthLeap Inc. Logo HealthLeap Inc.

Data Scientist

Artificial Intelligence • Healthtech • Machine Learning • Software
Hybrid
San Francisco, CA, USA
20 Employees
170K-215K Annually

Pika Logo Pika

Scientist

Information Technology
In-Office
Palo Alto, CA, USA
29 Employees
185K-400K Annually

AfterQuery Logo AfterQuery

Scientist

Artificial Intelligence • Big Data
In-Office
San Francisco, CA, USA
200 Employees
250K-450K Annually

ifm Logo ifm

Scientist

Information Technology • Automation • Manufacturing
In-Office
Sunnyvale, CA, USA
3924 Employees
150K-450K Annually

Similar Companies Hiring

Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account