Research Engineer Intern, Evaluations

Reposted 22 Days Ago
Be an Early Applicant
San Francisco, CA, USA
In-Office
Internship
Artificial Intelligence • Software • Automation
The Role
Intern will create evaluation frameworks for AI agents, benchmark models, and develop automated assessments for data-focused tasks in AI systems.
Summary Generated by Built In
Research Engineer Intern, Evaluations & Benchmarks
Location: San Francisco (Hybrid)

About TensorStax:

TensorStax is building fully autonomous AI systems to manage and optimize mission-critical data infrastructure. Our research integrates reinforcement learning and language models to enhance reasoning over large-scale data lakes and warehouses, detect failures in pipelines, and autonomously construct and optimize data workflows with high precision.

We are looking for a Research Engineer Intern to design evaluation frameworks and benchmarks that assess the autonomy, adaptability, and reliability of AI agents in data engineering environments. This role is ideal for candidates passionate about AI evaluations, language model benchmarking, and autonomous data systems.

What You’ll Do:

  • Develop evaluation environments to test AI agents' ability to reason, plan, and act autonomously within mission-critical data pipelines.
  • Design benchmarks to assess model capabilities in failure detection, pipeline optimization, and agentic decision-making in data workflows.
  • Implement automated assessment frameworks for language model-based agents operating over data lakes and warehouses.
  • Work with synthetic and real-world datasets to create robust testing environments for AI-driven data automation.
  • Collaborate with research engineers to refine reward shaping strategies, guiding models toward more efficient and agentic behaviors in data-intensive tasks.

What We’re Looking For:

  • Experience in language model research, with a focus on benchmarking LLMs in mission-critical domains.
  • Strong background in AI evaluation methodologies, reinforcement learning, and RLHF techniques.
  • Familiarity with benchmarking language models for structured and unstructured data tasks.
  • Proficiency in Python and experience with ML frameworks like PyTorch or JAX.
  • Hands-on experience with data lakes, warehouses, and data engineering tools (Snowflake, BigQuery, dbt, Spark, Kafka).
  • High agency—proactive, resourceful, and comfortable working in a fast-paced research environment with minimal supervision.
  • Attention to detail—ability to design rigorous, reproducible experiments and evaluations.

Bonus Points:

  • Contributions to open-source AI benchmarks (e.g., SweBench, BIRD, SPIDER).
  • Contributions to open-source agentic frameworks.
  • Experience developing custom RL environments for AI evaluation.
  • Strong understanding of ETL, ELT, and data transformation pipelines.

Benefits:

  • Competitive internship stipend.
  • 100% employer-covered health, dental, and vision insurance (for eligible interns).
  • Access to Bay Club or Equinox in San Francisco.
  • Opportunity to work at the cutting edge of AI evaluations and autonomous data engineering research.

Skills Required

  • Experience in language model research focusing on benchmarking LLMs
  • Strong background in AI evaluation methodologies and reinforcement learning
  • Proficiency in Python and experience with ML frameworks like PyTorch or JAX
  • Hands-on experience with data lakes and data engineering tools
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: San Francisco, CA
4 Employees

What We Do

Autonomous AI to help build and maintain data pipelines using your infrastructure.

Similar Jobs

Navan Logo Navan

Venue Find Specialist

Fintech • Information Technology • Payments • Productivity • Software • Travel • Automation
Easy Apply
Remote or Hybrid
USA
3300 Employees
55K-68K Annually

Navan Logo Navan

Event Coordinator

Fintech • Information Technology • Payments • Productivity • Software • Travel • Automation
Easy Apply
Remote or Hybrid
USA
3300 Employees

Tapestry - Coach and Kate Spade Logo Tapestry - Coach and Kate Spade

Sales Associate III

eCommerce • Fashion • Retail • Sales • Wearables • Design
Hybrid
Viejas Stop, CA, USA
16000 Employees
15-20 Hourly

Tapestry - Coach and Kate Spade Logo Tapestry - Coach and Kate Spade

Sales Support Associate II

eCommerce • Fashion • Retail • Sales • Wearables • Design
Hybrid
Los Cerritos, CA, USA
16000 Employees
15-22 Hourly

Similar Companies Hiring

Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
Kepler  Thumbnail
Fintech • Software
New York, New York
6 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account