Machine Learning Engineer

Reposted 23 Days Ago
7 Locations
Remote or Hybrid
Senior level
Artificial Intelligence • Information Technology • Software
The Role
Senior Machine Learning Engineer role focusing on building and improving AI evaluation infrastructures, collaborating with a multidisciplinary team, and contributing to the development of a scalable evaluation platform.
Summary Generated by Built In
About Arena Intelligence

Arena is the platform for evaluating how AI models perform in the real world. Founded by researchers from UC Berkeley's SkyLab, we're on a mission to measure and advance the frontier of AI for real-world use, and to build the foundation for everyone to understand, shape, and benefit from it.


Tens of millions of people use Arena each month to evaluate how frontier systems handle the work they actually do. The preferences they share power the most transparent, rigorous, and human-centered evaluations in AI. Leading AI labs, enterprises, and independent researchers rely on our work and open datasets to understand how models behave in real workflows: agentic coding, creative generation, professional productivity, and beyond. We go beyond leaderboards and decompose what human experience reveals about AI, so models advance toward the work people actually do.


We're a team of researchers, academics, builders, and creatives from UC Berkeley, Google, Stanford, and DeepMind. We seek truth, move fast, and value craftsmanship, curiosity, and impact over hierarchy. We're building a company where thoughtful, curious people from all backgrounds can do their best work together, in an office culture that radiates excellence, energy, and focus.

About the Role

Arena Intelligence is seeking a Senior Machine Learning Engineer to help scale and strengthen the core infrastructure that powers real-world AI evaluation. You’ll play a foundational role in shaping how we build, deploy, and improve our model benchmarking systems, working across data pipelines, inference APIs, and new evaluation methodologies. This is an opportunity to apply your technical expertise to a platform trusted by millions, and to help define how cutting-edge AI is assessed in the wild.

As one of the first ML engineers on the team, you’ll partner closely with researchers, engineers, and product leadership to turn new ideas into reliable systems. You’ll help us move fast while staying rigorous, improving reproducibility, scaling up to new modalities, and deepening our ability to understand and compare frontier models.

You’ll
  • Architect and build what will become our core modeling for data and evaluation products

  • Own the full stack data, model training, and eval pipelines

  • Help grow a culture of feedback and rapid product iteration as we build new features as a tight-nit team

  • Conduct research into state-of-the-art evaluation methods and contribute to the long-term vision for a centralized, scalable evaluation platform.

You’ll have
  • Strong programming skills with the ability to work across the stack in a typical recommendation system or LLM stack

  • Experience in deep learning, language models or reward model training

  • Experience in working with LLM for fine tuning, prompt engineering, function calling etc

  • Self-motivated with a willingness to take ownership of tasks

  • A passion for shipping quality products

  • 4+ years of industry experience or relevant projects

  • Solid understanding of statistics, and various tools and methodologies for evaluating uncertainty in a way that is specific to the given product being shipped

What we offer
  • We offer competitive compensation and equity aligned to the markets where our team members are based. The base salary range will depend on the candidate’s permanent work location.

  • Comprehensive health and wellness benefits, including medical, dental, vision, and additional support programs.

  • The opportunity to work on cutting-edge AI with a small, mission-driven team

  • A culture that values transparency, trust, and community impact

Come help build the space where anyone can explore and help shape the future of AI.

Arena Intelligence provides equal employment opportunities (EEO) to all employees and applicants for employment without regard to race, color, religion, sex, national origin, age, disability, genetics, sexual orientation, gender identity, or gender expression. We are committed to a diverse and inclusive workforce and welcome people from all backgrounds, experiences, perspectives, and abilities.

Skills Required

  • Strong programming skills with the ability to work across the stack in a typical recommendation system or LLM stack
  • Experience in deep learning, language models or reward model training
  • Experience in working with LLM for fine tuning, prompt engineering, function calling
  • 4+ years of industry experience or relevant projects
  • Solid understanding of statistics and methodologies for evaluating uncertainty
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: San Francisco, California
58 Employees
Year Founded: 2025

What We Do

Created by researchers from UC Berkeley, Arena (formerly LMArena) is a community-powered platform for understanding AI performance in the real world. Tens of millions of builders, researchers, and creative professionals come to Arena to use frontier models and give feedback on their responses, shaping a public leaderboard grounded in real-world use.

Similar Jobs

Block Logo Block

Machine Learning Engineer

Blockchain • eCommerce • Fintech • Payments • Software • Financial Services • Cryptocurrency
In-Office or Remote
8 Locations
12000 Employees
200K-415K Annually

Cash App Logo Cash App

Machine Learning Engineer

Blockchain • Fintech • Mobile • Payments • Software • Financial Services
Remote or Hybrid
8 Locations
3500 Employees
200K-415K Annually

Block Logo Block

Machine Learning Engineer

Blockchain • eCommerce • Fintech • Payments • Software • Financial Services • Cryptocurrency
In-Office or Remote
8 Locations
12000 Employees
277K-415K Annually

Block Logo Block

Machine Learning Engineer

Blockchain • eCommerce • Fintech • Payments • Software • Financial Services • Cryptocurrency
In-Office or Remote
8 Locations
12000 Employees
277K-415K Annually

Similar Companies Hiring

Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees
Kepler  Thumbnail
Fintech • Software
New York, New York
6 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account