Lead Machine Learning Engineer

Posted 2 Days Ago
Be an Early Applicant
2 Locations
Hybrid
170K-190K Annually
Expert/Leader
Artificial Intelligence • Machine Learning • Natural Language Processing • Software
The Role
Lead development of an evaluation platform for agentic AI and LLM systems. Design benchmarking, human review, LLM-as-judge, monitoring, annotation, dataset versioning, and data quality workflows. Own technical architecture, production infrastructure, stakeholder reporting, and platform health while partnering with research, product, and platform teams. Mentor engineers and support scalable AI experimentation and deployment.
Summary Generated by Built In
At ASAPP, our mission is simple: deliver the best AI-powered customer experience—faster than anyone else. To achieve that, we’re guided by principles that shape how we think, build, and execute. We value customer obsession, purposeful speed, ownership, and a relentless focus on outcomes. ASAPP’s AI Engineering team is seeking an enterprising, talented and curious machine learning engineer.
 
The AI Engineering team is responsible for working closely with the research and modeling teams to create state-of-the-art NLP models for specific tasks, and deploy them in a production setting designed to serve our customers at scale. We are looking for a Machine Learning Engineer to help build and evaluate the core intelligence behind our agentic AI systems. This role will play a key part in designing and owning evaluation frameworks that ensure quality, safety, and performance across complex agentic systems.
We're looking for a Lead Machine Learning Engineer to own and grow the evaluation platform that measures quality, safety, and performance across ASAPP's agentic AI systems- the infrastructure that tells us, with confidence, whether a model or agent change is actually an improvement before it reaches customers.
This a hybrid role with 10-12 days of in-office presence per month to balance flexibility with collaboration.

What you'll do

  • Help develop the technical roadmap and architecture for the evaluation platform, from offline benchmarking to online/production monitoring of agentic and LLM-based systems.

  • Design eval methodologies appropriate to different stages of the pipeline: golden/regression test sets, human-in-the-loop review workflows, LLM-as-judge approaches, and automated metrics for task success, safety, and hallucinations.

  • Build the data infrastructure evaluation depends on: annotation and labeling pipelines, dataset versioning, data quality checks, and tooling that lets researchers and product teams run and interpret experiments without needing platform team help.

  • Partner closely with Research, Product, and Platform teams to productize experiments into robust AI solutions

  • Represent the eval platform to stakeholders outside the immediate team- set expectations on what "good" looks like for a model/agent release, and report on platform health and coverage.

  • Stay current with advancements in ML, NLP, voice, and LLM systems, and contribute actively to technical discussions across teams.

  • Mentor and support other engineers through design reviews, feedback, and knowledge sharing.

What you'll need

  • Deep, hands-on experience building and operating evaluation systems for modern ML/LLM/agentic systems- not just consuming existing eval tools.

  • Demonstrated experience leading the technical direction of a project or small team: setting architecture, driving design reviews, and being accountable for a system's long-term health (not just shipping features).

  • Strong architectural skills, with proven experience designing complex, data-intensive software systems and production experience with Python, AWS, Kubernetes, and/or Docker.

  • Experience designing data pipelines for ML evaluation- labeling/annotation workflows, dataset versioning and quality control, and reproducible benchmarking.

  • A Bachelor’s Degree in CS or other related fields

  • Demonstrated technical mentorship of junior and mid-level engineers, driving adoption of best practices and architectural alignment for scalability and extensibility.

  • Desire to learn, teach, and collaborate closely with cross-functional peers.

What we'd like to see

  • Experience building and evaluating agentic systems at scale. 

  • Experience with voice/audio quality evaluations.

  • Production experience with LLM-centric services (e.g., inference, orchestration, evaluation, monitoring)

  • Familiarity with large-scale ML experimentation, benchmarking, or simulation frameworks.

  • Experience with conversational/customer-support AI domains (e.g., containment rate, conversation quality, goal completion).

  • Knowledge of techniques for optimizing model architectures for faster inference.

  • Experience with AWS, CI/CD, Kafka, Athena

Skills Required

  • Deep hands-on experience building and operating evaluation systems for modern machine learning, LLM, or agentic systems
  • Experience leading the technical direction of a project or small team, including architecture, design reviews, and long-term system ownership
  • Experience designing complex, data-intensive software systems
  • Production experience with Python, AWS, Kubernetes, and/or Docker
  • Experience designing ML evaluation data pipelines, including labeling workflows, dataset versioning, quality control, and reproducible benchmarking
  • Bachelor's degree in Computer Science or a related field
  • Technical mentorship of junior and mid-level engineers
  • Ability to learn, teach, and collaborate cross-functionally
  • Experience building and evaluating agentic systems at scale
  • Experience with voice or audio quality evaluations
  • Production experience with LLM-centric services such as inference, orchestration, evaluation, or monitoring
  • Familiarity with large-scale ML experimentation, benchmarking, or simulation frameworks
  • Experience with conversational or customer-support AI domains
  • Knowledge of techniques for optimizing model architectures for faster inference
  • Experience with AWS, CI/CD, Kafka, and Athena
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: New York, NY
389 Employees
Year Founded: 2014

What We Do

Our artificial intelligence and machine learning products deliver automation and human augmentation, allowing individuals and organizations to realize their full potential. Today, the world's largest organizations rely on ASAPP to provide amazingly efficient and effective customer experiences. Our Research & Development team is unparalleled, driving the advancement of AI, machine learning, speech recognition, robotic process automation, natural language processing and more.

Similar Jobs

Capital One Logo Capital One

Lead Machine Learning Engineer

Fintech • Machine Learning • Payments • Software • Financial Services
Hybrid
5 Locations
55000 Employees
197K-246K Annually

Capital One Logo Capital One

Lead Machine Learning Engineer

Fintech • Machine Learning • Payments • Software • Financial Services
Hybrid
5 Locations
55000 Employees
230K-286K Annually

Capital One Logo Capital One

Lead Machine Learning Engineer

Fintech • Machine Learning • Payments • Software • Financial Services
Hybrid
7 Locations
55000 Employees
179K-246K Annually

Capital One Logo Capital One

Lead Machine Learning Engineer

Fintech • Machine Learning • Payments • Software • Financial Services
Hybrid
2 Locations
55000 Employees
197K-246K Annually

Similar Companies Hiring

Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Revel.io Thumbnail
Aerospace • Hardware • Robotics • Software
US
50 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software • Productivity
US
15 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account