Machine Learning Engineer (Evals and Voice Models)

Posted 2 Days Ago
San Francisco, CA, USA
In-Office
181K-250K Annually
Senior level
Artificial Intelligence • Cloud • Mobile • Sales • Software
Aircall is the phone system for modern business.
The Role
Build evaluation infrastructure for Aircall’s AI voice, chat, and messaging agents. Responsibilities include designing benchmarks and metrics, training and fine-tuning TTS, ASR, and speech-to-speech models, developing automated regression and release gates, calibrating LLM-as-judge systems, analyzing failures, monitoring production quality and drift, and evaluating voice-specific scenarios such as accents, noise, barge-in, DTMF, latency, and tool failures.
Summary Generated by Built In

Aircall is a unicorn, AI-powered customer communications platform used by 22,000+ companies worldwide to drive revenue, resolve issues faster, and scale customer-facing teams. We’re redefining customer communications by bringing voice, SMS, WhatsApp, and AI together into one seamless workspace.

Our momentum comes from a simple idea: help teams work smarter, not harder. Aircall’s AI Voice Agent automates routine calls, AI Assist streamlines post-call work, and AI Assist Pro delivers real-time guidance so people can do their best work. The result is higher revenue, faster resolutions, and teams that scale with confidence.

Aircall is headquartered in Paris, our European HQ, with a strong North American presence anchored in Seattle, our North American HQ, and teams across Madrid, London, Berlin, San Francisco, New York City, Sydney, and Mexico City. We’ve built a product customers love and a business that’s scaling quickly, backed by world-class investors and driven by rapid AI innovation across multiple product lines.

At Aircall, you’ll join a company in motion. We’re ambitious, product-driven, and execution-focused, with visible impact, fast decisions, and real growth.


How we work at Aircall: We’re customer-obsessed, data-driven, and focused on delivering meaningful outcomes. We value ownership, continuous learning, and thoughtful speed. If you thrive in a collaborative, fast-moving environment where trust and impact matter, you’ll feel at home here.


Aircall's AI suite includes an AI Voice Agent and AI Messaging Agent that autonomously handle calls, WhatsApp, and SMS, plus AI Assist, which delivers real-time coaching, call summaries, and CRM automation for sales and support teams. We are looking for someone that can build out the evaluation foundation across all of these products and other agentic products. You'll work on voice models, agent capability evals, benchmark design, LLM-as-judge systems, failure analysis, and the infrastructure that ties it together by establishing shared metrics, test sets, and tooling to measure accuracy, resolution quality, and safety consistently across products. You will set up repeatable pipelines for regression testing and benchmarking as models and features evolve so teams can ship confidently without re-inventing evaluation methodology for each product.

Key Responsibilities

  • Design and document comprehensive evaluation frameworks for Aircall’s AI agents across voice, chat and messaging.
  • Train and fine-tune voice models (TTS, ASR, speech-to-speech) using production and synthetic data, iterating on architecture, data mix, and training strategy to improve accuracy, naturalness, and latency.
  • Assess AI generated solutions across training pipelines, experimentation setups, debugging processes, and optimization strategies. 
  • Analyze system design decisions and identify strengths, weaknesses, and potential failure points.
  • Design annotation guidelines and workflows for human-labeled evaluation data, and calibrate LLM-as-judge systems against human raters to ensure automated evals stay trustworthy over time.
  • Build and maintain live quality monitoring for deployed AI agents, tracking accuracy, resolution rate, and safety signals in production, and flagging model or data drift before it impacts customers.
  • Own the metric contract for every published AI metrics, including definition, population, grain, rollup, validity window.
  • Build release gates, the offline regression suite each AI surface must pass before a prompt, model, or config change ships, measuring reliability across repeated trials, not just average pass rates.
  • Build voice-specific evaluation: simulated callers across accents, languages, background noise, barge-in, DTMF, and tool failures, with latency and ASR accuracy as first-class quality metrics.

Minimum Qualifications

  • BS in Computer Science, Machine Learning, Statistics, or related field
  • 3+ years of experience in ML Engineering or Applied ML with 8+ years of overall experience
  • Strong experience in evaluating supervised, unsupervised, LLMs and deep learning models.
  • Hands-on experience in failure analysis and evaluating LLMs
  • Experience building automated evaluation systems
  • Strong communication skills to articulate complex technical concepts across technical and non-technical audiences
  • Hands-on experience training or fine-tuning voice/speech models (TTS, ASR, or speech-to-speech), including data pipeline construction and experimentation.

Preferred Qualifications

  • MS / PhD in Computer Science, Machine Learning, Statistics, or related field
  • Experience evaluating LLMs or agentic systems (e.g., LLM-as-a-judge, RAG evaluation)
  • Experience with synthetic data generation and prompt engineering
  • Experience training or fine-tuning voice models at scale, with familiarity in synthetic data generation, model distillation, or low-latency inference optimization for production voice agents.
Base salary range:
$181,000$250,000 USD

Why join us?

🚀 Key moment to join Aircall in terms of growth and opportunities

💆‍♀️ Our people matter, work-life balance is important at Aircall

📚 Fast-learning environment, entrepreneurial and strong team spirit

🌍 45+ Nationalities: cosmopolite & multi-cultural mindset

💶 Competitive salary package & benefits


DE&I Statement: 

At Aircall, we believe diversity, equity and inclusion – irrespective of origins, identity, background and orientations – are core to our journey. 

We pride ourselves on promoting active inclusion within our business to foster a strong sense of belonging for all. We’re working to create a place filled with diverse people who can enrich and learn from one another. We’re committed to ensuring that everyone not only has a seat at the table but is valued and respected at it by providing equal opportunities to develop and thrive.  

We will constantly challenge ourselves to make sure that we live up to our ambitions around diversity, equity and inclusion, and keep this conversation open. Above all else, we understand and acknowledge that we have work to do and much to learn.


Want to know more about candidate privacy? Find our Candidate Privacy Notice here.

We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.

Skills Required

  • Bachelor’s degree in Computer Science, Machine Learning, Statistics, or a related field
  • 3+ years of experience in ML Engineering or Applied ML
  • 8+ years of overall professional experience
  • Strong experience evaluating supervised, unsupervised, large language, and deep learning models
  • Hands-on experience with failure analysis and LLM evaluation
  • Experience building automated evaluation systems
  • Strong communication skills for technical and non-technical audiences
  • Hands-on experience training or fine-tuning voice or speech models, including TTS, ASR, or speech-to-speech models
  • Master’s or PhD in Computer Science, Machine Learning, Statistics, or a related field
  • Experience evaluating LLMs or agentic systems, including LLM-as-a-judge or RAG evaluation
  • Experience with synthetic data generation and prompt engineering
  • Experience training or fine-tuning voice models at scale
  • Familiarity with synthetic data generation, model distillation, or low-latency inference optimization for production voice agents

Aircall Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Aircall and has not been reviewed or approved by Aircall.

  • Healthcare Strength Employer-paid core medical, dental, and vision coverage is emphasized, alongside mental health and wellness support. Coverage is described as starting quickly after hire and complemented by wellness reimbursements and related programs.
  • Parental & Family Support Generous parental leave for primary and secondary caregivers is highlighted, with added childcare reimbursements. Family-oriented benefits are positioned as part of a supportive culture.
  • Leave & Time Off Breadth Unlimited PTO, wellness days, and paid volunteer time are offered. Time-off policies are framed to encourage rest and work-life balance.

Aircall Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Paris
700 Employees
Year Founded: 2014

What We Do

Aircall is the phone system for modern business. An entirely cloud-based voice platform that integrates seamlessly with popular productivity and helpdesk tools that workplaces are already using, Aircall was built to make phone support as easy to manage as any other business workflow—accessible, transparent, and collaborative.

Why Work With Us

At Aircall, we’re equally thrilled by our ambitious goals, and by the journey that will lead us there. Our culture is rooted in our mission: we believe that now more than ever, good communication has the ability to make a difference. We’re learning, trying, and improving every day.

Gallery

Gallery

Similar Jobs

CoreWeave Logo CoreWeave

Manager, Field Engineering

Cloud • Information Technology • Machine Learning
In-Office
4 Locations
1450 Employees

Shield AI Logo Shield AI

Capture Portfolio Analyst, X-BAT Family of Systems (FoS) (R6011)

Aerospace • Artificial Intelligence • Machine Learning • Robotics • Software
In-Office
6 Locations
130K-240K Annually

ServiceNow Logo ServiceNow

Architect

Artificial Intelligence • Cloud • HR Tech • Information Technology • Productivity • Software • Automation
Remote or Hybrid
San Diego, CA, USA
29000 Employees
148K-232K Annually

Eve Logo Eve

Enterprise Account Executive

Legal Tech • Software • Generative AI
Easy Apply
Remote or Hybrid
United States
180 Employees
300K-310K Annually

Similar Companies Hiring

Revel Thumbnail
Aerospace • Hardware • Robotics • Software
Marina Del Rey, California
60 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software
New York, New York
30 Employees
Vega Thumbnail
Artificial Intelligence • Automotive • Insurance • Transportation
US
43 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account