Machine Learning Intern

Posted 4 Days Ago
Be an Early Applicant
San Francisco, CA, USA
In-Office
Internship
Artificial Intelligence • Information Technology
The Role
Conduct an end-to-end machine learning research project focused on audio and voice AI. Responsibilities include literature review, implementation, experimentation, ablation design, model training and evaluation on large-scale telephony audio, presenting findings, and potentially moving successful results into production. Potential areas include text-to-speech, speech recognition, neural audio codecs, streaming inference, and conversational turn-taking.
Summary Generated by Built In
The Role: Machine Learning Research Intern, Audio

As a Research Intern at Bland, you will own a focused research project across our voice stack: speech-to-text, large language models, neural audio codecs, or text-to-speech. You will work alongside our research team on the same problems they are working on, not on a side track built to keep interns busy.

We scope internships around a single meaningful question that can be answered in the time you have. The goal is a result worth shipping, publishing, or both. Interns here regularly see their work reach production systems handling millions of calls.

What You Will Do

Own a research question end to end

  • Take one well-scoped problem from literature review through implementation, experimentation, and results.

  • Design ablations that isolate what actually caused an improvement.

  • Present your findings to the research team and defend the methodology.

Work on real systems

  • Train and evaluate models on large-scale, real-world telephony audio, including the accents, noise, and artifacts that make production speech hard.

  • Use our distributed GPU infrastructure rather than toy-scale setups.

  • Where the result warrants it, work with engineers to move it toward production.

Choose your depth
Depending on your background and interests, your project may focus on:

  • Expressive and controllable text-to-speech, including prosody and emotion modeling

  • Neural audio codecs and discrete or continuous speech representations

  • ASR robustness for telephony, accents, and code switching

  • Real-time and streaming inference under latency constraints

  • Full-duplex conversation and turn-taking dynamics

What Makes You a Great Fit

Research foundations

  • Currently pursuing a MS or PhD in ML, CS, EE, or a related field, or equivalent research experience.

  • Comfortable reading a paper and reimplementing it without hand-holding.

  • Experience with self-supervised, generative, or multimodal modeling.

Audio or speech grounding

  • Hands-on work with speech or audio models, whether TTS, ASR, codecs, or audio representation learning.

  • Strong intuition for audio quality and what makes synthetic speech sound wrong.

  • Prior publications or open source contributions in speech or language AI are a strong signal, though not required.

Engineering ability

  • Fluent in PyTorch and comfortable in a real codebase.

  • Able to run your own experiments on GPU clusters without waiting to be unblocked.

How You Show Up
  • You identify the single experiment that validates an idea in days, not months.

  • You measure everything and let data drive decisions.

  • You are honest about negative results, because they are how we narrow the search.

  • You are obsessed with making voice agents sound truly human.

  • You use AI tools aggressively to amplify your own impact.

Benefits
  • Competitive intern compensation

  • Mentorship from researchers working on frontier voice AI

  • Every tool you need to succeed

  • Beautiful office in Levi's Plaza, SF with rooftop views

  • A real shot at a return offer

Skills Required

  • Currently pursuing a master's or PhD in machine learning, computer science, electrical engineering, or a related field, or equivalent research experience
  • Ability to read research papers and reimplement methods independently
  • Experience with self-supervised, generative, or multimodal modeling
  • Hands-on experience with speech or audio models, including TTS, ASR, codecs, or audio representation learning
  • Strong intuition for audio quality and synthetic speech
  • Fluency in PyTorch
  • Comfort working in a real codebase
  • Ability to run experiments independently on GPU clusters
  • Prior publications or open-source contributions in speech or language AI
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Bakersfield, CA
27 Employees
Year Founded: 2023

What We Do

The enterprise first solution for AI phone calls

Similar Jobs

In-Office
Los Gatos, CA, USA
13212 Employees
40-85 Hourly

Netflix Logo Netflix

Scientist

News + Entertainment
In-Office
Los Gatos, CA, USA
13212 Employees
40-85 Hourly
In-Office
South San Francisco, CA, USA
20069 Employees
26-50 Hourly

Kodiak Robotics Logo Kodiak Robotics

Winter 2027 Intern, Artificial Intelligence/Machine Learning

Automotive • Robotics • Software • Transportation
In-Office
Mountain View, CA, USA
81 Employees
10K-10K Annually

Similar Companies Hiring

Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees
Kepler  Thumbnail
Artificial Intelligence • Fintech • Software
New York, New York
9 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account