Senior Applied ML Engineer, Evals & Data

Posted Yesterday
Be an Early Applicant
Bengaluru, Bengaluru Urban, Karnataka, IND
In-Office
Senior level
Artificial Intelligence • Digital Media • Marketing Tech • Software
The Role
Own the quality measurement and improvement of an AI video-editing agent. Build evaluation datasets, offline and online evaluations, automated checks, human reviews, regression gates, and release criteria. Analyze agent failures and improve results through better data, evaluation methods, model selection, and fine-tuning. Partner with product and engineering to ship reliable improvements while tracking quality, latency, and cost.
Summary Generated by Built In
About
  • We're building the future of storytelling and video editing.

  • We're a small team that moves fast and builds things we're proud of.

  • We care obsessively about taste: in design, in product, in every detail.

  • We're solving these hard problems.

  • We're backed by a Tier-1 global fund, YC, and founders of billion dollar companies.

Engineering

Video is the most powerful way humans tell stories. It always has been. But creating it today is still painfully hard. Fragmented tools, steep learning curves, and workflows that get in the way of the actual creative work. We're building Cardboard to change that.

Cardboard runs a real video editor in the browser, backed by a serious cloud media pipeline and an AI agent that actually understands footage. You will own how we measure and improve the quality of Cardboard’s AI agent.

You will study real agent runs, turn important failures into evaluation cases, and measure whether changes make the product better. You will also work with product and engineering to ship those improvements.

This is not a research-only, prompt-only, or QA role.

You’ll be working alongside a team of engineers who all care deeply about craft, including the founders. You like owning problems end to end, and you’d rather ship something great this week than something perfect next quarter

 
What you'll actually do
  • Define quality standards and build trusted evaluation datasets from real product usage.

  • Build offline and online evaluations, including automated checks and human review.

  • Analyze model and agent failure patterns, then improve quality through better data, evaluation methods, model selection, and, where useful, fine-tuning.

  • Add regression checks and release gates while tracking quality, latency, and cost.

  • Solve these hard problems.

What we are looking for
  • Experience shipping and operating an LLM or agent system used by real customers.

  • Strong software engineering skills in TypeScript or Python, with the ability to work across both.

  • Experience building evaluations, datasets, experiments, or AI quality systems.

  • Strong product judgment and the ability to turn unclear quality problems into measurable improvements.

You do not need a PhD or experience training foundation models. Evidence of building reliable AI products matters more than formal credentials or knowledge of a specific framework.

Nice to have
  • Experience with multimodal AI, video, media, or creative software.

  • Experience with human labeling, model graders, or fine-tuning.

  • Good knowledge of experiment design and statistics.

Within your first six months:
  • We have a trusted quality baseline for our main agent workflows.

  • Production failures regularly become new evaluation cases.

  • Important agent changes pass clear regression checks before release.

  • We can show measurable improvements in key editing workflows.

What you get

You'd be surrounded by people who are absurdly good at what they do. One started coding at 11 and shipped an app with 6M+ downloads in high school. One got into CS engineering at 14 and has been working on distributed systems for 8+ years. One's an ex-founder who took a company to 1.2M users and $300M+ in transactions. That's the team. We're looking for someone who'll raise the bar on technical craftsmanship and creative product quality. Apart from that you'd get:

  • Competitive salary and founding-team equity.

  • Unlimited tokens across every AI model. Use whatever you want, as much as you want.

  • A healthy budget for AI tools and any peripherals you need to do your best work.

Skills Required

  • Experience shipping and operating an LLM or agent system used by real customers
  • Strong software engineering skills in TypeScript or Python, with the ability to work across both
  • Experience building evaluations, datasets, experiments, or AI quality systems
  • Strong product judgment and ability to turn unclear quality problems into measurable improvements
  • Experience with multimodal AI, video, media, or creative software
  • Experience with human labeling, model graders, or fine-tuning
  • Knowledge of experiment design and statistics
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
5 Employees
Year Founded: 2025

What We Do

Cardboard, Inc. develops an agentic, browser-based AI video editor for growth and marketing teams and serious creators. Users can describe an intended edit in natural language, while the product cuts, captions, reframes, composes, and searches footage. Its browser-native workflow combines a multi-layer timeline, collaboration, auto-captioning, voiceover and music generation, and exports for professional finishing applications, helping teams produce and iterate videos faster.

Similar Jobs

Adyen Logo Adyen

Software Engineer

Fintech • Payments • Financial Services
Easy Apply
Hybrid
Bengaluru, Bengaluru Urban, Karnataka, IND
4771 Employees

LogicMonitor Logo LogicMonitor

Edwin GTM Engineer

Artificial Intelligence • Cloud • Information Technology • Machine Learning • Software
Easy Apply
Hybrid
Bangalore, Bengaluru Urban, Karnataka, IND
1100 Employees

Zscaler Logo Zscaler

Consultant

Cloud • Information Technology • Security • Software • Cybersecurity
Easy Apply
Hybrid
Bangalore, Bengaluru, Karnataka, IND
8697 Employees

ServiceNow Logo ServiceNow

Staff Software Engineer

Artificial Intelligence • Cloud • HR Tech • Information Technology • Productivity • Software • Automation
Hybrid
Bangalore, Bengaluru, Karnataka, IND
29000 Employees

Similar Companies Hiring

Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Revel.io Thumbnail
Aerospace • Hardware • Robotics • Software
US
50 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software
New York, New York
30 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account