Senior AI Engineer

Reposted 10 Days Ago
Be an Early Applicant
Bengaluru, Bengaluru Urban, Karnataka, IND
In-Office
Senior level
Artificial Intelligence • Productivity
Augment Market Research Operations
The Role
Design, build, and improve production-grade AI agent systems: synthetic data pipelines, evaluation datasets, monitoring and logging, long-context handling, cascading error mitigation, prompt engineering, and tooling to enable domain experts to refine training data and optimize model performance at scale.
Summary Generated by Built In
About Metaforms

Market research runs on 30-year-old survey platforms and armies of specialists hand-coding questionnaires in proprietary languages. Metaforms is the agent layer that does that work. Every survey is a program — full of skip logic, piping, quotas, and loops — and a single wrong number in a client report is unrecoverable. Our AI agents write production survey code, QA live deployments, process and clean large structured datasets, configure analysis, and generate client-ready reports, so agencies like Dynata, Savanta, and Borderless Access ship more projects with far less friction.

  • 1,000+ surveys processed monthly

  • Serving Fortune 500 companies across the globe

  • Rapid month-over-month growth

We’re Series A funded and scaling fast, aggressively growing our AI engineering team to build the next generation of production-grade AI agent systems.

The Role

We’re hiring a Senior AI Engineer to own the design, development, and continuous improvement of the AI agent systems that power modern research operations.

This is a high-ownership, high-impact role at the intersection of applied AI and systems engineering. You’ll work on genuinely hard problems: agent reliability at scale, long-context handling, cascading error mitigation, and evaluation infrastructure — like codegen agents that write in proprietary DSLs, computer-use agents that QA live deployments, data agents that clean tabular exports and configure multi-step analysis, and evals for outputs where “correct” is genuinely ambiguous. And you’ll do it on a team that ships fast and treats quality as non-negotiable.

What You’ll OwnAgent Harness and Architecture
  • Own the agent harness our production agents run on — the loop where agents plan, use tools, check their work, and recover from failures

  • Lead research and implementation for long-context handling and cascading-error challenges in multi-step agent pipelines

  • Drive context engineering strategy and experimentation frameworks across the team

Evaluation and Production Monitoring
  • Define structured rubrics for evaluating AI outputs on nuanced, ambiguous research tasks

  • Build continuous monitoring, tracing, and failure-mode analysis for agents in production — including the loop that turns production failures into test cases

  • Create tooling that lets domain experts refine and evolve the skill files, eval sets, and knowledge bases our agents consume

Reliability for High-Stakes Outputs
  • Build eval suites — regression sets, golden datasets, LLM-as-judge pipelines — that catch regressions before deploy

  • Develop evaluation datasets for DSLs, structured data transforms, and computed outputs to systematically find and close model weaknesses

  • Design human-in-the-loop and review workflows for outputs where a single wrong number in a client report is unrecoverable

What We’re Looking ForMust-Have
  • Built and operated agentic systems in production — multi-step pipelines, tool use, codegen, computer-use, or data and reporting agents — not just prototypes

  • 4+ years of engineering experience, with at least 1 year focused on LLM/agent systems in production

  • Deep hands-on experience with frontier model APIs (Anthropic, OpenAI, Gemini), evaluation frameworks, and AI system optimization

  • Strong Python skills; Go or TypeScript a plus

  • Solid grasp of context engineering and evaluation methodology

  • Strong instincts for debugging complex, non-deterministic system failures

  • High ownership: you drive problems to resolution independently and pull others in when it matters

Nice to Have
  • Experience with LLM observability and eval tooling (Braintrust, Langfuse, LangSmith, Weave, promptfoo, or in-house equivalents)

  • Background in semantic parsing, DSLs, or structured-output generation

  • Prior work on computer-use or browser agents

  • Experience with human-in-the-loop agent workflows where proposals are reviewed before apply, or agents over large structured datasets

Why Metaforms
  • Work at the frontier of production AI: systems handling 1,000+ research projects a month, with the reliability bar that implies

  • A small, senior team where your decisions carry real architectural weight

  • Zero-bureaucracy culture: high autonomy, fast feedback loops, direct access to leadership

  • Well-funded and financially stable, with a clear roadmap and the runway to execute on it

Benefits
  • Full family health insurance

  • $1,000 USD annual learning and development budget

  • Dedicated mentor and coaching support

  • Free snacks and dinner at the office

Skills Required

  • 4+ years of engineering experience, with at least 2 years focused on LLM systems or applied ML in production
  • Deep hands-on experience with LLMs, synthetic data pipelines, evaluation frameworks, and AI system optimization
  • Strong Python skills
  • Go or TypeScript experience
  • Solid understanding of prompt engineering, model fine-tuning, and evaluation methodology
  • Experience building and operating production AI systems (not just prototypes)
  • Strong instincts for debugging complex, non-deterministic system failures
  • High ownership: drive problems to resolution independently
  • Experience with model monitoring, bias evaluation, or dataset management tooling
  • Background in semantic parsing, DSLs, or structured output models
  • Prior work on agentic systems, multi-step reasoning pipelines, or tool-use frameworks
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Bengaluru , Karnataka
38 Employees
Year Founded: 2023

What We Do

Metaforms is the AI platform that helps market research agencies operate smarter, and win more business. We've built AI Agents that get your business; learning your exact standards, integrating with your tools, and augmenting your team's capabilities across every research workflow. • Generate survey code with AI • Process data with intelligent flagging • Respond to RFPs faster with automated structuring • Coordinate vendors with smart routing • Conduct voice research with AI moderation The result? Agencies win more business. Handle exponentially more projects. Scale without the complexity or quality trade-offs. Trusted by world's leading agencies; Metaforms is defining the future of market research.

Similar Jobs

Ericsson Logo Ericsson

Artificial Intelligence Engineer

Cloud • Information Technology • Internet of Things • Machine Learning • Software • Cybersecurity • Infrastructure as a Service (IaaS)
In-Office
Bangalore, Bengaluru Urban, Karnataka, IND
88000 Employees

JPMorganChase Logo JPMorganChase

Data Engineer

Financial Services
Hybrid
Bengaluru, Bengaluru Urban, Karnataka, IND
289097 Employees

Optum Logo Optum

Machine Learning Engineer

Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
In-Office
Bengaluru, Bengaluru Urban, Karnataka, IND
160000 Employees
Hybrid
Bengaluru, Bengaluru Urban, Karnataka, IND
289097 Employees

Similar Companies Hiring

Legora Thumbnail
Artificial Intelligence • Legal Tech • Software
New York, New York
700 Employees
Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account