Data Scientist– AI Infra & Evaluation Foundations

Posted Yesterday
Be an Early Applicant
Tel Aviv, ISR
Hybrid
Mid level
Artificial Intelligence • Productivity • Sales • Software
Power the most ambitious teams.
The Role
Design and own evaluation methodologies, metrics, datasets, test suites, and pipelines for production AI agents. Build end-to-end evaluation tools, analyze agent execution traces and failure modes, establish quality standards, and embed continuous evaluation into engineering workflows. Partner with engineering and product teams to drive adoption, improve agent reliability, and guide production decisions across the organization.
Summary Generated by Built In

About monday.com:

monday.com is the AI work platform powering the most ambitious teams. 250,000+ customers across departments use us to bring people, workflows, and AI agents together on one flexible platform where AI doesn't just assist, it executes. We move fast, build things that matter, and foster an ownership-driven culture where you're empowered to shape how organizations work and outpace their competition.

About the team

 

The AI Infra group builds the foundations, tools, and platforms that every team at monday relies on to ship intelligent, agentic features. We own the core infrastructure—including the AI Gateway and our centralized Evals framework—ensuring every AI feature deployed to production is secure, resilient, cost-effective, and above all, trustworthy.

 

Our focus is on the frontier of agentic AI: dissecting complex agent trajectories, building robust evaluation frameworks, and turning subjective notions of "good AI" into rigorous, actionable metrics. As a Data Scientist on this team, you will bridge the gap between AI research and production infrastructure. You'll partner closely with engineering and product teams across monday to design the judges, metrics, and error analysis workflows that allow us to ship cutting-edge AI agents with speed and confidence.

 
 
 

This position is based at our Tel Aviv office (Headquarters).

 

About the role:

As a Data Scientist in AI Infra, your goal goes far beyond simply building an evaluation framework—you will own the organizational impact of how monday evaluates and trusts AI. You will define how teams measure quality, influence engineering decisions across R&D, and turn fuzzy notions of "good AI" into numbers product teams rely on to ship with confidence.

  • Own the evaluation methodology: Design metrics, pipelines, and methodology that teams across monday trust and adopt as their source of truth.

  • Transform the AI agent lifecycle: Standardize how AI agents are built, regression-tested, and maintained across the org, embedding continuous evaluation into everyday engineering workflows and post-deployment monitoring.

  • Drive organizational impact & enablement: Partner with AI feature teams across monday to translate domain expectations into meaningful datasets, test suites, and continuous evaluation pipelines—leveling up engineers and product managers along the way.

  • Build hands-on tools: Prototype and stand up eval pipelines end-to-end, bridging the gap between ambiguous, high-level product requirements into clear, quantifiable evaluation standards that become central to how features are greenlit for production

  • Anticipate future failure modes: Stay ahead of evolving agent architectures by proactively designing next-generation evaluation strategies.

 

Requirements:

  • Agentic Systems & Architecture:

    • 3+ years of experience as a Data Scientist in non-academic settings working with complex production running AI systems. Familiarity with current agentic frameworks like LangGraph, LangChain, and SoTA SDKs.

    • Deep, practical understanding of how agents operate - models, context, capabilities and harnesses. Deep experience with agentic evaluation methodologies

  • Execution, Code & Trace-First Mindset:

    • Production-grade coding skills with a track record of building, prototyping, and shipping end-to-end data or eval pipelines,

    • A "trace-first" diagnostic mindset—comfortable diving into raw agent execution logs, inspecting failure modes, and constructing qualitative error taxonomies.

  • Product & Organizational Impact:

    • Strong product intuition and exceptional communication skills to translate complex evaluation data into clear, actionable guidelines.

    • Proven ability to partner closely with software engineers and product teams, taking ownership of driving adoption and raising the quality bar across the organization.

 

Preferred Qualifications

  • Direct experience designing evaluation strategies for complex agentic systems in production

  • Prior experience working within centralized platform/infra teams that support multiple product verticals.

  • Experience contributing to modern microservice architectures, GitHub workflows, and automated production CI/CD pipelines

  • Familiarity with modern agent and eval tooling and observability stacks (e.g., LangSmith, Langfuse, or custom internal platforms).

  • Knowledge of TypeScript or experience working within modern platform architectures.

  • Master’s degree in Computer Science, Data Science, Statistics, Engineering, or a related quantitative field.

#LI-DNI

Skills Required

  • 3+ years of experience as a Data Scientist in a non-academic setting
  • Experience working with complex production AI systems
  • Familiarity with agentic frameworks such as LangGraph and LangChain
  • Deep practical understanding of agent architecture, including models, context, capabilities, and harnesses
  • Deep experience with agentic evaluation methodologies
  • Production-grade coding skills
  • Experience building, prototyping, and shipping end-to-end data or evaluation pipelines
  • Ability to inspect raw agent execution logs, diagnose failure modes, and construct qualitative error taxonomies
  • Strong product intuition and exceptional communication skills
  • Ability to partner with software engineers and product teams and drive organizational adoption
  • Experience designing evaluation strategies for complex agentic systems in production
  • Experience working within centralized platform or infrastructure teams
  • Experience with modern microservice architectures
  • Experience with GitHub workflows and automated production CI/CD pipelines
  • Familiarity with LangSmith, Langfuse, or comparable agent evaluation and observability platforms
  • Knowledge of TypeScript or modern platform architectures

What the Team is Saying

Ruchita
Nate
Kyle
Brad Wisselman
Brad Wisselman
Bianca Collado

monday.com Compensation & Benefits Highlights

  • Healthcare Strength — Comprehensive medical, dental, and vision coverage is offered in the U.S., complemented by mental-health resources such as counseling sessions and Calm access.
  • Parental & Family Support — Up to 13 weeks of fully paid parental leave for all caregivers is provided, with adoption assistance and a structured return-to-work program that support families.
  • Retirement Support — A 401(k) with an automatic 3% company contribution regardless of employee deferral is included, strengthening long-term financial security.

monday.com Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: New York, NY
3,155 Employees
Year Founded: 2012

What We Do

At monday.com, we help teams get more work done. We are the best AI work platform that empowers teams to automate, build, and scale their impact end-to-end with tools that actually execute the work for you. With over $1B in ARR, 250,000+ customers, and a global team, we’re serious about building a product people love to use and giving our employees the same ownership and flexibility to shape the way the world works.

Why Work With Us

At monday.com we believe in transparency, accountability, and impact. Together, those values have lent themselves to create a strong culture of professional and creative autonomy where every team member is encouraged to share ideas and help bring them to life!

Gallery

Gallery
Gallery
Gallery
Gallery
Gallery
Gallery
Gallery
Gallery
Gallery
Gallery
Gallery
Gallery
Gallery
Gallery
Gallery
Gallery
Gallery
Gallery
Gallery
Gallery
Gallery
Gallery
Gallery
Gallery
Gallery

monday.com Teams

Team
Customer Experience
About our Teams

monday.com Offices

Hybrid Workspace

Employees engage in a combination of remote and on-site work.

monday.com embraces a flexible work environment with our hybrid model.

Typical time on-site: 3 days a week
HQNew York, NY
HQTel Aviv
Denver, CO
London
Melbourne
Munich
Paris, France
Sao Paolo
Singapore
Sydney
Tokyo
Warsaw
Learn more

Similar Jobs

monday.com Logo monday.com

Security Engineer

Artificial Intelligence • Productivity • Sales • Software
Hybrid
Tel Aviv, ISR
3155 Employees

monday.com Logo monday.com

Lead Product Designer

Artificial Intelligence • Productivity • Sales • Software
Hybrid
Tel Aviv, ISR
3155 Employees

monday.com Logo monday.com

AI Enablement Solutions Manager

Artificial Intelligence • Productivity • Sales • Software
Hybrid
Tel Aviv, ISR
3155 Employees

monday.com Logo monday.com

Software Engineering Tech Lead - AI Project Management

Artificial Intelligence • Productivity • Sales • Software
Hybrid
Tel Aviv, ISR
3155 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account