Sr. Evaluation Engineer

Posted 12 Hours Ago
Easy Apply
Be an Early Applicant
Bangalore, Bengaluru, Karnataka, IND
Hybrid
Senior level
Artificial Intelligence • Cloud • Information Technology • Machine Learning • Software
Hybrid Observability powered by AI
The Role
Design and build production-grade evaluation pipelines, golden datasets, automated graders, and regression frameworks for LLMs, agents, retrieval systems, and tool integrations. Define quality metrics, generate representative and adversarial test cases, calibrate LLM-based graders, monitor AI quality in production, and convert failures into tests and safeguards. Integrate evaluations with CI/CD, experimentation, and release gating while mentoring engineers and establishing evaluation-driven development practices.
Summary Generated by Built In

About Us:  

We love going to work and think you should too. Our team is dedicated to trust, customer obsession, agility, and striving to be better everyday. These values serve as the foundation of our culture, guiding our actions and driving us towards excellence. We foster a culture of performance and recognition, allowing us to transform growth as we enable our employees to do the best work of their careers.

This position is located in Bangalore. You'll be working in a major tech center of Pune, India. Across the globe, our Centers of Energy serve as hubs where we accelerate productivity and collaboration, inspire creativity, and cultivate a culture of connection and celebration. Our teams coordinate their time in Centers of Energy to reflect how they work best.

To learn more about life at LogicMonitor, check out our Careers Page.

What You'll Do:

LogicMonitor® is the AI-first hybrid observability platform powering the next generation of digital infrastructure. LogicMonitor delivers complete visibility and actionable intelligence across on-premises, cloud, and edge environments. By anticipating issues before they strike, optimizing resources in real time, and enabling faster, smarter decisions, LogicMonitor helps IT and business leaders protect margins, accelerate innovation, and deliver exceptional digital experiences without compromise.

Our customers love LogicMonitor's ability to bring cloud and traditional IT together into one view, as seen in minimal churn rates, expansion business, and exciting new customer references. In fact, LogicMonitor has received the highest Net Promoter Score of any IT Infrastructure Management provider. LogicMonitor also boasts high employee satisfaction. We have been certified as a Great Place To Work®, and named one of BuiltIn's Best Places to Work for the seventh year in a row! 

Edwin AI is LogicMonitor’s AI-powered observability and incident intelligence platform. It helps enterprise operations teams investigate incidents, identify root causes, recommend remediation, and automate operational workflows.

As a Senior AI Engineer, Evaluations, you will design and build the evaluation systems that guide how Edwin AI is developed, tested, and released. You will create production-grade evaluation pipelines, golden datasets, automated graders, and regression frameworks for AI agents, retrieval systems, tool integrations, and complex investigation workflows

Here's a closer look at this key role:

  • Define quality metrics for incident diagnostics, root-cause analysis, alert correlation, grounding, tool use, safety, and operational usefulness.
  • Build offline and online evaluation pipelines in Python and integrate them with CI/CD, experimentation, model selection, prompt iteration, and release gating.
  • Lead the creation and maintenance of golden datasets and regression suites using alerts, events, metrics, logs, traces, topology, configuration data, incident timelines, change records, ITSM workflows, and historical investigation outcomes.
  • Build representative, customer-specific scenarios covering different technologies, failure modes, operational patterns, and environmental constraints.
  • Use human-authored and AI-assisted methods to generate regression, edge, adversarial, rare, ambiguous, and incomplete-context test cases.
  • Treat evaluation datasets and test suites as first-class components that evolve alongside Edwin AI.
  • Design step-level and trajectory-level evaluations for multi-step and multi-agent workflows.
  • Assess both final outcomes and intermediate behavior, including planning, reasoning consistency, retrieval, evidence use, tool selection, tool parameters, state transitions, escalation decisions, and human-in-the-loop approvals.
  • Identify whether failures originate from models, prompts, retrieval, data quality, tools, agent logic, orchestration, or infrastructure.
  • Evaluate capabilities including incident investigation, on-call assistance, operational question answering, change-impact analysis, remediation recommendations, automated resolution, infrastructure operations, and ITSM and observability integrations.
  • Design and calibrate LLM-based graders against expert human judgment.
  • Monitor AI quality and behavioral drift in production, and convert failures and customer feedback into new tests and safeguards.
  • Establish evaluation-driven development practices and mentor other engineers.
What You'll Need:
  • 5+ years of experience in software engineering, machine learning, applied AI, or a related field.
  • Strong Python engineering skills and experience building production systems.
  • Hands-on experience with AI evaluation, experimentation, testing, and quality frameworks.
  • Experience using multiple LLM and agent evaluation frameworks, such as LangSmith, Arize Phoenix, Braintrust, DeepEval, Ragas, TruLens, OpenAI Evals, MLflow, or comparable platforms.
  • Ability to select, customize, and integrate evaluation frameworks for offline testing, online monitoring, regression analysis, experimentation, model and prompt comparison, and release gating.
  • Strong understanding of LLMs, agents, retrieval-augmented generation, prompt engineering, tool calling, and context engineering.
  • Experience evaluating non-deterministic, multi-step, or multi-agent AI systems.
  • Ability to translate human and domain-expert judgment into test cases, evaluation rubrics, scoring functions, and automated graders.
  • Experience with LLM-as-a-judge techniques, including grader design, calibration, reliability measurement, and alignment with expert human judgment.
  • Experience with regression testing, CI/CD, production monitoring, behavioral drift detection, and failure analysis.
  • Strong analytical, systems-thinking, and communication skills.

Click here to read our International Applicant Privacy Notice.

LogicMonitor is an Equal Opportunity Employer
At LogicMonitor, we believe that innovation thrives when every voice is heard and each individual is empowered to bring their unique perspective. We’re committed to creating a workplace where diversity is celebrated, and all employees feel inspired and supported to contribute their best.

For us, equal opportunity means fostering a truly inclusive culture where everyone has the chance to grow and succeed. We don’t just open doors; we invite you to step through and be part of something bigger. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, or status as a protected veteran.

Notice Regarding Use of AI in Hiring

We use artificial intelligence tools to assist with reviewing job applications, such as matching skills and experience to job requirements. These tools support, but do not replace, human review. All hiring decisions are made by our recruiting and hiring teams. You may opt out of AI processing at any time, and your application will still be reviewed. To opt out, please contact us at [email protected]

By submitting your application, you acknowledge this notice.

 

                                               

Our goal is to ensure an accessible and inclusive experience for every candidate.

If you need a reasonable accommodation during the application or interview process under applicable local law, please submit a request via this Accommodation Request Form.

Know your rights: workplace discrimination is illegal. Please click here to review LogicMonitor’s U.S. Pay Transparency Nondiscrimination Provision.

Skills Required

  • 5+ years of experience in software engineering, machine learning, applied AI, or a related field.
  • Strong Python engineering skills and experience building production systems.
  • Hands-on experience with AI evaluation, experimentation, testing, and quality frameworks.
  • Experience using multiple LLM and agent evaluation frameworks (e.g., LangSmith, Arize Phoenix, Braintrust, DeepEval, Ragas, TruLens, OpenAI Evals, MLflow).
  • Ability to select, customize, and integrate evaluation frameworks for offline testing, online monitoring, regression analysis, experimentation, and release gating.
  • Strong understanding of LLMs, agents, retrieval-augmented generation, prompt engineering, tool calling, and context engineering.
  • Experience evaluating non-deterministic, multi-step, or multi-agent AI systems.
  • Ability to translate human and domain-expert judgment into test cases, rubrics, scoring functions, and automated graders.
  • Experience with LLM-as-a-judge techniques, grader design, calibration, reliability measurement, and alignment with expert human judgment.
  • Experience with regression testing, CI/CD, production monitoring, behavioral drift detection, and failure analysis.
  • Strong analytical, systems-thinking, and communication skills.

What the Team is Saying

Kenyon
Franky
Kwame
Gisselle
Antonio
Carly
Chris
Rockel
Rob
David
Peyton
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Santa Barbara, CA
1,100 Employees
Year Founded: 2007

What We Do

LogicMonitor® is the AI-first hybrid observability platform powering the next generation of digital infrastructure. LogicMonitor delivers complete visibility and actionable intelligence across on-premises, cloud, and edge environments. By anticipating issues before they strike, optimizing resources in real time, and enabling faster, smarter decisions, LogicMonitor helps IT and business leaders protect margins, accelerate innovation, and deliver exceptional digital experiences without compromise. For more information, visit www.logicmonitor.com and our blog, or follow us on LinkedIn, X, Facebook, and YouTube.

Why Work With Us

We love going to work and think you should too. We are customer-obsessed, work as one agile team, and strive to be better every day while building trust. These are our core values. So it's no surprise that we work hard and genuinely have fun working with each other as we expand our global presence and achieve record-breaking success.

Gallery

Gallery
Gallery
Gallery
Gallery
Gallery
Gallery
Gallery
Gallery
Gallery
Gallery

LogicMonitor Offices

Hybrid Workspace

Employees engage in a combination of remote and on-site work.

We call our offices Centers of Energy, because they’re where we accelerate work, spark creativity, and ignite our culture of connection and celebration. Our teams coordinate their time in Centers of Energy to reflect how they work best.

Typical time on-site: Flexible
Company Office Image
HQSanta Barbara, CA
Company Office Image
Austin, TX
Company Office Image
Boston, MA
Company Office Image
London, UK
Company Office Image
Pune, IN
Company Office Image
San Francisco
Company Office Image
Singapore
Company Office Image
Sydney, Australia
Learn more

Similar Jobs

LogicMonitor Logo LogicMonitor

Sr. Forward Deployed Engineer

Artificial Intelligence • Cloud • Information Technology • Machine Learning • Software
Easy Apply
Hybrid
Bangalore, Bengaluru Urban, Karnataka, IND
1100 Employees

LogicMonitor Logo LogicMonitor

Principal Product Manager

Artificial Intelligence • Cloud • Information Technology • Machine Learning • Software
Easy Apply
Hybrid
Bangalore, Bengaluru Urban, Karnataka, IND
1100 Employees

LogicMonitor Logo LogicMonitor

Machine Learning Engineer

Artificial Intelligence • Cloud • Information Technology • Machine Learning • Software
Easy Apply
Hybrid
Bangalore, Bengaluru Urban, Karnataka, IND
1100 Employees

LogicMonitor Logo LogicMonitor

Systems Architect

Artificial Intelligence • Cloud • Information Technology • Machine Learning • Software
Easy Apply
Hybrid
2 Locations
1100 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account