Senior AI Engineer - Observability

Posted Yesterday
Hiring Remotely in United States
Remote
Senior level
Software
The Role
Build and improve production AI features using LLMs, RAG, semantic search, and agentic workflows. Design evaluation pipelines, golden datasets, scoring methodologies, regression tests, and quality gates. Instrument AI systems to monitor quality, latency, cost, safety, and reliability using observability platforms. Create reusable tooling and standards, analyze production traces, support compliant AI delivery, and coach product teams on dependable AI engineering practices.
Summary Generated by Built In

Sr AI Engineer (Observability) 

Function: Engineering 

Reports to: Director, Software Engineering 

Location: Remote - United States 

Position Summary 

OnBoard is a board intelligence platform trusted by over 6,000 organizations to simplify governance, and we are building AI-powered product experiences across RAG pipelines, semantic search, summarization, and emerging agentic workflows. We are looking for a Senior AI Engineer to help make those experiences reliable, measurable, cost-effective, and safe in production. 

This is a hands-on engineering role for someone who wants to build AI systems, not just observe them. You will partner with product engineering teams to design and improve AI features, instrument them deeply, and — critically — define how we know those features actually work. You will build the evaluation pipelines and datasets that tell us whether an AI experience is good enough to ship, analyze real-world behavior, and turn production signals into product and architecture improvements. 

You will help define how OnBoard ships AI: how we test prompts and retrieval quality, detect regressions, monitor cost and latency, evaluate user-facing quality, and safely evolve models, prompts, datasets, and providers over time. 

The right person is a strong software engineer with practical LLM application experience and a genuine quality mindset. You care not only that an AI feature works in a demo, but that it performs consistently for customers, degrades gracefully, provides traceable results, and improves through feedback loops. And you have real opinions about what makes an evaluation trustworthy — not just that one exists, but whether it measures the right thing. 

What You'll Do 

Build and improve AI-powered product systems 

- Partner with product engineers to design, build, and improve AI features using LLMs, RAG, semantic search, and agentic workflows 

- Contribute directly to production codebases in Python, C#/.NET, or related technologies 

- Improve prompt, retrieval, context assembly, ranking, grounding, and response-generation patterns 

- Help teams make practical architecture tradeoffs across quality, latency, cost, privacy, and maintainability 

- Support model and provider evaluations, migrations, fallback strategies, and rollout plans 

Make AI quality measurable — define what "good enough" means 

This is a defining pillar of the role, not an afterthought. It is not enough to have an evaluation; you will be responsible for whether our evaluations are adequate. 

- Design and implement evaluation pipelines for LLM-powered features, and build scoring methodologies from first principles rather than reaching for the nearest metric 

- Build and maintain versioned golden datasets covering real-world use cases, edge cases, failure modes, and customer-critical workflows 

- Implement LLM-as-judge, heuristic, human-feedback, and task-specific quality scoring approaches 

- Establish the criteria that determine whether an existing evaluation is sufficient for a given feature and risk profile — and identify gaps before they become production issues 

- Establish prompt and retrieval regression testing as part of the development lifecycle 

- Define quality gates and thresholds that help teams know when an AI feature is ready to ship 

Own observability, reliability, and cost signals 

- Instrument LLM interactions, RAG pipelines, tool calls, and agent workflows using observability platforms (OnBoard currently uses Arize; comparable tools include Langfuse, LangSmith, W&B, and OpenTelemetry-based stacks) 

- Track latency, token usage, cost, retrieval quality, groundedness, failure modes, safety signals, and user feedback 

- Build dashboards and alerts that surface meaningful product and engineering signals, not just raw telemetry 

- Analyze production traces to identify quality issues, cost spikes, regressions, and improvement opportunities 

- Create runbooks and response patterns for common LLM and AI-product failure modes 

Scale AI quality across teams 

- Build reusable libraries, SDKs, templates, and reference implementations that make correct AI instrumentation and evaluation easy 

- Document standards for tracing, metadata, prompt/version tracking, evaluation, cost reporting, and incident response 

- Coach product teams on AI quality, evaluation design, observability, and reliable release practices 

- Help establish shared patterns that let OnBoard scale AI development across product lines 

Support responsible and compliant AI delivery 

- Ensure AI systems handle sensitive data appropriately and align with security, privacy, SOC 2, ISO 27001, and data-residency requirements 

- Monitor guardrails, policy enforcement, content safety signals, and safety-related anomalies as a distinct observability concern 

- Support auditability and traceability of AI interactions where required 

  

What We're Looking For 

Required 

- 5+ years of software engineering experience building production systems 

- Hands-on experience building or operating LLM-powered features, RAG systems, AI workflows, or similar AI applications 

- Strong engineering ability in Python, C#/.NET, or both 

- Practical understanding of prompts, embeddings, vector search, retrieval quality, orchestration patterns, and LLM application architecture 

- Experience designing evaluations, quality metrics, or regression frameworks for AI or software systems — with the judgment to assess whether an evaluation actually measures what matters 

- Strong observability fundamentals: tracing, logging, metrics, alerting, and production debugging 

- Experience with CI/CD, git workflows, cloud environments, and production release practices (Azure DevOps preferred) 

- Ability to communicate clearly with engineering, product, QA, security, and business stakeholders 

- Strong product judgment and genuine curiosity about how AI systems behave with real users 

- Proficiency with AI-assisted development tools (e.g., Claude Code, PlayerZero) 

Preferred 

- Experience with LLMOps or AI observability tools such as Arize, Langfuse, LangSmith, W&B, Humanloop, or Helicone 

- Experience with OpenTelemetry, Azure Monitor, Application Insights, or similar observability platforms 

- Experience with Azure AI Search, Pinecone, Qdrant, Weaviate, pgvector, or other vector search platforms 

- Experience with Semantic Kernel, LangChain, LlamaIndex, AutoGen, or related frameworks 

- Experience with LLM-as-judge evaluation, RAG evaluation, semantic similarity metrics, hallucination detection, groundedness scoring, or human-feedback workflows 

- Experience with dedicated evaluation frameworks (e.g., DeepEval, LangTest) and benchmarking approaches for LLM outputs 

- Experience with A/B testing, online experimentation, or product analytics for AI features 

- Experience in regulated environments with SOC 2, ISO 27001, PII handling, or data-residency requirements 

- Background in QA, ML, or data engineering that informs a rigorous approach to quality measurement 

  

You'll Be Successful If You 

- Make AI feature quality measurable and visible 

- Help teams detect regressions before customers do 

- Improve reliability, latency, cost, and user trust in AI-powered experiences 

- Build reusable patterns that reduce friction for every product team 

- Translate ambiguous AI behavior into concrete engineering actions 

- Balance innovation with operational discipline 

  

Why This Role Matters 

AI is becoming a core part of the OnBoard product experience. This role helps determine whether that AI is merely impressive in demos or dependable for thousands of organizations making important governance decisions. 

You will have the opportunity to shape OnBoard's AI engineering standards, influence product architecture, and build the systems that let teams ship AI faster, safer, and with greater confidence. 


Competencies 

- Accountability 

- Adaptability 

- AI Curiosity / Innovation 

- Applied Learning 

- Business Acumen 

- Collaboration 

- Customer Focus 

- Dealing with Ambiguity 

- Decision Making 

- Driving for Results 

- Initiating Action 

- Planning and Organizing 

- Technical / Professional Knowledge 

About the Company:

Boards set the standard for what organizations can achieve. At OnBoard, our board management software helps boards function at a higher level so every organization can make a bigger difference in the world.

Launched in 2011, today, OnBoard serves as the board intelligence platform for more than 5,000 organizations and their 12,000 boards and committees in 60 countries worldwide. With customers in higher education, nonprofit, healthcare systems, government, and enterprise business, OnBoard is the leading board management provider.

OnBoard has grown from a class project at Purdue University in West Lafayette, Indiana in 2003 into the world’s leading board management software platform today. Backed by JMI Equity and the acquisitions of eScribe and Govenda, OnBoard is positioned to become the industry leader in Board Management and Meeting Solutions for private and public sector entities.

Benefits and Perks: 

  • Fully remote work with company provided equipment (laptop, software, etc.) 
  • Employment with a growing, casual, fun, philanthropic minded company
  • US Based Employees
    • Comprehensive, high-quality medical/prescription drug plan options, as well as dental and vision plan offerings.   
    • An employer contribution to your Health Savings Account (HSA) if you participate in a High Deductible Healthcare Plan.  
    • Medical Flexible Spending Accounts available.   
    • Dependent Care Flexible Spending Accounts available.  
    • Basic life insurance in the amount of $50,000 or 1 X’s your salary (whichever is higher) 
    • Short and long-term disability and Accidental Death and Dismemberment benefits at no cost to you.  
    • 401K Retirement Savings Plan with automatic enrollment at the first of the month following 60 days of employment at 5% to help you secure your financial freedom. We offer a generous company match that starts on the first of the month following 60 days of employment. The company match is dollar for dollar on the first 3% of your pay that you contribute and $0.50 on the dollar on the next 2%, for a total match of 4%. 
    • Paid Time Off (PTO)/Holiday 
  • CAN Based Employees
    • Employer paid Life and Accidental Death Insurance
    • Contribution to Health Care Spending Account
    • Dependent Life Insurance
    • Optional Life Insurance
    • LTD Insurance
    • Drug and Paramedical Coverage
    • Dental Insurance
    • Vision Insurance
    • EAP
  • AUS Based employees
    • Superannuation rate of 12% 
    • Monthly stipend of $400 AUD to purchase private medical insurance 
  • UK Based Employees (via EPG)
    • Pension - Aegon
      • Passageways/OnBoard contributes 8% of the employee's basic salary
      • Employees can contribute up to 100% of salary subject to max limits
      • Enrolled from Day 1 of employment
    • Private Medical Insurance
    • Life Assurance
    • Income Protection
    • Critical Illness
    • Employee Assistance Programme
    • Serious Illness Benefit
    • Help@Hand
    • Cashplan

Diversity Statement - Culture of Togetherness:  

At OnBoard, our mission is to encourage and celebrate a culture of togetherness. We acknowledge that uniqueness is powerful, and we welcome, foster, and appreciate all. Diversity, Equity, and Inclusiveness fuel the Pathfinder atmosphere and all our efforts. Our power is in our people and we Pledge 1% to give back to our communities and across the globe.

OnBoard is an equal opportunity employer and committed to a diverse and inclusive working environment. We  do not discriminate based on race, national origin, gender, gender identity, sexual orientation, protected veteran status, disability, age, or other legally protected status.

Interview Transparency & Technology Disclosure
We use video/audio recordings and artificial intelligence (AI) tools during our interview process to transcribe responses, evaluate skills, and streamline evaluations. Your data is processed securely and handled in line with our Privacy Policy and local data protection laws

Skills Required

  • 5+ years of software engineering experience building production systems
  • Hands-on experience building or operating LLM-powered features, RAG systems, AI workflows, or similar AI applications
  • Strong engineering ability in Python, C#/.NET, or both
  • Practical understanding of prompts, embeddings, vector search, retrieval quality, orchestration patterns, and LLM application architecture
  • Experience designing evaluations, quality metrics, or regression frameworks for AI or software systems
  • Strong observability fundamentals, including tracing, logging, metrics, alerting, and production debugging
  • Experience with CI/CD, Git workflows, cloud environments, and production release practices
  • Ability to communicate clearly with engineering, product, QA, security, and business stakeholders
  • Strong product judgment and curiosity about how AI systems behave with real users
  • Proficiency with AI-assisted development tools such as Claude Code or PlayerZero
  • Experience with LLMOps or AI observability tools such as Arize, Langfuse, LangSmith, W&B, Humanloop, or Helicone
  • Experience with OpenTelemetry, Azure Monitor, Application Insights, or similar observability platforms
  • Experience with Azure AI Search, Pinecone, Qdrant, Weaviate, pgvector, or other vector search platforms
  • Experience with Semantic Kernel, LangChain, LlamaIndex, AutoGen, or related frameworks
  • Experience with LLM-as-judge evaluation, RAG evaluation, semantic similarity metrics, hallucination detection, groundedness scoring, or human-feedback workflows
  • Experience with dedicated evaluation frameworks such as DeepEval or LangTest and benchmarking approaches for LLM outputs
  • Experience with A/B testing, online experimentation, or product analytics for AI features
  • Experience in regulated environments with SOC 2, ISO 27001, PII handling, or data-residency requirements
  • Background in QA, machine learning, or data engineering informing rigorous quality measurement
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Indianapolis, IN
200 Employees
Year Founded: 2003

What We Do

At OnBoard, we believe board meetings should be informed, effective, and uncomplicated. That’s why we give boards and leadership teams an elegant solution that simplifies governance. Launched in 2011, today, OnBoard serves as the board intelligence platform for more than 2,000 organizations and their 12,000 boards and committees in 32 countries worldwide. With customers in higher education, nonprofit, healthcare systems, government, and corporate enterprise business, OnBoard is the leading board management provider. We are always looking for smart people to join us - check out the current openings: https://www.onboardmeetings.com/about-us/careers With its headquarters in Indianapolis, Ind., OnBoard is a global company with offices in London, Sydney, and Montreal.

Similar Jobs

Lowe’s Logo Lowe’s

Project Coordinator

Consumer Web • eCommerce • Information Technology • Retail • Software • Analytics • App development
Remote or Hybrid
Wilkesboro, NC, USA
300000 Employees

Liberty Mutual Insurance Logo Liberty Mutual Insurance

Inside Sales Representative

Artificial Intelligence • Fintech • Insurance • Marketing Tech • Software • Analytics
Remote or Hybrid
10 Locations
40000 Employees
44K-100K Annually

Wells Fargo Logo Wells Fargo

Infrastructure Engineer

Fintech • Financial Services
Remote or Hybrid
Irving, TX, USA
205000 Employees

Wells Fargo Logo Wells Fargo

Infrastructure Engineer

Fintech • Financial Services
Remote or Hybrid
Charlotte, NC, USA
205000 Employees

Similar Companies Hiring

Kepler  Thumbnail
Artificial Intelligence • Fintech • Software
New York, New York
9 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Revel.io Thumbnail
Aerospace • Hardware • Robotics • Software
US
50 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account