Senior Software Engineer - LLM Ops & Evals

Posted Yesterday
Be an Early Applicant
Amsterdam, NLD
Hybrid
Senior level
Productivity • Software • Automation
The Role
Own and scale the LLM gateway and evaluation platform, including model routing, failover, rate limits, cost attribution, deployments, infrastructure as code, observability, and async workloads. Build versioned datasets and experiment tracking for AI quality evaluation. Lead reliability, security, compliance, incident response, on-call, and shared-service integrations across product and ML teams.
Summary Generated by Built In

Every AI call in DataSnipper goes through us. Product teams do not talk to model providers directly, they talk to our gateway. We are responsible for how inference is routed, how it fails over, what it costs, and how anyone can tell whether the output is any good.

The second half of the job is evaluation. We are building the platform teams use to measure AI quality: versioned datasets, experiment tracking, and evaluation runs they can act on. It is a hard problem and largely an open one, so you will have real influence over how we solve it.

This is a small team with a large blast radius. You will own real production systems, set the standards other teams build against, and see your work in front of hundreds of thousands of users in audit and finance.

Why DataSnipper

Audit and finance are still massively manual and we are changing that. DataSnipper is a $1B, bootstrapped unicorn with 600,000+ users across 180+ countries, already embedded in the daily workflows of top audit and accounting firms.
Now, we are taking things further with our Excel Agent, bringing AI directly into where the work actually happens. Unlike generic AI tools, we do not sit on the sidelines. Our AI operates inside Excel, with access to real documents and audit evidence, meaning it does not just generate answers, it does the work, with full traceability.
We are not just applying AI, we are redefining how audit gets done. If you want to build something category-defining at scale, this is the place.

What you will doTechnical Delivery
  • Own the LLM gateway: routing, provider failover, rate limits, retries, and cost attribution across multiple model providers

  • Deploy, version and deprecate models across clouds, regions and environments, including quota and capacity planning, managed as infrastructure as code

  • Build the shared evaluation platform: versioned datasets, experiment tracking, run and result schemas, reporting, and trace linkage back to the run

  • Own the infrastructure for async and long-running AI workloads

Reliability, Security & On-Call
  • Own observability for AI traffic: latency, retries and fallbacks, token usage, cost and errors, per team and per use case

  • Take part in the on-call rotation, run incidents, and close the follow-ups

  • Implement the security and compliance controls the platform is held to: retention, access control, RBAC and SSO

Collaboration & Impact
  • Define and maintain clean integration contracts between the platform and the teams that consume it

  • Partner with product and ML engineers to turn their requirements into platform capabilities that are self-service rather than a request queue

What you will bringMust-Have
  • 5+ years in backend or platform engineering, with strong production Python

  • Experience building or running LLM inference infrastructure: a gateway or routing layer with multiple providers, failover, rate limiting and cost attribution

  • Experience with cloud at the infrastructure level and infrastructure as code (we run across Azure and GCP with Terraform)

  • Experience running a shared service in production: on-call, incidents, postmortems, SLOs

  • Hands-on experience with observability tools (OpenTelemetry, Grafana), including instrumenting services and designing dashboards and alerts

  • Comfort with privacy and compliance work: PII handling, anonymisation, retention, access control

  • Experience building platform or shared-service capabilities consumed by multiple internal teams

Nice-to-Have
  • Experience with LLM or agent evaluation

  • Temporal or another durable workflow engine

  • Self-hosted inference, capacity planning, load testing

  • Synthetic data or document anonymisation pipelines

  • Document AI: VLMs, OCR, structured extraction and the metrics that go with it

  • Domain experience in audit, accounting or fintech

What We Expect
  • Ownership: You own work end-to-end, anticipate issues, and ensure high-quality delivery without close supervision

  • Growth Mindset: You encourage open feedback exchange and provide clear, balanced feedback that helps others grow

  • Collaboration: You build strong cross-functional relationships and influence peers through expertise, data, and empathy

  • Adaptability: You navigate ambiguity calmly, model positive behavior, and help peers adjust through clear communication

  • Judgment: You exercise sound judgment in ambiguous situations, balance speed and accuracy, and adjust priorities proactively

Recruitment steps
  • Recruiter screen

  • Hiring Manager interview

  • Peer programming session

  • System design interview

  • Final interviews with Engineering leadership

Skills Required

  • 5+ years of backend or platform engineering experience
  • Strong production Python experience
  • Experience building or operating LLM inference infrastructure with multiple providers, failover, rate limiting, and cost attribution
  • Cloud infrastructure experience and infrastructure as code, specifically Azure, GCP, and Terraform
  • Experience operating shared production services, including on-call, incident response, postmortems, and SLOs
  • Hands-on experience with observability tools such as OpenTelemetry and Grafana, including service instrumentation, dashboards, and alerts
  • Experience with privacy and compliance, including PII handling, anonymization, retention, and access control
  • Experience building platform or shared-service capabilities for multiple internal teams
  • Experience with LLM or agent evaluation
  • Experience with Temporal or another durable workflow engine
  • Experience with self-hosted inference, capacity planning, or load testing
  • Experience with synthetic data or document anonymization pipelines
  • Document AI experience with VLMs, OCR, structured extraction, and related metrics
  • Domain experience in audit, accounting, or fintech
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Amsterdam
215 Employees
Year Founded: 2017

What We Do

Accelerate your Audit and Finance teams’ productivity. Drive company growth and resilience with DataSnipper’s Intelligent Automation Platform in Excel.

Similar Jobs

Mastercard Logo Mastercard

Manager, Specialist Sales, SME Solutions Northern Europe

Blockchain • Fintech • Payments • Consulting • Cryptocurrency • Cybersecurity • Quantum Computing
Hybrid
Amsterdam, NLD
38800 Employees

Datadog Logo Datadog

Field Enablement Manager - Amsterdam

Artificial Intelligence • Cloud • Security • Software • Cybersecurity
Easy Apply
Hybrid
Amsterdam, NLD
6500 Employees

Adyen Logo Adyen

Business Analyst

Fintech • Payments • Financial Services
Easy Apply
Hybrid
Amsterdam, NLD
4771 Employees

Navan Logo Navan

Deal Desk Manager

Fintech • Information Technology • Payments • Productivity • Software • Travel • Automation
Easy Apply
Hybrid
Amsterdam, NLD
3300 Employees

Similar Companies Hiring

Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Revel.io Thumbnail
Aerospace • Hardware • Robotics • Software
US
50 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software
New York, New York
30 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account