Site Reliability Engineer, Provider Operations

Posted Yesterday
Be an Early Applicant
Hiring Remotely in US
Remote
Mid level
Software
The Role
Own the operational health, observability, reliability, and failover of AI model providers and endpoints. Build SLO monitoring, canaries, quality regression detection, automated provider-operations tooling, and load-testing systems. Lead incident response, communicate with external providers, produce reliability scorecards, and drive postmortems and corrective actions. Support capacity planning and launch readiness for high-volume LLM inference traffic.
Summary Generated by Built In
About OpenRouter

OpenRouter is the leading AI routing and infrastructure layer that enterprises use to access, manage, and optimize the best large language models across providers—without lock-in, capacity constraints, or unnecessary cost. We power the most advanced AI teams in the world by giving them the flexibility to move fast, scale confidently, and stay future-proof as models evolve.

As enterprise adoption of AI accelerates, OpenRouter sits at the center of how organizations operationalize LLMs across research, product, and production workloads.

About the Role

OpenRouter routes almost a billion requests and more than 20 trillion tokens a day, across 80+ providers and thousands of endpoints. Every one of those providers can degrade, rate-limit, change behavior, or go down without warning. Our customers count on us to absorb that chaos so their apps never notice.

We're hiring our first AI Inference SRE to own the operational health of our provider supply. You'll make sure every endpoint we route to is fast, correct, and available, and that we detect and route around problems before customers do. You'll sit on the Provider Operations team, reporting to the Provider Operations Manager.

What You'll Do
  • Provider health and observability. Build and own monitoring for every provider and endpoint: latency, throughput, error rates, uptime, and output correctness. Set SLOs per provider tier and alert on them.

  • Detection and failover. Improve how quickly we detect degraded endpoints, and work with the routing team so traffic shifts away from them automatically.

  • Incident response. Own on-call for provider incidents: triage, mitigate, communicate with providers, run postmortems, and drive follow-ups to closure.

  • Provider accountability. Turn telemetry into scorecards and SLO reporting that providers act on, and be the technical escalation point when a provider's endpoint is misbehaving.

  • Quality regression detection. Build continuous canaries and evals that catch silent regressions (quantization changes, broken tool calling, truncated streams, pricing or usage-reporting mismatches), not just outright downtime.

  • Automate the toil. Replace manual provider-ops work (disabling endpoints, capacity changes, deprecations, rate-limit tuning) with safe, auditable tooling.

  • Capacity and launch readiness. Build tooling to load-test endpoints before big launches so day-zero traffic doesn't take them down.

About You
  • 4+ years in SRE, production engineering, or infrastructure roles running high-traffic, customer-facing systems.

  • Strong with observability tooling and practice: metrics, tracing, logs, SLOs/error budgets, alerting that is always actionable.

  • Capable software engineer who prefers writing tools over executing runbooks. TypeScript and/or Python.

  • Experienced with distributed systems failure modes: timeouts, retries, backpressure, partial outages, noisy neighbors.

  • Calm, clear incident commander who communicates well with external partners under pressure.

  • Understands, or is eager to learn deeply, how LLM inference is served: streaming, tool calling, prompt caching, throughput/latency tradeoffs, and how provider APIs differ.

Nice to Have
  • Experience at an inference provider, model lab, GPU cloud, or API gateway/CDN company.

  • Experience with our stack: TypeScript, Cloudflare Workers, Postgres, ClickHouse, GCP, Vercel.

  • Background in routing, load balancing, or traffic management systems.

  • Experience with evals or synthetic monitoring for ML systems.

Skills Required

  • 4+ years of experience in SRE, production engineering, or infrastructure roles running high-traffic, customer-facing systems
  • Strong experience with observability, including metrics, tracing, logs, SLOs, error budgets, and actionable alerting
  • Software engineering ability with TypeScript and/or Python
  • Experience with distributed-systems failure modes, including timeouts, retries, backpressure, partial outages, and noisy neighbors
  • Ability to act as an incident commander and communicate clearly with external partners under pressure
  • Understanding of, or willingness to deeply learn, LLM inference concepts including streaming, tool calling, prompt caching, and throughput-latency tradeoffs
  • Experience at an inference provider, model lab, GPU cloud, or API gateway/CDN company
  • Experience with TypeScript, Cloudflare Workers, Postgres, ClickHouse, GCP, or Vercel
  • Background in routing, load balancing, or traffic management systems
  • Experience with evaluations or synthetic monitoring for ML systems
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
8 Employees
Year Founded: 2023

What We Do

A router for LLMs. 180+ models, explorable data, private chat, & a unified API. https://openrouter.ai/discord

Similar Jobs

Comcast Logo Comcast

Account Executive

Digital Media • Information Technology • News + Entertainment
Remote or Hybrid
Oregon, USA
115000 Employees

Comcast Logo Comcast

Comcast Cybersecurity: Cyber Incident Response Triage Analyst

Digital Media • Information Technology • News + Entertainment
Remote or Hybrid
Pennsylvania, USA
115000 Employees
60K-139K Annually

Liberty Mutual Insurance Logo Liberty Mutual Insurance

Senior Technical Claims Specialist, Energy Complex

Artificial Intelligence • Fintech • Insurance • Marketing Tech • Software • Analytics
Remote or Hybrid
5 Locations
40000 Employees
83K-139K Annually

Samsara Logo Samsara

Senior Software Engineer

Artificial Intelligence • Cloud • Computer Vision • Hardware • Internet of Things • Software
Easy Apply
Remote or Hybrid
United States
4000 Employees
155K-260K Annually

Similar Companies Hiring

Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Revel Thumbnail
Aerospace • Hardware • Robotics • Software
Marina Del Rey, California
70 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software
New York, New York
30 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account