AI Engineer, Decision Intelligence

Posted 12 Days Ago
Hiring Remotely in San Francisco, CA, USA
In-Office or Remote
180K-220K Annually
Senior level
Artificial Intelligence • Healthtech • Software • Analytics
The Role
Design and own the decision-intelligence layer: build models and LLM systems to decide which claims to contest and offer amounts, create evaluation pipelines and golden datasets, architect deterministic vs learned components, close feedback loops from outcomes, and produce explainable recommendations for customers.
Summary Generated by Built In
Every dispute is one move in a repeated game against an adaptive opponent. Build the system that gets better at playing it.


Why This Exists

A federal arbitration system called Independent Dispute Resolution, or IDR, now determines billions of dollars in healthcare payments each year. Providers win the vast majority of disputes, yet most eligible claims are never filed. The process is manual, fragmented, and resource-intensive, and most providers don't have the infrastructure to pursue what they're owed.

The No Surprises Act created the framework, and the market already exists. Today it runs on spreadsheets, consultants, and static playbooks. We're building the first intelligent system designed to operate inside it.

Automating the paperwork is table stakes. The interesting part is the second half: IDR is baseball-style arbitration, where each side submits one number and an arbitrator picks one. No splitting the difference. That means every submission is a bet, and every outcome is a signal about how to bet better next time.

This role owns that.


Why This Is Hard (and Interesting)

You're building a decision system in an adversarial, regulated, sparse-data environment.

The same payers appear repeatedly and they change behavior when they notice patterns. Arbitrators rotate and have their own tendencies. Regulations shift under you. You have to make a call on every claim before you have statistically comfortable data, and being wrong costs a provider real money.

The system also has to be defensible. When a customer asks why we submitted a particular number, "the model said so" is not an acceptable answer. You need reasoning you can explain to a hospital CFO.

Concretely: which claims to contest, what to offer, what evidence to include, how to write the argument, and how all of that should change based on the payer, the arbitrator, the procedure codes, and everything we've learned so far. Deterministic rules and LLM reasoning both have a role. Figuring out which does what is your call.


Who We Are

Recourse is being built in partnership with 25M Health, a healthtech venture firm. We have institutional backing, a shared platform team spanning engineering, strategy, design, and back-office, and early access to large provider systems.

We are actively filing disputes for real customers, including a large multi-facility health system and a litigation-finance partner with hundreds of millions in claim value. This is a funded, validated opportunity with real customers and real data.

We are a small, nimble team. We move quickly and we value clarity over theater. We want this to be the best work of your career. The stretch you look back on as the one where you shipped real things, with people who raised your game, on something that mattered.

We care about clear thinking, high ownership, intellectual honesty, and direct communication. We believe operations, product, and engineering should operate as one pod, not three functions. We want the machines to do machine work, and the humans to do their best work.


The Role

You'll own the intelligence layer of the product. You will:

  • Build the decision engine that determines which claims to contest and at what offer amount
  • Design the LLM systems that generate arguments: medical necessity, patient acuity, market comparables, the full submission
  • Build evaluation infrastructure. Golden datasets, offline evals, and the ability to know whether a change actually improved outcomes
  • Architect the split between deterministic rules and model reasoning, and defend where you drew the line
  • Close the feedback loop from arbitration outcomes back into the system, so every decision makes the next one better
  • Build for explainability. Every recommendation needs reasoning a customer would accept
  • Partner with operations and payer strategy, who see patterns in the claims before the data does

We build with Claude Code daily and run a Codex review pass on every slice. The stack is TypeScript, Next.js, Prisma, and PostgreSQL on Google Cloud, with LLM reasoning throughout.


Who You Are

You've shipped LLM systems that people depend on. Not demos. Production systems with evals, guardrails, and a real answer for what happens when the model is wrong.

You think about systems, not prompts. Prompt engineering is a component. The interesting work is architecture: what's deterministic, what's learned, how they interact, and how the whole thing improves over time.

You're rigorous about evaluation. You know that "it seems better" isn't evidence. You build the measurement before you build the feature.

You are AI-pilled and current. The landscape moves monthly. You track it because you want to see the next shift before anyone else, and you have opinions about what's real versus hype.

You have a bias to action. You don't default to no. Speed of iteration over polish of iteration. You start, you learn, you fix things in motion. Most decisions are reversible and do not need extensive study.

You are intellectually honest. You seek out evidence that disconfirms your approach. You say so when you're wrong. You use plain language. You respectfully challenge decisions you disagree with, and once a decision is made, you commit.

You put the team first. You are reliable and fully invested. You take your vacations. You check on your teammates. You help build a culture where people do their best work because they are supported, not squeezed.


What You Bring

  • 5+ years engineering, with meaningful recent time building LLM-powered or ML-driven products in production
  • Real experience with evaluation: building golden datasets, running offline evals, measuring whether changes helped
  • Comfort across the stack. You can ship the thing end to end, not just the model layer
  • Strong instincts for where probabilistic reasoning helps and where it introduces unacceptable risk
  • Comfort with TypeScript or Python. The specific stack matters less than the ability to pick things up
  • Experience building in regulated or security-sensitive environments (HIPAA, SOC 2, PCI, financial controls) is a plus
  • A preference for small teams and early-stage chaos over mature org charts

Strong plus, not required: healthcare data, claims, or any adversarial decision domain (fraud, risk, pricing, trading). If you have it, you'll move faster. If you don't, we'll teach you.

Sound judgment, technical depth, and ownership mindset are required. Grit matters more than pedigree.


Why This Role

Most applied AI jobs are wrapping a model around an existing workflow. This one is different: the intelligence is the product, the feedback loop is real and measurable, and the outcome is dollars recovered rather than an engagement metric.

You'd also be first here. Nobody in this market has built a genuine decision-intelligence layer on arbitration outcomes. The data exists and almost nobody is doing the work. You'd have unusual latitude to decide what this becomes.

If you want this to be the most memorable stretch of your career, where you shipped something real, with a team you respected, in a domain that actually matters, this is the seat. Please apply even if you don't fit 100% of these requirements. We would like to talk.


Equal Opportunity

Recourse is an equal opportunity employer. We do not discriminate on the basis of race, color, religion, sex, sexual orientation, gender identity, national origin, age, disability, veteran status, or any other characteristic protected by law. We believe the best teams are built from people with different backgrounds and perspectives, and we're committed to creating an environment where everyone can do their best work.

Compensation
The base pay range for this role is $180,000 – $220,000 per year.

Skills Required

  • 5+ years engineering experience
  • Recent experience building LLM-powered or ML-driven products in production with guardrails
  • Experience building evaluation infrastructure: golden datasets, offline evaluations, and metrics
  • Ability to ship end-to-end across the stack (model layer through product)
  • Comfort with TypeScript or Python
  • Experience building in regulated or security-sensitive environments (HIPAA, SOC 2, PCI)
  • Experience with healthcare data, claims, or adversarial decision domains (fraud, risk, pricing, trading)
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company

What We Do

Recourse Health builds claims-intelligence software for healthcare providers, automating the manual, resource-intensive IDR and out-of-network dispute workflows created by the No Surprises Act. The company uses AI and LLM-powered reasoning to ingest claims, package evidence, and operate end-to-end dispute workflows so providers can pursue disputed payments at scale while preserving auditability and regulatory compliance.

Similar Jobs

Liftoff Logo Liftoff

Product Analyst

AdTech • Artificial Intelligence • Big Data • Machine Learning • Marketing Tech • Mobile • Software
Easy Apply
Remote
2 Locations
645 Employees
126K-170K Annually

Affirm Logo Affirm

Senior Product Manager

Big Data • Fintech • Mobile • Payments • Financial Services
Easy Apply
Remote
United States
2200 Employees
173K-255K Annually

RTB House Logo RTB House

Field Marketing Manager

AdTech • Artificial Intelligence • Big Data • Digital Media • eCommerce • Machine Learning • Marketing Tech
Remote
United States
1300 Employees
110K-130K Annually

Samsara Logo Samsara

Product Manager

Artificial Intelligence • Cloud • Computer Vision • Hardware • Internet of Things • Software
Easy Apply
Remote or Hybrid
6 Locations
4000 Employees
166K-196K Annually

Similar Companies Hiring

Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
Kepler  Thumbnail
Fintech • Software
New York, New York
6 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account