Senior ML Data Engineer - Matching & Recommendation

Reposted 3 Hours Ago
Be an Early Applicant
Nashville, TN, USA
Hybrid
Senior level
Healthtech • Information Technology
The Role
Own and advance a production physician-organization matching engine from rules and weights to semantic matching, embeddings, and vector search. Build rigorous precision/recall evaluation, entity resolution across 8M+ NPI records, and a reliable lightweight data platform using healthcare data. Independently own ingestion, enrichment, modeling, serving, and architecture decisions while improving recommendation quality and supporting customer-facing reporting.
Summary Generated by Built In

Physicians make some of the biggest decisions of their careers with almost no trustworthy information: job posts, recruiter calls, and vague promises about culture. Tessellate is building something better: a physician-first platform that helps doctors see where they'd actually fit, using real clinical data instead of recruiter intuition.

At the center of the product is a matching engine that pairs physicians with organizations. Making it genuinely intelligent is your job.

 

About Tessellate

We're an early-stage healthcare tech company, physicians-first in everything we do. We combine claims data, provider and facility data, physician preferences, and AI-assisted enrichment to help doctors evaluate opportunities with more clarity and less noise. For employers, we're a better way to understand fit; no placement fees, no transactional recruiting games. This isn't a whiteboard idea. We have a working product, real users, a live matching engine, and early commercial demand.

 

The Work

Our matching engine works today, but it's mostly rules and weights. Over the next year we want to make it genuinely smart; real semantic matching and vector search over physicians, organizations, and clinical profiles - with the rigor to actually measure whether it's getting better. That's the heart of this role.

Two things make it hard. The matching is only as good as the identity underneath it: if one health system shows up as three fragmented records, you get three bad embeddings instead of one good one. And healthcare data is a mess, fragmented across sources that don't agree, so reasoning about fit from it is a genuinely deep problem.

You'll own:

  • The matching engine. Take it from rules-and-weights toward semantic matching and vector search. Decide how we represent physicians and organizations as embeddings, and make the recommendations measurably better.
  • Match quality as a discipline. Precision/recall, evaluation sets, real comparison so "better" means something.
  • The identity layer underneath it. Entity resolution across 8M+ NPI records into one trustworthy record per organization. The bar is "good enough for the model to trust," not perfection, and knowing the difference.
  • The data platform, kept simple. Ingest, master, enrich, serve - reliable enough to trust without babysitting. Our source data refreshes roughly quarterly, so we don't want heavyweight batch tooling the cadence doesn't justify.

 

Our Stack

Python 3.12 · DuckDB · Athena + Glue · S3 / Parquet · Aurora PostgreSQL · AWS (Lambda, EC2 Graviton, ECS) · Terraform · GitHub Actions

Vector search and embedding infrastructure are still open decisions, likely among your first. The platform is young, which means little legacy and real architecture calls that are genuinely yours.

 

You're Probably a Fit If You

  • Have built and shipped matching, ranking, or recommendation systems in production, owned a model that had to get measurably better, not just run.
  • Have real applied ML depth: you can frame a problem as precision/recall or ranking, build and validate a model rigorously, and know when a heuristic beats a model.
  • Know embeddings and vector search in practice, or clearly can and want to get us there.
  • Write strong Python and expert SQL, with hands-on time in a columnar/OLAP engine at scale (DuckDB, Spark, Trino/Presto/Athena, BigQuery, Snowflake).
  • Can own your own pipeline without a platform team behind you, but you're not looking for a job that's mostly building DAGs.
  • Have first-principles instincts for entity resolution, you get why identity quality makes or breaks everything above it, and can invent and test approaches rather than reach for a framework.
  • Are comfortable being the data/ML team: small company, no handoffs, you ship and own it.

 

Nice to Have

  • Healthcare data: NPPES/NPI, NUCC taxonomy codes, claims, CMS files, provider directories.
  • Production LLM/RAG or agentic AI, and honest views on where it helps versus where it's theater.
  • Vector database experience (pgvector and the tradeoffs of dedicated stores vs. Postgres-native).
  • Customer-facing reporting: dashboards or reports people outside engineering rely on. Our customers increasingly want this; BI tooling (Looker, Metabase, Mode, Power BI) helps.
  • Graph-based entity clustering and hierarchy inference.
  • Early-stage startup experience.

 

First 90 Days

1.  Get deep on the current model and data, and give us an honest read on where match quality is strong, weak, and why.

2.  Stand up a first vector-search prototype with a real way to measure it against what we have now.

3.  Tell us whether organization identity is good enough for the model to trust, or quietly degrading it.

4.  Propose where the matching engine goes next, and how we'll know it's working.

 

Why This One's Worth It

Most ML jobs are tuning a model nobody ships, or maintaining someone else's DAGs. This is the model at the center of a real product, on a hard and messy dataset, with room to take it somewhere genuinely better and the data underneath it yours to shape. You'd build from a working platform, not a blank page, and directly shape a product that decides where physicians spend their careers.

We're early, and startups aren't for everyone. But you'll learn a lot, own work that matters, and share in the upside if we win.

 

Details

  • Location. Nashville, TN, hybrid preferred. We'll consider strong remote candidates with directly relevant matching/ML or healthcare-data experience.
  • Experience. Senior-level: we care about demonstrated depth in matching and applied ML more than a year count.
  • Work authorization. You must be authorized to work in the US without sponsorship, now and in the future. We are not able to sponsor visas for this role. No agency submissions.
  • Compensation. Competitive base benchmarked to Midwest/South markets and set by experience, plus meaningful early-stage equity.
  • Reports to. Head of Product Engineering.

Skills Required

  • Production experience building and shipping matching, ranking, or recommendation systems
  • Applied machine learning expertise, including precision/recall, ranking, model validation, and heuristic evaluation
  • Practical experience with embeddings and vector search, or the ability to implement them effectively
  • Strong Python programming and expert SQL skills
  • Hands-on experience with a columnar or OLAP engine at scale, such as DuckDB, Spark, Trino, Presto, Athena, BigQuery, or Snowflake
  • Ability to independently own data pipelines without a dedicated platform team
  • First-principles experience with entity resolution and identity quality
  • Ability to operate as the primary data and machine learning team member in an early-stage company
  • Authorization to work in the United States without current or future sponsorship
  • Healthcare data experience, including NPPES/NPI, NUCC taxonomy, claims, CMS files, or provider directories
  • Production LLM, RAG, or agentic AI experience
  • Vector database experience, including pgvector and database tradeoffs
  • Customer-facing reporting or BI dashboard experience with tools such as Looker, Metabase, Mode, or Power BI
  • Graph-based entity clustering and hierarchy inference experience
  • Early-stage startup experience
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
6 Employees
Year Founded: 2025

Similar Jobs

Tessellate Logo Tessellate

Data Engineer

Healthtech • Information Technology
Hybrid
Nashville, TN, USA
6 Employees

Luxury Presence Logo Luxury Presence

Senior Product Designer

Marketing Tech • Real Estate • Software • PropTech • SEO
Easy Apply
Remote or Hybrid
United States
500 Employees

Capco Logo Capco

Software Engineer

Fintech • Professional Services • Consulting • Energy • Financial Services • Cybersecurity • Generative AI
Remote or Hybrid
US
6000 Employees
30-30 Hourly

HiBob Logo HiBob

Enterprise Account Executive

HR Tech • Information Technology • Professional Services • Sales • Software
Remote or Hybrid
US
1350 Employees
123K-158K Annually

Similar Companies Hiring

OneImaging Thumbnail
Healthtech
Miami, FL
62 Employees
NODA AI Thumbnail
Artificial Intelligence • Information Technology • Software • Cybersecurity
Sydney, AU
54 Employees
Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account