Staff Data Engineer- RWE

Reposted 12 Days Ago
New York, NY, USA
Hybrid
170K-190K Annually
Senior level
Artificial Intelligence • Big Data • Healthtech • Database • Business Intelligence
We are creating a healthier future by connecting the world with the right doctors.
The Role
As a Staff Data Engineer, lead significant data projects, ensuring scalable architectures and optimized data pipelines, while mentoring others and collaborating cross-functionally.
Summary Generated by Built In

H1's mission is to connect the world to the right doctor. We've built one of the largest healthcare datasets in the world: profiles of more than 10 million physicians, assembled from billions of insurance claims, 25 million publications, and nearly 500K clinical trials, with hundreds of sources feeding each doctor's profile.

85% of the top 20 pharma companies use it to decide who should run their clinical trials. Nine of the ten largest US health insurers use it to build their networks. And increasingly, the leading frontier models use our data to help patients find the right doctor. Most people we help never see our name. The engineering behind this is genuinely hard: thousands of sources, no shared identifiers, strict privacy rules, and answers that have to be right.

Data Engineering builds that dataset. The team owns the pipelines that turn raw sources into the physician profiles our customers query: hundreds of terabytes on PySpark, EMR, and Hudi. The day-to-day is deciding how the truth gets computed: which source wins when two disagree, how the same doctor gets matched across systems that have never heard of each other, and how it all stays fresh and affordable at scale. Our data is growing faster than our team, so we're hiring engineers to own the hardest pipelines end to end.

WHAT YOU'LL DO AT H1
You'll be the senior-most data engineer on the Real World Evidence (RWE) team, owning the claims pipelines that are becoming H1's largest and most important datasets. This is a hands-on role: you'll be in the codebase every day.
 
What's claims data? Nearly every doctor visit in America generates an insurance claim: who was treated, for what, by whom. Billions of them, from thousands of sources that share no common format or identifier. Assembled correctly, they reveal how medicine is actually practiced, which no other dataset can show.
 
You will:
- Claims pipelines: turn billions of US insurance claims into the physician-level data our customers query every day. This is the largest dataset H1 has ever built, and it's yours.
- Patient journeys: reconstruct treatment timelines from claims and health records, so customers can see how patients actually move through care and which doctors to reach.
- Performance and cost: own the hundreds-of-terabytes Spark workloads end to end, and make them faster, more reliable, and cheaper.
- Roadmap: set the technical direction for RWE data with Product, Data Science, and downstream teams, and make the architecture calls.
- Mentorship: raise the team's bar through design reviews, pairing, and deep domain teaching.
 
ABOUT YOU
You've run data products at a serious scale, and you liked owning all of it: the design, implementation, the tradeoffs, the incident at 2am, the cost curve. You're at your best when the problem is ambiguous and the data is messy, and you'd rather ship a pragmatic call this quarter than a perfect one next year. You've probably never worked in healthcare. That's fine. The engineers who do this well came for the problem.
 
REQUIREMENTS
- 8+ years building production data or backend systems, and you still write code daily and want to keep it that way.
- You've designed and run Spark pipelines at scale and owned the performance, cost, and reliability tradeoffs yourself.
- Leadership: you've driven multi-quarter, cross-team technical work without formal authority.
- Stack: ours is PySpark on EMR, Hudi/Delta, Airflow, with SQL and Python everywhere. Deep in something comparable, fast ramp on the rest.
- Operations: you can deploy, debug, and un-break distributed workloads in the cloud without waiting for another team.
- Nice to have: streaming (Kafka or Kinesis); healthcare, claims, or other regulated-data experience.
 - Experience using AI-assisted coding tools (e.g., GitHub Copilot, Claude Code) to accelerate development while maintaining quality is encouraged 
 
COMPENSATION
This role pays $190,000 to $220,000 per year, based on experience, in addition to stock options.

Anticipated role close date: 9/15/2026

H1 OFFERS
- Full suite of health insurance options, in addition to generous paid time off
- Pre-planned company-wide wellness holidays
- Retirement options
- Health & charitable donation stipends
- Impactful Business Resource Groups
- Flexible work hours & the opportunity to work from anywhere
- The opportunity to work with leading biotech and life sciences companies in an innovative industry with a mission to improve healthcare around the globe
 
 
H1 is proud to be an equal opportunity employer that celebrates diversity and is committed to creating an inclusive workplace with equal opportunity for all applicants and teammates. Our goal is to recruit the most talented people from a diverse candidate pool regardless of race, color, ancestry, national origin, religion, disability, sex (including pregnancy), age, gender, gender identity, sexual orientation, marital status, veteran status, or any other characteristic protected by law.
 
H1 is committed to working with and providing access and reasonable accommodation to applicants with mental and/or physical disabilities. If you require an accommodation, please reach out to your recruiter once you've begun the interview process. All requests for accommodations are treated discreetly and confidentially, as practical and permitted by law.

Skills Required

  • 8+ years as a software, data, or backend engineer building and operating scalable, production-grade systems
  • Experience with large-scale data processing or scalable distributed backend systems
  • Strong proficiency in SQL, including writing and optimizing complex queries over large datasets
  • Strong programming experience in Python or a modern language with ability to ramp up in Python
  • Experience designing systems or large-scale datasets/pipelines with attention to performance, reliability, and maintainability
  • Hands-on experience with modern engineering workflows and tooling such as Git, JIRA, and CI/CD systems
  • Comfort deploying and troubleshooting distributed workloads in cloud environments
  • Experience with workflow orchestration or job scheduling tools
  • Demonstrated ability to independently drive complex, cross-team technical initiatives
  • Experience with streaming/messaging technologies
  • Background in RWE, healthcare data, or other complex/regulated data domains
  • Experience using AI-assisted coding tools
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: New York , NY
500 Employees
Year Founded: 2017

What We Do

Access to medicine and healthcare is a basic human right. At H1, we believe access to the best healthcare information is also a basic human right, one that will be more important in the 21st century than ever before. Our commitment to creating a healthier future for everyone drives us to build and maintain the most current, accurate, and comprehensive healthcare knowledge base available, as well as the tools and intelligence to extract unparalleled insights to carry global healthcare forward.

Why Work With Us

We’re a team of people building products that help solve difficult problems in healthcare. We work through complex challenges every day, navigating ambiguity, wrestling with uncertainty, and pushing the boundaries of what’s possible–all while caring deeply about one another and the people we seek to help.

Gallery

Gallery

Similar Jobs

H1 Logo H1

Back-end Engineer

Artificial Intelligence • Big Data • Healthtech • Database • Business Intelligence
Hybrid
New York, NY, USA
500 Employees
170K-190K Annually
Hybrid
2 Locations
289097 Employees
Hybrid
2 Locations
289097 Employees

Datadog Logo Datadog

Technical Support

Artificial Intelligence • Cloud • Security • Software • Cybersecurity
Easy Apply
Hybrid
3 Locations
6500 Employees
66K-88K Annually

Similar Companies Hiring

Legora Thumbnail
Artificial Intelligence • Legal Tech • Software
New York, New York
700 Employees
Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account