Senior Data Engineer

Posted Yesterday
Chicago, IL, USA
Hybrid
125K-180K Annually
Senior level
Artificial Intelligence • Big Data • Healthtech • Machine Learning • Analytics • Biotech • Generative AI
Tempus is a technology company leading the adoption of AI to advance precision medicine and patient care.
The Role
Build and operate scalable healthcare data pipelines, dbt models, event-driven services, APIs, and Google Cloud infrastructure. Model and govern multimodal EHR, genomic, and imaging data for AI agents across federated hospital networks. Ensure data quality, lineage, observability, security, HIPAA compliance, and production reliability while contributing to TypeScript application services and agentic workflows.
Summary Generated by Built In

Passionate about precision medicine and advancing the healthcare industry?

Recent advancements in underlying technology have finally made it possible for AI to impact clinical care in a meaningful way. Tempus' proprietary platform connects an entire ecosystem of real-world evidence to deliver real-time, actionable insights to physicians, providing critical information about the right treatments for the right patients, at the right time.

We are building the Patient Evaluation Engine: a high-scale, multi-modal healthcare platform where autonomous AI agents reason over clinical data to drive real-time clinical evaluation across federated networks of hospitals. We are looking for a Senior Data Engineer to build and own the data platform underneath it — the pipelines, models, and services that make EHR records, genomic results, and cardiovascular imaging discoverable, trustworthy, and usable by agents.

This is a data engineering role at its core, and it asks for two things beyond the usual scope. First, you should be a capable software engineer: the person who builds the pipeline here is the person who writes the service that exposes it, and you will regularly work in our TypeScript application and service code rather than handing that off. Second, you should know cloud infrastructure well, specifically Google Cloud — you will make real decisions about how this platform is deployed, scaled, secured, and paid for, not just what runs on it.

Our goal is to move beyond static data warehousing toward a dynamic, "agent-ready" data fabric that supports real-time clinical evaluation at enterprise scale, in a HIPAA-regulated environment. The platform is early and much of it is still being built, which is why we are looking for someone with high ownership and a strong self-starting instinct rather than someone waiting for a fully specified backlog.

What You'll Do

  • Build the pipelines that feed the agents. Develop the systems that fetch, parse, and serve both structured and unstructured data in formats optimized for real-time inference — spanning clinical EHR records, high-throughput genomics (NGS), and cardiovascular imaging (Echo, Cath, ECG).

  • Own the warehouse and its transformations. Build and maintain our dbt models on BigQuery, along with the tests, documentation, and SQL standards that keep a growing model layer trustworthy.

  • Model the multi-modal patient record. Shape the data model across those domains, applying normalized and dimensional design as each one demands, and write the code that enforces it.

  • Move data through event-driven services. Build and operate the Pub/Sub topics, subscriptions, and dead-letter handling that connect ingestion, evaluation, and result delivery, with the retry and idempotency behavior that reliability at scale requires.

  • Make the data agent-ready. Build the data access patterns and metadata layers that let AI agents autonomously discover, query, and reason over structured and unstructured datasets, and the retrieval services those agents call.

  • Write the software on top. Build the TypeScript services and APIs that handle agent input and output and coordinate specialized agents, meeting the platform's performance and scalability demands. You are expected to be comfortable in the application codebase, not only in the data layer.

  • Own the infrastructure your platform runs on. Extend and operate the platform's Google Cloud footprint in Terraform — BigQuery datasets, Pub/Sub, Cloud SQL, Cloud Storage, Memorystore, Secret Manager, and the service-account and IAM model that governs access to clinical data.

  • Scale across hospital networks. Build for federated networks of hospitals: multi-tenancy, high availability, and performance across hybrid on-prem and cloud environments built for sensitive health-system integrations.

  • Guarantee ground truth. Implement automated solutions to monitor data quality and lineage with strict traceability back to source systems, ensuring "ground truth" for agentic evaluations.

  • Instrument for trust. Build the observability, error tracking, and human-in-the-loop checkpoints that make automated clinical evaluation transparent and debuggable.

  • Raise the standard around you. Partner with clinical, analytics, and platform engineering teams on data modeling standards, governance, and practices for maintaining data integrity in a HIPAA-regulated environment.

How You Work

We care about these as much as the technical checklist.

  • High ownership. You own what you build all the way into production — you care whether it stays up, you chase root causes instead of symptoms, and you do not treat the deploy boundary as the end of your responsibility.

  • Self-starter. The problem space is genuinely open. You are comfortable identifying the most valuable next thing and starting on it without a fully specified ticket, and you surface ambiguity early rather than stalling on it.

  • Collaborative. You work directly with clinical, analytics, and platform engineering partners. You write things down, you explain trade-offs to non-specialists, and you make the people around you faster.

  • Quick to add impact and value. You bias toward shipping something real and incremental early over long design cycles, and you look for the change that moves the platform now.

Our Stack

You will not have used all of this, and we do not expect you to have. It is here so you know what you would be working in.

  • Warehouse and transformation: BigQuery, dbt

  • Operational data stores: Cloud SQL (PostgreSQL), Memorystore (Redis), Cloud Storage

  • Messaging: Pub/Sub with dead-letter queues

  • Healthcare data: Google Cloud Healthcare API FHIR stores, HL7/FHIR, DICOM, Avro

  • Languages: Python for data pipelines and transforms; TypeScript on Node for platform services and APIs

  • Application frameworks: NestJS, TypeORM

  • Infrastructure: Terraform, Docker, Secret Manager, service-account and IAM-based access control

  • Decisioning: GoRules ZEN engine for versioned decision models

  • Cloud: primarily Google Cloud, with some AWS at the edges

What We're Looking For

  • Data engineering depth. Proven track record building and operating production data pipelines that handle structured and unstructured data at scale, with real ownership of reliability and correctness.

  • Google Cloud fluency. Hands-on experience designing and running workloads on GCP — BigQuery, Pub/Sub, Cloud Storage, Cloud SQL, and Secret Manager — including the IAM and service-account model that controls access to sensitive data.

  • Analytics engineering. Strong SQL and practical experience with dbt or an equivalent transformation framework, including testing, documentation, and managing a model layer as it grows.

  • Infrastructure practice. Comfort owning infrastructure as code in Terraform, working in containers, and taking responsibility for the operational characteristics of what you deploy.

  • Software engineering ability. You write production-quality application and service code, not just pipeline glue — including APIs, tests, and the design work that goes with them.

  • Python and TypeScript. Python strong enough for production pipelines as well as hands-on data profiling and debugging, plus enough TypeScript or another statically typed language to work confidently in our service and application code.

  • Event-driven systems. Experience with pub/sub or queue-based architectures and the failure modes that come with them — retries, ordering, idempotency, and dead-letter handling.

  • Interoperability standards. Working knowledge of HL7, FHIR, and Epic/Cerner data structures, along with DICOM and genomic data formats.

  • Regulatory fluency. Familiarity with building secure, resilient systems under HIPAA and SOC 2.

Experience Requirements

  • Total Professional Experience: 5+ years building data-intensive software systems in production.

  • Data Engineering: 3+ years focused on data engineering, pipeline ownership, or data modeling, ideally in the healthcare or life sciences domain.

  • Cloud Infrastructure: 2+ years hands-on building and operating on Google Cloud, with demonstrated ownership of infrastructure decisions rather than consuming someone else's.

  • Healthcare Domain: 2+ years in HIPAA-regulated environments, with hands-on exposure to EMR integrations (Epic, Cerner) and healthcare data standards.

  • AI/ML Orchestration: 1+ years hands-on building with Large Language Models — agentic workflows, RAG, or autonomous tool use.

  • Data at Scale: Demonstrated experience managing structured (SQL, NoSQL) and unstructured data at a scale of millions of records, ensuring data integrity for downstream AI consumption.

Education

  • Primary Requirement: Bachelor's degree in Computer Science, Software Engineering, Data Science, Health Informatics, or a related technical field.

  • Preferred: Master's degree or Ph.D. in Computer Science (AI/ML or distributed systems focus) or Biomedical Informatics.

  • Alternative Background: Equivalent professional experience — including a portfolio of significant open-source contributions or industry-recognized technical writing — will be considered.

Bonus Points

  • Google Cloud Healthcare API. Direct experience with managed FHIR or DICOM stores.

  • Specialized clinical data. Direct experience with OMOP, DICOM, genomic data models, or longitudinal patient records.

  • Kubernetes. Experience running containerized workloads on Kubernetes.

  • AI data engineering. Experience with vector databases (Pinecone, Weaviate, pgvector) or graph databases to support RAG and agentic memory.

  • AWS. Experience with AWS services alongside GCP in a multi-cloud environment.

  • Advanced modeling techniques. Experience with Data Vault 2.0, Master Data Management, or comparable enterprise modeling methodologies.

CHI: $125,000-$180,000

The expected salary range may vary for other locations. Actual salary may vary based on qualifications and experience. Tempus offers a full range of benefits, which may include incentive compensation, restricted stock units, medical and other benefits depending on the position.

We are an equal opportunity employer. We do not discriminate on the basis of race, religion, color, national origin, gender, sexual orientation, age, marital status, veteran status, or disability status. 

Skills Required

  • 5+ years building data-intensive software systems in production
  • 3+ years focused on data engineering, pipeline ownership, or data modeling
  • 2+ years hands-on building and operating workloads on Google Cloud
  • 2+ years working in HIPAA-regulated environments
  • Hands-on exposure to EMR integrations, including Epic or Cerner, and healthcare data standards
  • 1+ year building with large language models, agentic workflows, RAG, or autonomous tool use
  • Experience managing structured and unstructured data at a scale of millions of records
  • Strong SQL and practical experience with dbt or an equivalent transformation framework
  • Production experience with Python and TypeScript or another statically typed language
  • Experience with event-driven or queue-based architectures, retries, ordering, idempotency, and dead-letter handling
  • Experience with Terraform, containers, and infrastructure ownership
  • Familiarity with HIPAA and SOC 2 secure system requirements
  • Bachelor's degree in Computer Science, Software Engineering, Data Science, Health Informatics, or a related technical field
  • Master's degree or Ph.D. in Computer Science or Biomedical Informatics
  • Direct experience with Google Cloud Healthcare API, managed FHIR stores, or DICOM stores
  • Experience with OMOP, DICOM, genomic data models, or longitudinal patient records
  • Experience running containerized workloads on Kubernetes
  • Experience with vector or graph databases supporting RAG and agentic memory
  • AWS experience in a multi-cloud environment
  • Experience with Data Vault 2.0, Master Data Management, or comparable enterprise modeling methodologies

What the Team is Saying

Rachel
Louis
Anita
Alexis
Hala
Aaron
Alexis
Ash
Emma
Anita
Mile

Tempus AI Compensation & Benefits Highlights

  • Healthcare Strength Health coverage is described as comprehensive, including medical, dental, vision, life and disability insurance, plus mental‑health resources and HSA/FSA options, with core items employer‑verified in mid‑2026.
  • Leave & Time Off Breadth Paid time off and parental leave are consistently part of the package, with parental leave commonly cited around three months and related pages noted as recently employer‑verified.
  • Wellbeing & Lifestyle Benefits Everyday perks such as commuter benefits, on‑site meals and coffee programs, gym discounts, charitable matching, ERGs, pet insurance, and team events are widely highlighted, with some amenities available in select offices.

Tempus AI Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Chicago, IL
3,775 Employees
Year Founded: 2015

What We Do

We bring together one of the world’s largest libraries of multimodal clinical and molecular data with a robust suite of AI tools to help physicians personalize care in real time, connect patients with therapies and clinical trials, and enable partners to accelerate discovery and development of new treatments. With ~8 million de-identified research records and 350+ petabytes of data, Tempus partners with more than half of U.S. oncologists and the majority of the top 20 global pharma companies. Our teams are pioneering work across oncology, neurology, psychiatry, cardiology, and beyond—transforming how care is delivered and therapies are developed. At Tempus, every role contributes to our mission: to help each patient benefit from the experiences of those who came before. For more information, visit tempus.com.

Why Work With Us

We’re looking for people who can change the world. People who question the status quo and refuse to shy away from tough problems. For builders who are never done building, and the learners who are never done learning. Passionate individuals with undying curiosity who want to take on one of the greatest challenges humanity has ever faced—head on.

Gallery

Gallery
Gallery
Gallery
Gallery
Gallery
Gallery
Gallery
Gallery

Tempus AI Offices

Hybrid Workspace

Employees engage in a combination of remote and on-site work.

Most of the team follows a hybrid policy, with some roles allowing for a fully remote arrangement and some roles being onsite only.

Typical time on-site: 3 days a week
Company Office Image
HQChicago - Tempus Headquarters & Lab
Company Office Image
RTP - Tempus Lab
Company Office Image
Boston - Tempus Office
Company Office Image
Seattle - Tempus Office
Company Office Image
Lewisburg - Tempus Office
Company Office Image
Madison - Tempus Office
Company Office Image
Milwaukee - Tempus Office
Company Office Image
New York City - Tempus Office
Company Office Image
Atlanta - Tempus Lab
Company Office Image
Bay Area - Tempus Office
Company Office Image
Washington DC - Tempus Office
Learn more

Similar Jobs

Tempus AI Logo Tempus AI

TIME Operations Alliance Manager

Artificial Intelligence • Big Data • Healthtech • Machine Learning • Analytics • Biotech • Generative AI
Remote or Hybrid
Illinois, USA
3775 Employees
90K-118K Annually

Tempus AI Logo Tempus AI

Scientist

Artificial Intelligence • Big Data • Healthtech • Machine Learning • Analytics • Biotech • Generative AI
Hybrid
4 Locations
3775 Employees
90K-150K Annually

Tempus AI Logo Tempus AI

Scientist

Artificial Intelligence • Big Data • Healthtech • Machine Learning • Analytics • Biotech • Generative AI
Hybrid
3 Locations
3775 Employees
155K-210K Annually

Tempus AI Logo Tempus AI

Scientist

Artificial Intelligence • Big Data • Healthtech • Machine Learning • Analytics • Biotech • Generative AI
Hybrid
4 Locations
3775 Employees
140K-210K Annually

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account