Senior CE Data Engineer

Sorry, this job was removed at 04:16 a.m. (UTC) on Monday, Aug 31, 2026
Be an Early Applicant
Indianapolis, IN, USA
In-Office
63K-150K Annually
Senior level
Healthtech • Biotech • Pharmaceutical
The Role
Build and operate Lilly’s covered-entity pharmacy data platform on Databricks. Develop medallion architecture pipelines, data contracts, identity crosswalks, ingestion workflows, security controls, and governed data products handling PHI within a HIPAA trust boundary. Implement CI/CD, data quality, testing, monitoring, tokenization, and access controls while partnering with architects, analytics engineers, product owners, and policy engineers. Translate architectural designs into reliable, scalable production code and support governed AI and agent data consumption.
Summary Generated by Built In

At Lilly, the work is demanding because patients are waiting. We unite caring with discovery to help make life better for people around the world, knowing that every decision, every detail, and every day matters. Headquartered in Indianapolis, Indiana, our over 50,000 employees around the globe take on complex challenges to discover and deliver life-changing medicines, strengthen how health is understood and managed, and support the communities we serve. This is hard, urgent, selfless work—but it’s work worth doing. If you’re driven by purpose and ready to bring your best to work that truly matters for patients, we invite you to join us. 


Organization Overview

Lilly CIA powers the enterprise with governed, AI-ready data products, a converged semantic layer, and the identity and policy infrastructure behind Lilly's digital experiences. Within CIA, the LillyDirect Covered Entity (CE) Pharmacy pod builds and operates the data platform that sits inside the pharmacy's covered-entity trust boundary — handling identified PHI under HIPAA with strict isolation from parent-side systems. This role implements that platform on Databricks.

Position Summary

The CE Data Engineer builds and operationalizes the CE Pharmacy data platform on Databricks — implementing the medallion pipelines, data contracts, identity crosswalk logic, and isolation controls that the CE Data Architect designs. This is a hands-on engineering role: writing production pipelines, evaluating and applying Databricks capabilities and integration patterns, and holding the line on data quality, testing, and monitoring so the platform is reliable, auditable, and built to scale. Works in close, day-to-day partnership with the CE Data Architect and the Associate Director of CE Data Engineering — turning architectural intent into running code — and with Analytics Engineers and the DataHub Product Owner to ship governed data products.

Key Responsibilities
  • Design and build scalable, efficient Databricks pipelines implementing the CE Architect's canonical data models across the medallion (bronze → silver → gold), scoped entirely within the CE trust boundary.

  • Evaluate and apply Databricks capabilities and integration patterns — Unity Catalog, Delta Lake, Databricks Workflows, serverless compute, Lakebase, ingestion connectors — selecting the right tool for each pipeline given performance, cost, and scalability constraints.

  • Implement the identity crosswalk and PHI classification logic defined by the CE Architect (Reltio MDM, Auth0/Passport, Datavant tokens, Scriptly, Genesys, Prescryptive, Transcend), safeguarding merge/split integrity in code.

  • Build to ODCS data contracts — implement schema, quality, freshness/SLA, and lineage requirements as enforced pipeline logic, not documentation.

  • Implement row/column-level security, masking, and tokenization boundaries so PHI isolation is enforced at the platform layer, in partnership with the Policy-as-Code Engineer's OPA/Rego policies.

  • Own end-to-end pipeline development lifecycle — requirements to prototyping to production deployment and maintenance — for CE data products, in partnership with Analytics Engineers and the DataHub Product Owner.

  • Build and maintain CI/CD pipelines (Github Actions [CI], Git-based promotion dev → test → prod) for CE data, contract, policy, and agent artifacts.

  • Automate data ingestion and product creation to reduce manual pipeline maintenance and onboarding time for new CE data sources.

  • Partner with the CE Data Architect on reference architecture and patterns, providing implementation feedback that keeps designs buildable and performant at scale.

  • Build the data pipelines that lets CE Skills and agents consume governed data — ensuring PHI classification and consent travel with the data into agentic consumption paths.

Basic Requirements
  • Bachelor's degree in Computer Science, Engineering, Information Technology, or similar degree.

  • Experience in data engineering, with a focus on building production data pipelines and data products.

  • Advanced SQL and Python; hands-on Databricks / Unity Catalog fluency (notebooks, Delta Lake, catalog/schema/grants, Workflows).

  • Qualified applicants must be authorized to work in the United States on a full-time basis. Lilly will not provide support for or sponsor work authorization or visas for this role, including but not limited to F-1 CPT, F-1 OPT, F-1 STEM OPT, J-1, H-1B, TN, O-1, E-3, H-1B1, or L-1.

Additional Skills / Preferences
  • Experience defining and executing data ingestion pipelines at enterprise scale.

  • Working knowledge of data governance, classification, and access control (RBAC/ABAC, row- and column-level security).

  • Proficiency with Git-based CI/CD workflows (Git Actions or equivalent) for versioned data and pipeline artifacts.

  • Excellent problem-solving skills; ability to translate architectural designs into working, tested pipelines.

  • Good communication and collaboration skills — able to work effectively with architects, product owners, and multi-functional partners.

  • Experience in regulated or healthcare data environments; familiarity with HIPAA, PHI handling, and covered-entity constructs.

  • Data contract frameworks (ODCS or comparable) and policy-as-code (OPA/Rego) awareness.

  • Tokenization / de-identification pattern implementation (Datavant or comparable).

  • PySpark and transformation-as-code (DLT); test frameworks (pytest, DQX / Great Expectations).

  • Infrastructure-as-code (Terraform) for provisioning Unity Catalog objects and grants.

  • MDM / identity resolution implementation experience (Reltio or comparable).

  • Experience with Agile/Scrum methodologies and Jira.

  • Experience building or integrating AI Skills and agents against governed data platforms.

Lilly is dedicated to helping individuals with disabilities to actively engage in the workforce, ensuring equal opportunities when vying for positions. If you require accommodation to submit a resume for a position at Lilly, please complete the accommodation request form (https://careers.lilly.com/us/en/workplace-accommodation) for further assistance. Please note this is for individuals to request an accommodation as part of the application process and any other correspondence will not receive a response.


Lilly is proud to be an EEO Employer and does not discriminate on the basis of age, race, color, religion, gender identity, sex, gender expression, sexual orientation, genetic information, ancestry, national origin, protected veteran status, disability, or any other legally protected status.


Our employee resource groups (ERGs) offer strong support networks for their members and are open to all employees. Our current groups include: Africa, Middle East, Central Asia (AMECA), Black Employees at Lilly (BE@Lilly), Chinese Culture Network (CCN), EnAble, Evolve, Lilly Indian Network (LIN), Organization of Latinx at Lilly (OLA), Pride (LGBTQ+ Allies), Veterans Leadership Network (VLN) and Women’s Initiative for Leading at Lilly (WILL).


Actual compensation will depend on a candidate’s education, experience, skills, and geographic location.  The anticipated wage for this position is

$63,000 - $149,600

Full-time equivalent employees also will be eligible for a company bonus (depending, in part, on company and individual performance). In addition, Lilly offers a comprehensive benefit program to eligible employees, including eligibility to participate in a company-sponsored 401(k); pension; vacation benefits; eligibility for medical, dental, vision and prescription drug benefits; flexible benefits (e.g., healthcare and/or dependent day care flexible spending accounts); life insurance and death benefits; certain time off and leave of absence benefits; and well-being benefits (e.g., employee assistance program, fitness benefits, and employee clubs and activities).Lilly reserves the right to amend, modify, or terminate its compensation and benefit programs in its sole discretion and Lilly’s compensation practices and guidelines will apply regarding the details of any promotion or transfer of Lilly employees.

#WeAreLilly

Skills Required

  • Bachelor’s degree in Computer Science, Engineering, Information Technology, or a similar field
  • Experience in data engineering, focused on building production data pipelines and data products
  • Advanced SQL proficiency
  • Advanced Python proficiency
  • Hands-on Databricks and Unity Catalog experience, including notebooks, Delta Lake, catalog/schema/grants, and Workflows
  • Authorization to work full-time in the United States without employer sponsorship
  • Enterprise-scale data ingestion pipeline experience
  • Data governance, classification, and access-control experience, including RBAC/ABAC and row- and column-level security
  • Git-based CI/CD workflow experience for versioned data and pipeline artifacts
  • Experience working with regulated or healthcare data, including HIPAA, PHI, and covered-entity constructs
  • Experience with data contract frameworks such as ODCS
  • Policy-as-code experience or awareness, including OPA/Rego
  • Tokenization or de-identification implementation experience, such as Datavant
  • PySpark and transformation-as-code experience, including Delta Live Tables
  • Experience with testing or data-quality frameworks such as pytest, DQX, or Great Expectations
  • Infrastructure-as-code experience with Terraform for Unity Catalog objects and grants
  • MDM or identity-resolution implementation experience, such as Reltio
  • Agile/Scrum and Jira experience
  • Experience integrating AI skills or agents with governed data platforms

Eli Lilly and Company Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Eli Lilly and Company and has not been reviewed or approved by Eli Lilly and Company.

  • Retirement Support Feedback suggests long-term savings are bolstered by a defined-benefit pension alongside a company 401(k) match and retiree health options. These elements make total compensation feel strong beyond base salary.
  • Leave & Time Off Breadth Feedback suggests paid time off is expansive, with substantial vacation, company shutdown days, and milestone time. This breadth of leave is viewed as a meaningful part of overall rewards.
  • Parental & Family Support Feedback suggests family-building and caregiving support are robust, including paid parental leave, adoption or surrogacy assistance, and backup care. These programs enhance the perceived value of benefits across life stages.

Eli Lilly and Company Insights

Similar Jobs

Pfizer Logo Pfizer

Director, Quality Management Systems/Validation & Compliance Lead

Artificial Intelligence • Healthtech • Machine Learning • Natural Language Processing • Biotech • Pharmaceutical
Remote or Hybrid
7 Locations
121990 Employees
163K-272K Annually

Liberty Mutual Insurance Logo Liberty Mutual Insurance

Casualty Senior Claims Manager — Casualty Specialized Claims

Artificial Intelligence • Fintech • Insurance • Marketing Tech • Software • Analytics
Remote or Hybrid
9 Locations
40000 Employees
179K-322K Annually

CrowdStrike Logo CrowdStrike

Infrastructure Engineer

Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity
Remote or Hybrid
USA
11000 Employees
100K-155K Annually

CrowdStrike Logo CrowdStrike

Senior Platform Engineer

Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity
Remote or Hybrid
USA
11000 Employees
140K-215K Annually
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Indianapolis, IN
39,451 Employees
Year Founded: 1876

What We Do

Eli Lilly and Company engages in the discovery, development, manufacture, and sale of products in pharmaceutical products business segment. For more than a century, we have stayed true to a core set of values – excellence, integrity, and respect for people – that guide us in all we do: discovering medicines that meet real needs, improving the understanding and management of disease, and giving back to communities through philanthropy and volunteerism.

Similar Companies Hiring

Sailor Health Thumbnail
Healthtech • Social Impact • Telehealth
New York City, NY
20 Employees
Granted Thumbnail
Artificial Intelligence • Healthtech • Insurance • Mobile • Financial Services
New York, New York
23 Employees
OneImaging Thumbnail
Healthtech
Miami, FL
62 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account