Senior Data Architect – AWS & Databricks Modernization

Posted Yesterday
Be an Early Applicant
Bengaluru, Bengaluru Urban, Karnataka, IND
In-Office
Senior level
Healthtech • Biotech • Pharmaceutical
The Role
Leads large-scale migration from legacy warehouses and platforms to AWS Databricks lakehouses for clinical and non-clinical data. Builds production pipelines, Delta Lake architectures, governance through Unity Catalog, semantic models, reusable data products, and clinical data-quality frameworks. Owns modernization roadmaps, migration validation, performance and cost optimization, AI-assisted engineering accelerators, compliance automation, and stakeholder communication. The role combines approximately 70% hands-on coding and technical execution with 30% strategy.
Summary Generated by Built In

At Lilly, the work is demanding because patients are waiting. We unite caring with discovery to help make life better for people around the world, knowing that every decision, every detail, and every day matters. Headquartered in Indianapolis, Indiana, our over 50,000 employees around the globe take on complex challenges to discover and deliver life-changing medicines, strengthen how health is understood and managed, and support the communities we serve. This is hard, urgent, selfless work—but it’s work worth doing. If you’re driven by purpose and ready to bring your best to work that truly matters for patients, we invite you to join us. 


CTI MD Tech@Lilly 
Senior Data Architect — AWS & Databricks Modernization at Scale 

Position Description 

About Lilly 

At Lilly, everything we do starts with patients. We unite caring with discovery to make life better for people around the world. Headquartered in Indianapolis, Indiana, our global team of over 50,000 employees work with urgency and purpose to discover and deliver life-changing medicines, strengthen how health is understood and managed, and support the communities we serve. We bring our best to this work because people depend on it. If you're driven by purpose and determined to make a meaningful difference for patients, we invite you to bring your skill and your commitment to Lilly. 

About Technology@Lilly 

At Lilly, technology is not a support function. It is how a global medicine company operates, innovates, and delivers. Lilly in Bengaluru builds the capabilities that make this possible, cloud platforms, AI systems, and automation at enterprise scale, all in service of a purpose that makes this technology work genuinely distinctive, from advancing drug discovery to enabling connected clinical trials to keeping a global medicine company running at the standard patients deserve. 

About the Business Function: 

The Clinical & Non-Clinical Data Organization at Eli Lilly and Company is responsible for the design, build, and operation of enterprise data platforms that power drug discovery, clinical development, and regulatory submissions. Data Hub is building a robust Data Strategy to make Lilly's Clinical and Non-Clinical data AI-ready and audit-ready, delivering scalable, governed, and reusable data products that accelerate how medicines reach patients. Our data engineering organization sits at the intersection of science, technology, and patient impact — connecting Clinical and Non-Clinical data across the full chain, from ingestion to consumption. 

Role: 

Senior Data Architect — AWS + Databricks Modernization at Scale 

The Senior Data Architect (R5) is a hands-on Databricks and lakehouse modernization leader who owns the architecture and execution of large-scale migrations from legacy warehouses and point platforms onto a unified Databricks lakehouse for the Clinical and Non-Clinical data domain. This is a builder role: the architect writes code, builds migration pipelines, and personally ships production lakehouse assets — in addition to defining the modernization roadmap and rollout patterns that scale across dozens of domains. The role also owns the roadmap for semantic modeling and analytics-ready data, builds reusable data products the organization can adopt, drives data quality on clinical data specifically, and defines the metrics that enable data-driven decision making. AI-assisted and agentic ways of working are the expected default. Domain knowledge of the clinical data landscape is an added advantage. This role is split 70% hands-on technical execution including coding and 30% strategy, and shapes the future technology landscape. 

Key Responsibilities: 

Lakehouse Modernization & Migration at Scale (Hands-On) 

 Own the technical roadmap and execution for migrating legacy warehouses, on-prem databases, and point data platforms onto Databricks — sequencing dozens of Clinical and Non-Clinical domains with minimal disruption. 

 Personally build migration pipelines and re-platforming accelerators (schema conversion, historical backfill, dual-run validation, cutover automation) that move data at scale with parity and auditability. 

 Define reusable modernization patterns — landing zone design, medallion (bronze/silver/gold) conventions, workspace/catalog topology — that scale consistently as new domains onboard. 

 Right-size and standardize the Databricks platform footprint across environments (dev/test/prod, multiple workspaces) for cost, performance, and governance at enterprise scale. 

 Establish cutover, rollback, and data-reconciliation practices that de-risk large-scale migrations in a regulated environment. 

Databricks Lakehouse Architecture & Engineering (Hands-On) 

 Architect and build the Clinical/Non-Clinical lakehouse on Databricks, applying medallion design across Delta Lake tables at scale across multiple domains. 

 Personally build and optimize Delta Live Tables (DLT) pipelines, Databricks Workflows, and PySpark/Spark SQL jobs for high-volume ingestion, transformation, and curation. 

 Own Unity Catalog design and rollout — catalogs, schemas, access control, lineage, and data sharing — as the governance backbone across an expanding domain footprint. 

 Tune performance and cost at scale: Photon, cluster policies, job/task orchestration, auto-scaling, and Databricks SQL Serverless warehouses across many concurrent workloads. 

 Package and deploy pipelines using Databricks Asset Bundles and CI/CD; evaluate Lakehouse Federation and cross-workspace patterns for enterprise-wide access. 

AI-Native & Agentic Data Engineering 

 Use AI-assisted and agentic tooling by default — migration gap analysis, pipeline scaffolding, code conversion, and documentation — to accelerate modernization at scale. 

 Apply Databricks Mosaic AI / MLflow and LLM-based tooling to automate schema mapping, ontology alignment, and data-quality scoring during migration. 

 Build reusable AI-assisted accelerators that other engineers and architects reuse to modernize additional domains faster. 

Semantic Modeling, Analytics-Ready Data & Reusable Data Products 

 Define and own the multi-quarter roadmap for semantic modeling and analytics-ready data — ontologies, taxonomies, dimensional models, and semantic layers — across Clinical and Non-Clinical domains. 

 Design and build reusable, governed data products (documented, discoverable, versioned) on the lakehouse that other squads can adopt directly rather than rebuilding equivalents. 

 Own data quality specifically for clinical data — define quality rules, thresholds, and remediation workflows for clinical data sets, and track quality trends over time. 

 Define the metrics and KPIs (data quality scores, product adoption/reuse rates, pipeline reliability, time-to-insight) that enable data-driven decision making across the India data organization, and report progress against the roadmap. 

Governance, Standards & Automation 

 Implement role-based and attribute-based access control and encryption for data at rest and in transit within Unity Catalog, consistently across every migrated domain. 

 Stand up and maintain the data standards platform — the reference implementation, templates, and lineage/quality practices for how Clinical and Non-Clinical data is modeled and migrated. 

 Automate lineage capture, schema-drift detection, and post-migration data-quality validation so the team scales modernization without proportional manual effort. 

Collaboration & Stakeholder Engagement 

 Partner with business SMEs, solution architects, and engineering teams to sequence and translate legacy platform needs into Databricks-based modernization plans. 

 Communicate migration risk, trade-offs, and rollout progress clearly to technical and non-technical stakeholders, and own the India lakehouse modernization roadmap. 

Qualifications Required: 

Required — Lakehouse Modernization at Scale (Must-Have, Hands-On) 

 5+ years hands-on experience migrating or modernizing legacy data warehouses/on-prem platforms (e.g., Oracle, Teradata, SQL Server, Hadoop) onto a Databricks lakehouse in production. 

 Demonstrated experience sequencing and delivering multi-domain migrations at enterprise scale — not a single one-off project. 

 Hands-on experience with schema conversion, historical backfill, dual-run/parallel validation, and cutover automation for large-scale data migrations. 

 Experience standardizing workspace, catalog, and environment topology (dev/test/prod, multi-workspace) across a growing platform footprint. 

Required — Databricks & Lakehouse Engineering (Must-Have) 

 5+ years hands-on building on Databricks: Delta Lake, Delta Live Tables, Unity Catalog, Databricks Workflows, Databricks SQL, and MLflow, with real production pipelines shipped, not design-only. 

 Strong PySpark and Spark SQL engineering skill, including performance tuning, cluster sizing, and cost optimization at scale. 

 Experience with Delta table optimization (Z-ordering, OPTIMIZE/VACUUM, liquid clustering) and Databricks Asset Bundles / Git-based CI/CD / infrastructure-as-code. 

Required — Semantic Modeling, Data Products & Quality 

 Proven experience defining and delivering a semantic modeling / analytics-ready data roadmap: ontologies, taxonomies, dimensional models, and semantic layers, built and shipped, not design-on-paper. 

 Track record designing and shipping reusable, governed data products that other teams adopt, with clear documentation and ownership. 

 Hands-on experience defining and driving data quality frameworks — rules, thresholds, scorecards, remediation — specifically on clinical data sets. 

 Experience defining metrics/KPIs that support data-driven decision making and communicating them to business and technical stakeholders. 

Required — AI-Assisted & Platform Governance 

 Demonstrated, current use of AI-assisted/agentic tools as the default method for migration analysis, pipeline design, and documentation. 

 Hands-on experience with cloud platforms (AWS) and Unity Catalog-based access control, encryption, and lineage/compliance patterns for regulated data. 

Preferred 

 Databricks certifications (Data Engineer Professional, Databricks Certified Associate/Professional); prior experience leading an enterprise-wide legacy-to-lakehouse modernization program in a regulated industry. 

 Domain knowledge of the clinical data landscape (e.g., CDISC/SDTM/ADaM, EDC systems, clinical trial data structures) is an added advantage; familiarity with GxP / 21 CFR Part 11 compliance in a validated data environment. 

Education & Experience: 

Minimum 

 Bachelor's degree in Computer Science, Information Systems, or related discipline. 

 12+ years in data architecture/engineering, including 3+ years hands-on leading Databricks-based lakehouse modernization or migration at scale. 

 Demonstrated track record shipping large-scale migration and modernization work to production on Clinical and/or Non-Clinical data, not solely design-on-paper. 

 Demonstrated, current use of AI-assisted tools in day-to-day design and engineering work, with concrete examples. 

Preferred 

 Master's degree or equivalent in a quantitative or engineering discipline. 

 Experience in a regulated industry (pharma, life sciences, finance) with GxP or equivalent compliance requirements. 

 Prior contribution to clinical or non-clinical data programs supporting IND, NDA, or BLA submissions; working knowledge of the clinical data landscape is an added advantage. 

Preferred: 

 Master's degree or equivalent in a quantitative or engineering discipline. 

 Databricks certifications such as Data Engineer Professional, Databricks Certified Associate/Professional. 

 Prior experience leading an enterprise-wide legacy-to-lakehouse modernization program in a regulated industry. 

 Domain knowledge of the clinical data landscape including CDISC, SDTM, ADaM, EDC systems, and clinical trial data structures. 

 Experience in a regulated industry such as pharma, life sciences, or finance with GxP or equivalent compliance requirements. 

 Prior contribution to clinical or non-clinical data programs supporting IND, NDA, or BLA submissions. 

At Lilly, caring is not only what we do for patients. It is how we work. We believe the people who dedicate themselves to making medicines better deserve an environment that makes their lives better too, one where they are supported, respected, and given the space to do their best work. This is not just a policy. It is who we are. 

Equal Opportunity & Accommodation 

Lilly is dedicated to helping individuals with disabilities to actively engage in the workforce, ensuring equal opportunities when vying for positions. If you require accommodation to submit a resume for a position at Lilly, please complete the accommodation request form at https://careers.lilly.com/us/en/workplace-accommodation for further assistance. Please note this is for individuals to request an accommodation as part of the application process and any other correspondence will not receive a response. 

Lilly is an EEO/Affirmative Action Employer and does not discriminate on the basis of age, race, colour, religion, gender, sexual orientation, gender identity, gender expression, national origin, protected veteran status, disability or any other legally protected status. 

Lilly is dedicated to helping individuals with disabilities to actively engage in the workforce, ensuring equal opportunities when vying for positions. If you require accommodation to submit a resume for a position at Lilly, please complete the accommodation request form (https://careers.lilly.com/us/en/workplace-accommodation) for further assistance. Please note this is for individuals to request an accommodation as part of the application process and any other correspondence will not receive a response.

Lilly does not discriminate on the basis of age, race, color, religion, gender, sexual orientation, gender identity, gender expression, national origin, protected veteran status, disability or any other legally protected status.

#WeAreLilly

Skills Required

  • 5+ years of hands-on experience migrating or modernizing legacy data warehouses or on-premises platforms to Databricks in production
  • Experience sequencing and delivering multi-domain enterprise migrations
  • Hands-on experience with schema conversion, historical backfill, dual-run validation, and cutover automation
  • Experience standardizing workspace, catalog, and environment topology across development, test, and production
  • 5+ years of hands-on Databricks experience with Delta Lake, Delta Live Tables, Unity Catalog, Databricks Workflows, Databricks SQL, and MLflow
  • Strong PySpark and Spark SQL engineering skills, including performance tuning, cluster sizing, and cost optimization
  • Experience with Delta table optimization, Databricks Asset Bundles, Git-based CI/CD, and infrastructure as code
  • Experience defining and delivering semantic modeling and analytics-ready data roadmaps
  • Track record designing and shipping reusable, governed data products adopted by other teams
  • Hands-on experience defining and driving data-quality frameworks for clinical datasets
  • Experience defining metrics and KPIs and communicating them to business and technical stakeholders
  • Current use of AI-assisted or agentic tools for migration analysis, pipeline design, and documentation
  • Hands-on AWS experience and Unity Catalog-based access control, encryption, lineage, and compliance patterns
  • Bachelor's degree in Computer Science, Information Systems, or a related discipline
  • 12+ years in data architecture or engineering, including 3+ years leading Databricks lakehouse modernization or migration at scale
  • Production experience with large-scale migration and modernization involving clinical or non-clinical data
  • Master's degree or equivalent in a quantitative or engineering discipline
  • Databricks Data Engineer Professional or Associate/Professional certification
  • Experience leading enterprise legacy-to-lakehouse modernization in a regulated industry
  • Clinical data knowledge including CDISC, SDTM, ADaM, EDC systems, and clinical trial data structures
  • Experience with GxP or equivalent compliance requirements in pharma, life sciences, finance, or another regulated industry
  • Experience supporting IND, NDA, or BLA submissions

Eli Lilly and Company Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Eli Lilly and Company and has not been reviewed or approved by Eli Lilly and Company.

  • Retirement Support Feedback suggests long-term savings are bolstered by a defined-benefit pension alongside a company 401(k) match and retiree health options. These elements make total compensation feel strong beyond base salary.
  • Leave & Time Off Breadth Feedback suggests paid time off is expansive, with substantial vacation, company shutdown days, and milestone time. This breadth of leave is viewed as a meaningful part of overall rewards.
  • Parental & Family Support Feedback suggests family-building and caregiving support are robust, including paid parental leave, adoption or surrogacy assistance, and backup care. These programs enhance the perceived value of benefits across life stages.

Eli Lilly and Company Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Indianapolis, IN
39,451 Employees
Year Founded: 1876

What We Do

Eli Lilly and Company engages in the discovery, development, manufacture, and sale of products in pharmaceutical products business segment. For more than a century, we have stayed true to a core set of values – excellence, integrity, and respect for people – that guide us in all we do: discovering medicines that meet real needs, improving the understanding and management of disease, and giving back to communities through philanthropy and volunteerism.

Similar Jobs

GitLab Logo GitLab

Senior Solutions Architect

Cloud • Security • Software • Cybersecurity • Automation
Easy Apply
In-Office or Remote
Bangalore, Bengaluru Urban, Karnataka, IND
2500 Employees

CSC Logo CSC

Accountant

Fintech • Legal Tech • Software • Financial Services • Cybersecurity • Data Privacy
Hybrid
Bangalore, Bengaluru Urban, Karnataka, IND
8500 Employees

Optum Logo Optum

Senior Data Engineer

Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
In-Office
Bengaluru, Bengaluru Urban, Karnataka, IND
160000 Employees

Optum Logo Optum

Designer

Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
In-Office
Bengaluru, Bengaluru Urban, Karnataka, IND
160000 Employees

Similar Companies Hiring

Sailor Health Thumbnail
Healthtech • Social Impact • Telehealth
New York City, NY
20 Employees
Granted Thumbnail
Artificial Intelligence • Healthtech • Insurance • Mobile • Financial Services
New York, New York
23 Employees
OneImaging Thumbnail
Healthtech
Miami, FL
62 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account