At Lilly, the work is demanding because patients are waiting. We unite caring with discovery to help make life better for people around the world, knowing that every decision, every detail, and every day matters. Headquartered in Indianapolis, Indiana, our over 50,000 employees around the globe take on complex challenges to discover and deliver life-changing medicines, strengthen how health is understood and managed, and support the communities we serve. This is hard, urgent, selfless work—but it’s work worth doing. If you’re driven by purpose and ready to bring your best to work that truly matters for patients, we invite you to join us.
About Lilly
At Lilly, everything we do starts with patients. We unite caring with discovery to make life better for people around the world. Headquartered in Indianapolis, Indiana, our global team of over 50,000 employees work with urgency and purpose to discover and deliver life-changing medicines, strengthen how health is understood and managed, and support the communities we serve. We bring our best to this work because people depend on it. If you're driven by purpose and determined to make a meaningful difference for patients, we invite you to bring your skill and your commitment to Lilly.
About Tech@Lilly
At Lilly, technology is not a support function. It is how a global medicine company operates, innovates, and delivers. Lilly in Bengaluru builds the capabilities that make this possible — cloud platforms, AI systems, and automation at enterprise scale — all in service of a purpose that makes this technology work genuinely distinctive, from advancing drug discovery to enabling connected clinical trials to keeping a global medicine company running at the standard patients deserve.
About the Organization
The Clinical & Non-Clinical Data Organization at Eli Lilly and Company is responsible for the design, build, and operation of enterprise data platforms that power drug discovery, clinical development, and regulatory submissions. Data Hub is building a robust Data Strategy to make Lilly's Clinical and Non-Clinical data AI-ready and audit-ready, delivering scalable, governed, and reusable data products that accelerate how medicines reach patients. The data engineering organization sits at the intersection of science, technology, and patient impact — connecting Clinical and Non-Clinical data across the full chain, from ingestion to consumption.
Path/Level: R5 (Senior Data Architect)
Position Summary
The Senior Data Architect (R5) is a hands-on leader who designs and personally builds AI and agentic solutions that operate at enterprise scale across Lilly's Clinical and Non-Clinical data domain — multi-agent systems, LLM-orchestrated pipelines, and retrieval/reasoning architectures built on governed, semantic data foundations. This is a builder role: the architect writes code, stands up agent frameworks, and ships production AI systems personally, not just specifications. This role is split 70% hands-on technical execution including coding and 30% strategy, and shapes the future technology landscape.
A defining mandate of this role is Right Model, Right Task — routing every agentic and LLM workload to the model best suited to it on cost, latency, and accuracy grounds, integrated directly with Lilly's internal data platform so routing and reasoning are grounded in enterprise-native lineage and knowledge graphs rather than bespoke, disconnected metadata. Agentic and AI-assisted ways of working are the expected default across every design, analysis, and documentation activity — and this leader is the reference point for how the broader India team scales AI-native architecture.
Key Responsibilities
AI & Agentic Solution Architecture (Hands-On)
Architect and personally build multi-agent and LLM-orchestrated solutions — planning/tool-calling agents, retrieval-augmented generation (RAG), and agent-to-agent workflows — for Clinical and Non-Clinical use cases.
Design for scale from the start: agent orchestration, state management, concurrency, cost/latency budgets, and failure/retry handling across high-volume production workloads.
Implement evaluation, guardrails, observability, and human-in-the-loop patterns so agentic systems are safe, auditable, and production-ready in a regulated environment.
Right Model, Right Task – Platform-Integrated Routing, Lineage & Knowledge Graphs
Design and implement a right-model-right-task routing layer that selects the optimal model — by size, provider, and fine-tuned vs. general-purpose — for each agentic task based on complexity, cost, latency, and accuracy requirements, rather than defaulting to one model for every job.
Integrate directly with Lilly's internal data platform (Data Hub catalog, lineage, and metadata services) rather than building parallel metadata stores, so routing and agent reasoning are grounded in the enterprise's single source of truth.
Build and maintain lineage-aware enterprise knowledge graphs sourced from the internal data platform, capturing data provenance, sensitivity, and domain context that ground both agentic reasoning and model-selection decisions.
Use lineage and knowledge-graph context to enforce routing policy — for example, regulatory-sensitive Clinical data routes only to approved/validated models, while low-sensitivity exploratory tasks use lighter-weight, lower-cost models.
Continuously benchmark and re-tune the model portfolio and routing rules as new models become available, avoiding both over-provisioning expensive frontier models and under-serving tasks that genuinely need them.
AI-Native Data Engineering & Semantic Modeling
Use AI-assisted and agentic tooling by default across the data lifecycle — pipeline generation, schema/ontology alignment, data-quality scoring, and documentation — not as an occasional accelerator.
Design conceptual, logical, physical, and semantic models — ontologies, taxonomies, knowledge graphs — that ground both self-service analytics and agentic reasoning, built on and reconciled with the platform's native lineage.
Build analytics-ready dimensional models and semantic/vector layers that power self-service BI and AI-driven consumption for scientists, statisticians, and business analysts.
Automation & Team Reusability at Scale
Build reusable agentic accelerators — agent templates, orchestration patterns, evaluation harnesses, model-routing configs — that the broader India team adopts directly to scale AI-native delivery.
Automate lineage capture, schema-drift detection, and data/agent quality checks so the team spends more time on judgment calls, less on manual upkeep.
Stand up and maintain the AI/agentic standards platform — the enterprise reference for how agentic solutions, model routing, and knowledge graphs are designed, governed, and delivered across Clinical and Non-Clinical squads.
Governance & Responsible AI
Implement role-based and attribute-based access control, encryption, and prompt/data-leakage safeguards for agentic systems operating on regulated data.
Embed data quality, lineage, model/agent risk, and compliance requirements directly into every agentic design and routing decision.
Collaboration & Stakeholder Engagement
Partner with business SMEs, solution architects, platform/data-catalog teams, and engineering teams to translate Clinical and Non-Clinical needs into scaled agentic solutions.
Communicate technical trade-offs and AI/agentic architecture recommendations clearly to technical and non-technical stakeholders, and shape the India AI/Agentic roadmap.
Technical Skills & Qualifications
Required — AI & Agentic Solutions at Scale (Must-Have, Hands-On)
5+ years hands-on building and shipping AI/agentic systems in production: multi-agent orchestration, RAG, tool-calling/function-calling agents, and LLM application frameworks (e.g., LangGraph, LlamaIndex, Semantic Kernel, or equivalent).
Hands-on experience designing model-routing/orchestration layers (right-model-right-task) across a multi-model portfolio, balancing cost, latency, and accuracy.
Hands-on experience with vector databases/semantic search, embeddings, and retrieval architectures grounding LLMs in governed data.
Experience building evaluation frameworks, guardrails, and observability/monitoring for LLM and agentic applications.
Required — Data Platform Integration, Lineage & Semantic Modeling
Hands-on experience with cloud platforms (AWS) and Unity Catalog-based access control, encryption, and lineage/compliance patterns for regulated data.
Proven, hands-on semantic data modeling experience: ontologies, taxonomies, knowledge graphs (RDF/OWL or equivalent), built for real Clinical or Non-Clinical data sets.
Experience integrating agentic/AI systems with enterprise data-catalog and lineage platforms to build knowledge graphs that ground agent reasoning and model-routing decisions, rather than standing up disconnected metadata stores.
Hands-on experience with cloud-based lakehouse/data platforms (Databricks, or equivalent) as the governed data foundation feeding agentic and AI systems.
Required — AI-Assisted Ways of Working & Governance
Demonstrated, current practice of using AI-assisted and agentic tools as the default method for analysis, design, and delivery — not occasional use.
Experience implementing access control, encryption, and responsible-AI/compliance patterns for agentic systems on regulated data.
Preferred
Certifications or equivalent depth in LLMOps/MLOps platforms, or Databricks Mosaic AI / cloud AI-agent services.
Exposure to data mesh and data product architecture; API-first design for agent-consumable services.
Familiarity with GxP / 21 CFR Part 11 compliance and responsible-AI governance in a validated data environment.
Education & Experience
Minimum
Bachelor's degree in Computer Science, Information Systems, or related discipline.
12+ years of hands-on data architecture experience, including 5+ years building and shipping AI/agentic solutions to production, with direct experience on Clinical and/or Non-Clinical data.
Demonstrated track record of AI/agentic architecture work that shipped to production and operates at scale, including integration with an enterprise data platform for lineage and knowledge graphs.
Preferred
Master's degree or equivalent in a quantitative or engineering discipline.
Experience in a regulated industry (pharma, life sciences, finance) with GxP or equivalent compliance requirements.
Prior contribution to clinical or non-clinical data programs supporting IND, NDA, or BLA submissions.
At Lilly, caring is not only what we do for patients. It is how we work. We believe the people who dedicate themselves to making medicines better deserve an environment that makes their lives better too, one where they are supported, respected, and given the space to do their best work. This is not just a policy. It is who we are.
Equal Opportunity & Accommodation
Lilly is dedicated to helping individuals with disabilities to actively engage in the workforce, ensuring equal opportunities when vying for positions. If you require accommodation to submit a resume for a position at Lilly, please complete the accommodation request form at https://careers.lilly.com/us/en/workplace-accommodation for further assistance. Please note this is for individuals to request an accommodation as part of the application process and any other correspondence will not receive a response.
Lilly is an EEO/Affirmative Action Employer and does not discriminate on the basis of age, race, colour, religion, gender, sexual orientation, gender identity, gender expression, national origin, protected veteran status, disability or any other legally protected status.
Lilly is dedicated to helping individuals with disabilities to actively engage in the workforce, ensuring equal opportunities when vying for positions. If you require accommodation to submit a resume for a position at Lilly, please complete the accommodation request form (https://careers.lilly.com/us/en/workplace-accommodation) for further assistance. Please note this is for individuals to request an accommodation as part of the application process and any other correspondence will not receive a response.
Lilly does not discriminate on the basis of age, race, color, religion, gender, sexual orientation, gender identity, gender expression, national origin, protected veteran status, disability or any other legally protected status.
#WeAreLillySkills Required
- Bachelor's degree in Computer Science, Information Systems, or a related discipline
- 12+ years of hands-on data architecture experience
- 5+ years hands-on building and shipping AI or agentic systems in production
- Experience with multi-agent orchestration, RAG, tool-calling or function-calling agents, and LLM application frameworks
- Hands-on experience designing model-routing or orchestration layers across multi-model portfolios
- Hands-on experience with vector databases, semantic search, embeddings, and retrieval architectures
- Experience building evaluation frameworks, guardrails, and observability for LLM and agentic applications
- Hands-on experience with AWS and Unity Catalog-based access control, encryption, lineage, and compliance patterns
- Proven hands-on semantic data modeling experience with ontologies, taxonomies, and knowledge graphs
- Experience applying semantic modeling to real clinical or non-clinical datasets
- Experience integrating AI systems with enterprise data-catalog and lineage platforms
- Hands-on experience with Databricks or equivalent cloud-based lakehouse platforms
- Current demonstrated practice using AI-assisted and agentic tools as the default method for analysis, design, and delivery
- Experience implementing access control, encryption, responsible-AI, and compliance patterns for regulated data
- Direct experience with clinical and/or non-clinical data
- Certifications or equivalent depth in LLMOps, MLOps, Databricks Mosaic AI, or cloud AI-agent services
- Exposure to data mesh and data product architecture
- Experience with API-first design for agent-consumable services
- Familiarity with GxP, 21 CFR Part 11, and responsible-AI governance
- Experience in a regulated industry such as pharma, life sciences, or finance
- Experience supporting IND, NDA, or BLA submission programs
- Master's degree or equivalent in a quantitative or engineering discipline
Eli Lilly and Company Compensation & Benefits Highlights
The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Eli Lilly and Company and has not been reviewed or approved by Eli Lilly and Company.
-
Retirement Support — Feedback suggests long-term savings are bolstered by a defined-benefit pension alongside a company 401(k) match and retiree health options. These elements make total compensation feel strong beyond base salary.
-
Leave & Time Off Breadth — Feedback suggests paid time off is expansive, with substantial vacation, company shutdown days, and milestone time. This breadth of leave is viewed as a meaningful part of overall rewards.
-
Parental & Family Support — Feedback suggests family-building and caregiving support are robust, including paid parental leave, adoption or surrogacy assistance, and backup care. These programs enhance the perceived value of benefits across life stages.
Eli Lilly and Company Insights
What We Do
Eli Lilly and Company engages in the discovery, development, manufacture, and sale of products in pharmaceutical products business segment. For more than a century, we have stayed true to a core set of values – excellence, integrity, and respect for people – that guide us in all we do: discovering medicines that meet real needs, improving the understanding and management of disease, and giving back to communities through philanthropy and volunteerism.
.png)





