The Role
Lead and build scalable data architecture for an AI-driven product platform across customer cloud environments. Design data platforms, establish engineering standards, review code, mentor engineers, and deliver hands-on solutions for complex data and AI workloads. The role requires expertise in Databricks, Spark, Python, SQL, dimensional modeling, cloud platforms, LLM/RAG pipelines, governance, security, observability, infrastructure automation, and CI/CD.
Summary Generated by Built In
We’re looking for a hands-on Lead Data Engineer to own and build the data architecture for an AI-driven product platform deployed across customer cloud environments.
This is an architect-who-still-ships role — you’ll design scalable data platforms, set engineering standards, review code, mentor engineers, and work hands-on with complex data and AI workloads.
What We’re Looking For
• 10+ years of experience in production data engineering
• 3+ years in Lead/Staff/Principal-level ownership
• Expert in Databricks, Delta Lake, Unity Catalog & Medallion Architecture
• Strong PySpark, Spark SQL, Python & SQL
• Expertise in Kimball dimensional modelling & dbt
• Working experience with Snowflake & Data Vault 2.0
• Experience across Azure & AWS
• Hands-on experience with LLM/RAG pipelines, embeddings & vector search
• Strong understanding of data governance, security, observability & cost optimization
• Experience with Kubernetes, Terraform/Bicep & CI/CD
Experience in regulated/compliance-heavy environments is a plus
Skills Required
- 10+ years of experience in production data engineering
- 3+ years of Lead, Staff, or Principal-level ownership
- Expertise in Databricks, Delta Lake, Unity Catalog, and Medallion Architecture
- Strong PySpark, Spark SQL, Python, and SQL skills
- Expertise in Kimball dimensional modeling and dbt
- Working experience with Snowflake and Data Vault 2.0
- Experience across Azure and AWS
- Hands-on experience with LLM/RAG pipelines, embeddings, and vector search
- Strong understanding of data governance, security, observability, and cost optimization
- Experience with Kubernetes, Terraform/Bicep, and CI/CD
- Experience in regulated or compliance-heavy environments
Am I A Good Fit?
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.
Success! Refresh the page to see how your skills align with this role.
The Company
What We Do
Statsby Solutions is a pharma-native AI and data company specializing in the pharmaceutical and clinical research industry. They provide end-to-end data and AI capabilities, including AI-powered platforms like Revectra for protocol intelligence and Veractra for Clinical Study Report generation, as well as consulting services in data engineering, Generative AI, and MLOps designed for high-stakes, regulated environments.








