The Role
Lead design, develop, and maintain scalable Databricks-based data pipelines using PySpark/Python and Delta Lake. Build and optimize ETL/ELT workflows, ensure data quality, troubleshoot pipelines, and collaborate with analytics and engineering teams.
Summary Generated by Built In
We are looking for an experienced Data Engineer (Lead) with strong expertise in Databricks to design, develop, and maintain scalable data pipelines and data processing solutions.
- Design, develop, and maintain scalable data pipelines using Databricks.
- Develop data processing solutions using PySpark and Python.
- Work with Apache Spark for large-scale data processing.
- Build and optimize ETL/ELT pipelines.
- Work with Delta Lake for data storage and processing.
- Develop and manage workflows using Databricks Workflows.
- Perform data transformation, cleansing, and validation.
- Optimize Spark jobs and improve pipeline performance.
- Collaborate with data engineers, analysts, and other technical teams.
- Troubleshoot data pipeline issues and ensure data quality.
- 3–5 years of experience in Data Engineering.
- Strong hands-on experience with Databricks.
- Good knowledge of PySpark and Python.
- Strong understanding of Apache Spark.
- Experience with Delta Lake and data lake architecture.
- Good knowledge of SQL.
- Experience in building ETL/ELT data pipelines.
- Understanding of cloud platforms such as Azure, AWS, or GCP.
- Experience with Git and CI/CD is an added advantage.
- Experience with Azure Databricks.
- Knowledge of Azure Data Factory, ADLS, or similar cloud data services.
- Experience with DBT, Snowflake, or other modern data platforms.
- Databricks certification is a plus.
- Opportunity to work on real-world Data Engineering and AI projects.
- Exposure to modern data technologies and cloud platforms.
- Collaborative and learning-focused work environment.
- Opportunity to work with a growing technology team.
Skills Required
- 3-5 years of experience in Data Engineering
- Hands-on experience with Databricks
- Proficiency with PySpark and Python
- Strong understanding of Apache Spark
- Experience with Delta Lake and data lake architecture
- Good knowledge of SQL
- Experience building ETL/ELT data pipelines
- Understanding of cloud platforms (Azure, AWS, or GCP)
- Experience with Git and CI/CD
- Experience with Azure Databricks
- Knowledge of Azure Data Factory, ADLS, or similar cloud data services
- Experience with DBT, Snowflake, or other modern data platforms
- Databricks certification
Am I A Good Fit?
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.
Success! Refresh the page to see how your skills align with this role.
The Company
What We Do
Statsby Solutions is a pharma-native AI and data company specializing in the pharmaceutical and clinical research industry. They provide end-to-end data and AI capabilities, including AI-powered platforms like Revectra for protocol intelligence and Veractra for Clinical Study Report generation, as well as consulting services in data engineering, Generative AI, and MLOps designed for high-stakes, regulated environments.








