The Role
Design, build, and maintain scalable Databricks data pipelines using PySpark/Python and Apache Spark. Implement ETL/ELT workflows with Delta Lake, optimize Spark jobs, ensure data quality, and collaborate with engineering and analytics teams to troubleshoot and improve pipeline performance.
Summary Generated by Built In
About the Role
Key Responsibilities
We are looking for an experienced Data Engineer with strong expertise in Databricks to design, develop, and maintain scalable data pipelines and data processing solutions.
- Design, develop, and maintain scalable data pipelines using Databricks.
- Develop data processing solutions using PySpark and Python.
- Work with Apache Spark for large-scale data processing.
- Build and optimize ETL/ELT pipelines.
- Work with Delta Lake for data storage and processing.
- Develop and manage workflows using Databricks Workflows.
- Perform data transformation, cleansing, and validation.
- Optimize Spark jobs and improve pipeline performance.
- Collaborate with data engineers, analysts, and other technical teams.
- Troubleshoot data pipeline issues and ensure data quality.
- 3–5 years of experience in Data Engineering.
- Strong hands-on experience with Databricks.
- Good knowledge of PySpark and Python.
- Strong understanding of Apache Spark.
- Experience with Delta Lake and data lake architecture.
- Good knowledge of SQL.
- Experience in building ETL/ELT data pipelines.
- Understanding of cloud platforms such as Azure, AWS, or GCP.
- Experience with Git and CI/CD is an added advantage.
- Experience with Azure Databricks.
- Knowledge of Azure Data Factory, ADLS, or similar cloud data services.
- Experience with DBT, Snowflake, or other modern data platforms.
- Databricks certification is a plus.
- Opportunity to work on real-world Data Engineering and AI projects.
- Exposure to modern data technologies and cloud platforms.
- Collaborative and learning-focused work environment.
- Opportunity to work with a growing technology team.
Skills Required
- 3-5 years of experience in Data Engineering
- Strong hands-on experience with Databricks
- Good knowledge of PySpark and Python
- Strong understanding of Apache Spark
- Experience with Delta Lake and data lake architecture
- Good knowledge of SQL
- Experience in building ETL/ELT data pipelines
- Understanding of cloud platforms such as Azure, AWS, or GCP
- Experience with Git and CI/CD
Am I A Good Fit?
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.
Success! Refresh the page to see how your skills align with this role.
The Company
What We Do
Statsby Solutions is a pharma-native AI and data company specializing in the pharmaceutical and clinical research industry. They provide end-to-end data and AI capabilities, including AI-powered platforms like Revectra for protocol intelligence and Veractra for Clinical Study Report generation, as well as consulting services in data engineering, Generative AI, and MLOps designed for high-stakes, regulated environments.









