The Role
Designs and maintains scalable Databricks and Spark data pipelines, including ETL/ELT workflows, Delta Lake transformations, data quality checks, and performance optimization. Works with cloud platforms, data warehousing, modeling, CI/CD, monitoring, and production support. Troubleshoots large-scale data processing issues and collaborates with engineers, data scientists, analysts, and business teams. Preferred experience includes Unity Catalog, orchestration tools, streaming technologies, infrastructure automation, and Databricks certifications.
Summary Generated by Built In
We are looking for a Senior Databricks Engineer with 7+ years of experience in data engineering and strong hands-on expertise in Databricks and Apache Spark. The ideal candidate will design and build scalable data pipelines, optimize large-scale data processing, and work with cloud-based data platforms.
- Design, develop, and maintain scalable data pipelines using Databricks and Apache Spark.
- Develop and optimize PySpark/Spark jobs for large-scale data processing.
- Build reliable ETL/ELT pipelines using Databricks workflows, Delta Lake, and SQL.
- Implement data transformations, data quality checks, and performance optimization.
- Work with cloud data platforms such as AWS, Azure, or GCP.
- Optimize Spark jobs, cluster configurations, partitioning, joins, and data processing performance.
- Collaborate with data engineers, data scientists, analysts, and business teams to deliver data solutions.
- Implement best practices for CI/CD, version control, monitoring, and production support.
- Troubleshoot data pipeline and performance issues in production environments.
- 7+ years of experience in Data Engineering.
- Strong hands-on experience with Databricks.
- Strong expertise in Apache Spark / PySpark.
- Experience with Delta Lake, Databricks Workflows, and Spark SQL.
- Strong proficiency in Python and SQL.
- Experience building and optimizing large-scale ETL/ELT pipelines.
- Experience with at least one cloud platform: AWS, Azure, or GCP.
- Good understanding of data warehousing, data modeling, and distributed data processing.
- Experience with Git and CI/CD practices.
- Experience with Unity Catalog and Databricks governance.
- Experience with Azure Data Factory, AWS Glue, or similar orchestration tools.
- Knowledge of streaming technologies such as Kafka or Spark Structured Streaming.
- Experience with Infrastructure as Code or cloud automation.
- Databricks certifications are a plus.
Skills Required
- 7+ years of experience in data engineering
- Strong hands-on experience with Databricks
- Strong expertise with Apache Spark and PySpark
- Experience with Delta Lake, Databricks Workflows, and Spark SQL
- Strong proficiency in Python and SQL
- Experience building and optimizing large-scale ETL/ELT pipelines
- Experience with at least one cloud platform: AWS, Azure, or GCP
- Understanding of data warehousing, data modeling, and distributed data processing
- Experience with Git and CI/CD practices
- Experience with Unity Catalog and Databricks governance
- Experience with Azure Data Factory, AWS Glue, or similar orchestration tools
- Knowledge of Kafka or Spark Structured Streaming
- Experience with Infrastructure as Code or cloud automation
- Databricks certifications
Am I A Good Fit?
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.
Success! Refresh the page to see how your skills align with this role.
The Company
What We Do
Numentica is a consulting firm specializing in software product development, business intelligence, data management, data analytics, AI, and cloud services, focusing on digital transformation.








