The Role
Design and maintain scalable batch and streaming data pipelines using Azure Databricks, Spark, Delta Lake, Azure Data Factory, and related Azure services. Build governed Lakehouse architectures across Bronze, Silver, and Gold layers; optimize storage, partitioning, queries, and clusters; implement infrastructure as code; and support CI/CD deployments. Develop reusable pipeline components and integrate streaming, governance, security, and access-control solutions across enterprise data platforms.
Summary Generated by Built In
- Design and implement scalable batch and streaming data pipelines using Azure Databricks, Delta Live Tables, and Apache Spark.
- Build and maintain ETL/ELT workflows orchestrated through Azure Data Factory and Databricks Workflows.
- Develop reusable and modular pipeline components following software engineering best practices.
- Architect and manage Lakehouse solutions using Delta Lake across Bronze, Silver, and Gold layers.
- Design and enforce data models, schemas, and governance policies using Unity Catalog.
- Optimize storage, partitioning, and query performance for large-scale datasets on ADLS Gen2.
- Manage Databricks clusters, compute policies, and job scheduling.
- Implement Infrastructure as Code (IaC) using Terraform or ARM templates.
- Integrate Databricks with Azure services including Synapse, Event Hubs, Key Vault, and Azure DevOps.
Requirements
- Strong proficiency in PySpark, Python, and SQL for Big Data processing.
- Hands-on experience with Delta Lake, Delta Live Tables (DLT), and Medallion Architecture.
- Strong experience with Azure Data Services including:
- Azure Data Lake Storage Gen2 (ADLS Gen2)
- Azure Data Factory (ADF)
- Azure Synapse Analytics
- Azure Event Hubs
- Experience with Databricks Unity Catalog for data governance and access control.
- Experience implementing CI/CD pipelines using Azure DevOps or GitHub Actions for Databricks deployments.
- Strong understanding of distributed computing concepts, Spark optimization, partitioning, and performance tuning.
- Experience with streaming data processing using Structured Streaming, Kafka, or Event Hubs.
Skills Required
- Strong proficiency in PySpark, Python, and SQL for big data processing
- Hands-on experience with Delta Lake, Delta Live Tables, and Medallion Architecture
- Strong experience with Azure Data Lake Storage Gen2, Azure Data Factory, Azure Synapse Analytics, and Azure Event Hubs
- Experience with Databricks Unity Catalog for data governance and access control
- Experience implementing CI/CD pipelines using Azure DevOps or GitHub Actions for Databricks deployments
- Strong understanding of distributed computing, Spark optimization, partitioning, and performance tuning
- Experience with streaming data processing using Structured Streaming, Kafka, or Event Hubs
Am I A Good Fit?
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.
Success! Refresh the page to see how your skills align with this role.
The Company
What We Do
InFynd is a cloud-based B2B data and prospecting platform that helps sales, marketing, and recruiting teams identify and engage potential customers. It provides contact information, company profiles, buyer-intent data, social media handles, and verified business email addresses and phone numbers. The platform supports target identification, company research, lead generation, and more effective sales conversion through accurate, GDPR-compliant business intelligence.








