Cloudera Data Engineer

Posted 3 Days Ago
Be an Early Applicant
Bengaluru, Bengaluru Urban, Karnataka, IND
In-Office
Expert/Leader
Artificial Intelligence • Cloud • Information Technology • Cybersecurity
The Role
Design, build, and optimize end-to-end Databricks ETL/ELT pipelines using Delta Lake, DLT, Auto Loader, PySpark and Spark SQL. Implement medallion lakehouse architecture, governance with Unity Catalog, data quality, CDC, performance optimization, orchestration, CI/CD, monitoring, and deliver production-grade, tested increments for high-volume regulated data workloads.
Summary Generated by Built In

Key Responsibilities

  • Design, build, and optimize end-to-end ETL/ELT pipelines in Databricks using Delta Lake, Delta Live Tables (DLT), Auto Loader, PySpark, and Spark SQL for high-volume, multi-format partner ingestion.
  • Implement Medallion (zoned) architecture – Raw (bronze), Standardized (silver) with advanced validation, quarantine/reject logic, schema enforcement, and Curated (gold) consumer-ready datasets optimized for downstream COB/PI analytics.
  • Leverage Unity Catalog for data governance, access control, lineage, and secure multi-tenant data management.
  • Develop incremental processing, change data capture (CDC), backfill strategies, late-arriving data handling, and partitioning/optimization techniques (Z-Ordering, Liquid Clustering, Auto-Optimize) to eliminate performance bottlenecks.
  • Build robust data quality frameworks using Delta constraints, expectations, and monitoring to ensure clean, reliable data for downstream consumption.
  • Create production-grade Databricks Workflows, Jobs, and orchestration for reliable batch and near-real-time processing using Spark Structured Streaming.
  • Perform data profiling, mapping, reconciliation, and performance tuning of large-scale Spark jobs on Databricks clusters.
  • Collaborate with Senior Data Architect and Data Modeller to translate target-state lakehouse design into implementable, testable increments.
  • Deliver shippable, production-ready increments in Agile sprints within the implementation window, including CI/CD integration, unit/integration testing, and operational runbooks.
  • Establish comprehensive observability using Databricks Lakehouse Monitoring, SQL Alerts, and dashboards for pipeline health and SLA compliance.


Requirements

Required Qualifications & Experience

  • 8+ years of hands-on data engineering experience
  • 5+ years building enterprise-scale solutions on Databricks (Unity Catalog, Delta Lake, Delta Live Tables)
  • Proven track record delivering Medallion/zonal lakehouse architectures in production
  • Strong experience with high-volume, regulated data workloads (claims, financial, or healthcare data highly preferred)

Technical Skills – Databricks Expertise (Core)

  • Databricks Platform: Unity Catalog, Delta Lake, Delta Live Tables (DLT), Auto Loader, Workflows, Jobs, Repos, Lakehouse Monitoring
  • Core Technologies: PySpark, Spark SQL, Spark Structured Streaming, Delta constraints & expectations
  • Optimization & Performance: Liquid Clustering, Z-Ordering, Auto-Optimize, Dynamic Partition Overwrite, Photon engine
  • Governance & Quality: Unity Catalog ACLs, data lineage, schema evolution, Great Expectations (or equivalent)
  • Orchestration & CI/CD: Databricks Workflows, dbt on Databricks, Git integration, Azure DevOps / Jenkins
  • Languages: Expert Python (PySpark), SQL
  • Cloud: AWS/Azure/GCP (Databricks on any cloud)

Preferred Qualifications

  • Prior experience modernizing healthcare claims data lakes (COB, Payment Integrity, Medicaid/Medicare)
  • Exposure to partner ingestion patterns, multi-format data (EDI, flat files, APIs), and downstream analytical workloads
  • Familiarity with CMS/HIPAA data handling and compliance in Databricks environments


Skills Required

  • 8+ years hands-on data engineering experience
  • 5+ years building enterprise-scale solutions on Databricks (Unity Catalog, Delta Lake, Delta Live Tables)
  • Proven track record delivering Medallion/zonal lakehouse architectures in production
  • Experience with high-volume, regulated data workloads (claims, financial, or healthcare)
  • Expert Python (PySpark) and SQL
  • Databricks platform expertise: Unity Catalog, Delta Lake, Delta Live Tables, Auto Loader, Workflows, Jobs, Repos, Lakehouse Monitoring
  • Experience with Spark Structured Streaming, Delta constraints & expectations
  • Optimization experience: Liquid Clustering, Z-Ordering, Auto-Optimize, Dynamic Partition Overwrite, Photon engine
  • Governance & quality: Unity Catalog ACLs, data lineage, schema evolution, Great Expectations (or equivalent)
  • Orchestration & CI/CD: Databricks Workflows, dbt on Databricks, Git integration, Azure DevOps or Jenkins
  • Experience deploying Databricks on cloud (AWS, Azure, or GCP)
  • Prior experience modernizing healthcare claims data lakes (COB, Payment Integrity, Medicaid/Medicare)
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
2,500 Employees
Year Founded: 1932

What We Do

McLaren Strategic Solutions specializes in AI, Cloud, Cybersecurity, and Data Engineering to drive digital transformation. The company provides cutting-edge automation solutions to enhance efficiency and innovation, utilizing platform-based service delivery to help clients accelerate their digital future. By leveraging expertise in intelligent automation and digital engineering, they empower organizations to optimize their operations and achieve sustainable growth.

Similar Jobs

Capital One Logo Capital One

Tax Manager

Fintech • Machine Learning • Payments • Software • Financial Services
Hybrid
Bengaluru, Bengaluru Urban, Karnataka, IND
55000 Employees

Capital One Logo Capital One

Senior Associate, Accounting

Fintech • Machine Learning • Payments • Software • Financial Services
Hybrid
Bengaluru, Bengaluru Urban, Karnataka, IND
55000 Employees

Optum Logo Optum

Senior Software Engineering Lead

Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
In-Office
Bengaluru, Bengaluru Urban, Karnataka, IND
160000 Employees

Optum Logo Optum

Senior Quality Engineer II

Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
In-Office
Bengaluru, Bengaluru Urban, Karnataka, IND
160000 Employees

Similar Companies Hiring

Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account