Data Engineer

Posted 9 Hours Ago
Be an Early Applicant
Gurugram, Haryana, IND
Hybrid
Mid level
Information Technology • Database • Consulting
The Role
Build and operate Microsoft Fabric data pipelines for six heterogeneous sources. Responsibilities include ingestion, mirroring, CDC, incremental loads, Delta Lake raw-layer management, PySpark standardization, entity-resolution support, data quality checks, monitoring, troubleshooting, documentation, lineage, and Spark/Delta performance optimization. The role owns production ingestion pipelines, source-to-target mappings, exception handling, and operational run documentation.
Summary Generated by Built In

Build and operate the data pipelines that feed the Entity Hub. This role lands all six in-scope sources into Fabric, implements standardization and transformation logic, and maintains the data quality checks and monitoring that the entity resolution engine depends on. Reliable, observable ingestion is the foundation the entire programmed rests on.

Responsibilities
  • Ingestion development — build and maintain pipelines to land the six in-scope sources (Secretary of State, D&B, ARROW, E1, hCue, DocCentral) into the Fabric Bronze/raw layer.
  • Mirroring & CDC — implement Fabric Mirroring for supported structured sources and establish change-data-capture patterns; implement watermark/incremental load logic where mirroring is unavailable.
  • Raw layer management — maintain one Delta table per source on an append-only basis, retaining evidence records and full source provenance.
  • Standardization & transformation — implement name normalization, address parsing and attribute standardization logic in Spark notebooks; support identifier-spine construction.
  • Data quality — implement data quality checks, validation rules, threshold alerts and exception handling; support reconciliation against source.
  • Pipeline operations — schedule, monitor and troubleshoot pipeline runs; investigate failures and performance issues; maintain run documentation.
  • Performance tuning — optimise Spark jobs, Delta file sizes, partitioning and pipeline efficiency to manage Fabric capacity consumption.

Documentation — produce and maintain source-to-target mappings, transformation logic documentation and lineage records

Qualifications Skill Area Specific Requirements Core Engineering Python, PySpark, advanced SQL, Delta Lake, distributed data processing Microsoft Fabric Data Factory pipelines and Copy Activity, Lakehouse, OneLake, Spark notebooks, Environments, Mirroring, Shortcuts Data Integration Batch and incremental ingestion, CDC patterns, watermarking, reprocessing strategies, schema-on-read for varied formats Data Quality Validation rule implementation, completeness/accuracy checks, alerting, exception workflows, reconciliation Modelling Bronze/Silver/Gold medallion layering, cleansing and conformance, standardization of names, addresses, dates and codes Ops & Governance Pipeline monitoring, lineage and metadata capture, access controls, technical documentation

Must-Have Qualifications

  • 4+ years hands-on data engineering with strong PySpark and SQL
  • Production experience building ingestion pipelines from multiple heterogeneous sources
  • Working knowledge of Delta Lake and medallion/lakehouse architecture
  • Experience implementing incremental loads and CDC-style processing
  • Experience implementing data quality checks and troubleshooting pipeline failures

Nice-to-Have

  • Microsoft Fabric hands-on experience (Mirroring, Copy Jobs, Environments)
  • Exposure to entity/master data standardization (name and address parsing)
  • Familiarity with libraries such as Great Expectations for data quality
  • Experience optimising for Fabric capacity/CU consumption

Key Deliverables Owned

  • Operational ingestion pipelines for all agreed sources
  • Bronze/raw layer with one Delta table per source and CDC retained
  • Standardization and parsing transformation logic
  • Data quality checks, monitoring and exception handling
  • Source-to-target mapping and run documentation

Dual Role / Complementary Skills

Complementary with the Entity Resolution engineering workstream — both are PySpark-on-Fabric disciplines, so this role can cross-train on Splink tuning and candidate-pair generation to provide cover. Also supports the Sr. Data Engineer (Lead) on identifier-spine construction, and can assist the VectorDB Engineer with document/attribute preparation in Phase 2.

Skills Required

  • 4+ years of hands-on data engineering experience with strong PySpark and SQL
  • Production experience building ingestion pipelines from multiple heterogeneous sources
  • Working knowledge of Delta Lake and medallion/lakehouse architecture
  • Experience implementing incremental loads and CDC-style processing
  • Experience implementing data quality checks and troubleshooting pipeline failures
  • Hands-on Microsoft Fabric experience, including Mirroring, Copy Jobs, and Environments
  • Exposure to entity or master data standardization, including name and address parsing
  • Familiarity with data quality libraries such as Great Expectations
  • Experience optimizing Fabric capacity or CU consumption
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: New York, NY
30,246 Employees
Year Founded: 1999

What We Do

Choosing a digital partner is about more than capabilities — it’s about collaboration and character. Unrealistic overhauls and off-the-shelf products ignore what matters most — your unique needs, culture, goals, and your legacy data and technology environments. At EXL, our collaboration is built on ongoing listening and learning to adapt our methodologies. We’re your business evolution partner—tailoring solutions that make the most of data to make better business decisions and drive more intelligence into your increasingly digital operations. Whether your goals are scaling the use of AI and digital, redesign operating models, or driving better and faster decisions, we’re here to partner with you to help you gain—and maintain—competitive advantage with efficient, sustainable models at scale. Our expertise in transformation, data science, and change management helps make your business more efficient and effective, improve customer relationships and enhance revenue growth. Instead of focusing on multi-year, resource- and time-intensive platform designs or migrations, we look deeper at your entire value chain to integrate strategies with impact. We use our specialization in analytics, digital interventions, and operations management—alongside deep industry expertise — to deliver solutions that help you outperform the competition. At EXL, it’s all about outcomes—your outcomes—and delivering success on your terms. Share your goals with us and together, we’ll optimize how you leverage data to drive your business forward. For more information, visit www.exlservice.com.

Similar Jobs

CrowdStrike Logo CrowdStrike

Data Engineer

Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity
Remote or Hybrid
India
11000 Employees
In-Office or Remote
2 Locations
5000 Employees

EXL Logo EXL

Data Engineer

Information Technology • Database • Consulting
Hybrid
Gurugram, Haryana, IND
30246 Employees
In-Office
Gurugram, Haryana, IND

Similar Companies Hiring

Axle Health Thumbnail
Artificial Intelligence • Healthtech • Information Technology • Logistics
Santa Monica, CA
25 Employees
NODA AI Thumbnail
Artificial Intelligence • Information Technology • Software • Cybersecurity
Sydney, AU
54 Employees
Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account