The Role
Design, build, and maintain scalable batch and real-time ingestion pipelines using Medallion/Lakehouse patterns. Implement CDC, checkpointing, watermarking, data quality, reconciliation, identity resolution/MDM, schema validation/versioning, and source-to-canonical mappings. Collaborate cross-functionally to onboard sources and monitor pipeline health.
Summary Generated by Built In
Location: India
Employment Type: Full-Time
Experience: Senior (7+ Years)
Roles & Responsibilities
- Design, build, and maintain scalable batch and real-time data ingestion pipelines.
- Develop and manage Bronze, Silver, and Gold (Medallion) data layers.
- Implement CDC, watermarking, checkpointing, and batch-to-stream data processing.
- Build robust data quality, validation, reconciliation, and monitoring frameworks.
- Develop identity resolution, deduplication, and Golden Record (MDM) solutions.
- Create and maintain source-to-canonical data mappings and crosswalks.
- Ensure schema validation, versioning, and data contract enforcement.
- Collaborate with cross-functional teams to onboard new data sources and optimize data pipelines.
Mandatory Skills
- 7+ years of experience in Data Engineering / Data Pipeline development.
- Strong experience with Apache Kafka (Producers, Consumers, Replay, DLQ, Exactly-once/Idempotent processing).
- Strong SQL and ETL/ELT fundamentals.
- Hands-on experience in Java and/or Python.
- Experience with CDC, Batch & Streaming pipelines, and Medallion/Lakehouse architecture.
- Experience implementing Data Quality, Validation, and Reconciliation frameworks.
- Knowledge of Master Data Management (MDM), Identity Resolution, Deduplication, and Golden Record concepts.
- Experience with source-to-target mapping, canonical data models, YAML/JSON configuration, and Git.
Good to Have
- Experience with probabilistic record matching and record linkage.
- Schema Registry (Avro/Protobuf).
- Experience extracting data from legacy/Mainframe systems.
- Financial reconciliation experience.
- Healthcare or Benefits Administration domain experience.
Skills Required
- 7+ years of experience in Data Engineering / Data Pipeline development.
- Strong experience with Apache Kafka (Producers, Consumers, Replay, DLQ, Exactly-once/Idempotent processing).
- Strong SQL and ETL/ELT fundamentals.
- Hands-on experience in Java and/or Python.
- Experience with CDC, Batch & Streaming pipelines, and Medallion/Lakehouse architecture.
- Experience implementing Data Quality, Validation, and Reconciliation frameworks.
- Knowledge of Master Data Management (MDM), Identity Resolution, Deduplication, and Golden Record concepts.
- Experience with source-to-target mapping, canonical data models, YAML/JSON configuration, and Git.
- Experience with probabilistic record matching and record linkage.
- Schema Registry (Avro/Protobuf).
- Experience extracting data from legacy/Mainframe systems.
- Financial reconciliation experience.
- Healthcare or Benefits Administration domain experience.
Am I A Good Fit?
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.
Success! Refresh the page to see how your skills align with this role.
The Company
What We Do
PeerIslands is an AI-native enterprise transformation partner that helps global organizations modernize by streamlining workflows and accelerating software delivery. Using its proprietary PeerAI platform, the company leverages generative AI, multi-agent orchestration, and automation. It specializes in cloud-native application modernization and data engineering, serving Fortune 100 clients in industries such as BFSI, healthcare, and telecommunications.








