Data Pipeline & Ingestion Engineer

Posted 2 Hours Ago
Hiring Remotely in United States
Remote
Senior level
Artificial Intelligence • Cloud • Information Technology • Consulting
The Role
Design, build, and operate large-scale batch and streaming ingestion pipelines, medallion/lakehouse layering, identity resolution and golden-record consolidation, data-quality gates and reconciliation, source-to-canonical mappings and crosswalks, schema enforcement, and observability for client onboarding and production data flows.
Summary Generated by Built In

About the Program

The Operational Data Layer (ODL) program is a large-scale data-platform build for a leading benefitsadministration

platform. The platform ingests data from multiple legacy benefits systems, masters it into

golden records, transforms it into a canonical data model, and serves it through modern APIs — built on an

AWS / Java / Kafka stack in a HIPAA/SOX-regulated benefits domain spanning health, wealth/401(k), spending

accounts, and leaves.

About the Role

You will build and operate the data backbone of ODL: bulk and streaming ingestion from legacy source

systems, medallion-layered storage (Bronze/Silver/Gold), identity resolution and golden-record consolidation,

source-to-canonical mapping and crosswalks, and the data-quality and reconciliation gates that prove data is

complete and correct before it is published. This is the volume engine of the program — every new client

onboarded flows through the pipelines you build.

What You’ll Do

• Build batch-seed and event-tail ingestion per source system, including seed→tail watermark hand-off,

idempotent upserts, and dedup ledgers

• Build and operate medallion layers with reprocess-from-Bronze, pipeline orchestration (checkpoints,

retry/backoff, DLQ), and full observability

• Build data-quality gates (quarantine / pass-with-flag), quality scoring, and a reconciliation engine

covering count, record, and financial reconciliation — financial is zero-tolerance

• Build identity matching combining deterministic rules with probabilistic scoring and confidence bands;

deliver deduplication, golden-record materialization, and survivorship rules, calibrating match thresholds

with labelled data

• Author and maintain source→canonical structural mappings and value crosswalks (e.g., collapsing

1,800+ raw employment-status values to ~20 standard ones) as governed, versioned configuration

• Enforce data contracts at the boundary: schema registry, fail-fast validation, and semver-compatible

schema evolution

What We’re Looking For

• 5+ years building production data pipelines at scale

• Kafka depth: consumers/producers, replay, DLQ, exactly-once / idempotent processing patterns

• Strong SQL and solid ETL fundamentals

• Java and/or Python in production

• Medallion / lakehouse layering, CDC, watermark/checkpoint patterns, and batch–stream hand-off

• Data-quality frameworks: validation rules, quarantine and re-entry, quality scoring, reconciliation

• Entity resolution / MDM exposure: record matching, dedup, survivorship — via commercial tools

(Informatica MDM, Reltio) or custom builds

• Data mapping and crosswalk discipline: profiling messy datasets, authoring governed reference data,

config-as-code (YAML/JSON, Git)

Bonus Points

• Probabilistic record linkage at depth — blocking/candidate generation, scoring models, threshold

calibration (expected at senior level)

• Schema registry experience (Avro/Protobuf)

• Extracting from mainframe or older RDBMS sources with limited CDC support

• Financial reconciliation in finance-adjacent domains

• Benefits administration or healthcare domain knowledge

Skills Required

  • 5+ years building production data pipelines at scale
  • Kafka depth: consumers/producers, replay, DLQ, exactly-once/idempotent processing patterns
  • Strong SQL and solid ETL fundamentals
  • Java and/or Python in production
  • AWS (cloud platform)
  • Medallion / lakehouse layering, CDC, watermark/checkpoint patterns, batch-stream hand-off
  • Data-quality frameworks: validation rules, quarantine and re-entry, quality scoring, reconciliation (including financial reconciliation)
  • Entity resolution / MDM exposure: record matching, deduplication, survivorship (Informatica MDM, Reltio, or custom)
  • Data mapping and crosswalk discipline: profiling messy datasets, governed reference data, config-as-code (YAML/JSON), Git
  • Schema registry experience (Avro/Protobuf)
  • Probabilistic record linkage at depth: blocking/candidate generation, scoring, threshold calibration
  • Experience extracting from mainframe or older RDBMS sources with limited CDC support
  • Benefits administration or healthcare domain knowledge
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
148 Employees
Year Founded: 2018

What We Do

PeerIslands is an AI-native enterprise transformation partner that helps global organizations modernize by streamlining workflows and accelerating software delivery. Using its proprietary PeerAI platform, the company leverages generative AI, multi-agent orchestration, and automation. It specializes in cloud-native application modernization and data engineering, serving Fortune 100 clients in industries such as BFSI, healthcare, and telecommunications.

Similar Jobs

TransUnion Logo TransUnion

Vice President Media & Entertainment Sales

Big Data • Fintech • Information Technology • Business Intelligence • Financial Services • Cybersecurity • Big Data Analytics
Remote or Hybrid
14 Locations
13000 Employees
134K-296K Annually

People Inc. Logo People Inc.

Manager, Growth Marketing Platforms

AdTech • Consumer Web • Digital Media • eCommerce • Marketing Tech
Remote or Hybrid
New York, NY, USA
3500 Employees
60K-85K Annually

The Aerospace Corporation Logo The Aerospace Corporation

Sr Supplier Dev Engineer

Aerospace • Artificial Intelligence • Cloud • Machine Learning • Software • Cybersecurity • Defense
Remote or Hybrid
Neenah, WI, USA
4600 Employees

Samsara Logo Samsara

Customer Success Manager

Artificial Intelligence • Cloud • Computer Vision • Hardware • Internet of Things • Software
Easy Apply
Remote or Hybrid
United States
4000 Employees
66K-88K Annually

Similar Companies Hiring

Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account