We are looking for a Senior Technical Product Manager to drive the next generation of our high-throughput, enterprise data ingestion and data management architecture. In this role, you will lead the strategic transformation of our Data Ingress Engine using technologies like Apache Spark, Databricks, and Medallion Lakehouse Architecture
You will also own the product roadmap for our core Master Data Management (MDM) and Data Remediation Engines, defining how millions of disparate clinical records are cleansed, matched into a single "Golden Patient Record," and safely reprocessed when errors occur.
If you thrive at the intersection of distributed systems, high-volume streaming/batch data pipelines, and complex data models, this role is for you.
Responsibilities:
- Lead the product vision and roadmap to decouple core data ingestion modules from operational API boundaries, turning them into highly scalable, standalone micro-services.
- Productize event-driven and streaming ingestion pipelines (e.g., Kafka/Spark Structured Streaming) capable of processing billions of high-velocity records.
- Establish clear API contracts, event interfaces, and integration specs for external sources feeding raw payloads (HL7 v2, C-CDA, FHIR JSON, Claims) into our platform.
- Define requirements for transitioning mass analytics pipelines onto a Databricks / Medallion Architecture using Apache Spark.
- Drive the design of optimized Silver and Gold layer datasets, ensuring deeply nested JSON formats (e.g., FHIR) are efficiently flattened and structured into high-performance columnar models (Delta Lake/Parquet) for OLAP analytics.
- Partner with engineering architects to balance real-time, low-latency micro-batches against cost-optimized, high-throughput daily batch workloads.
- Own the product requirements for our FHIR-native MDM Engine, defining rules for deterministic and probabilistic patient matching at scale.
- Productize Survivorship Logic to establish rules for creating and maintaining the "Golden Record" across fragmented data feeds.
- Design crosswalk management capabilities to track source-to-target entity linkages across millions of patient lives.
- Build out the end-to-end exception lifecycle for data validation failures—from isolation in Dead-Letter Queues (DLQs) to structured error reporting.
- Define requirements for Data Steward interfaces and APIs, allowing users to review validation exceptions, perform manual record merges/unmerges, and correct bad data payloads.
- Architect Idempotent Replay Mechanisms to re-ingest corrected payloads through Spark processing pipelines without generating duplicate data or corrupting state.
1. Data Ingress & Decoupled Architecture
2. Big Data & Medallion Analytics Strategy
3. Master Data Management (MDM) & Identity Resolution
4. Error Remediation & Exception Management
Requirements:
- 5+ years of Technical Product Management experience leading complex backend data platforms, distributed data engineering products, or big-data ingestion systems.
- Hands-on Distributed Computing Knowledge: Proven track record productizing solutions powered by Apache Spark (batch or streaming) and distributed data processing frameworks.
- Databricks / Lakehouse Expertise: Deep familiarity with Medallion Architecture (Bronze/Silver/Gold) and Delta Lake/Parquet performance patterns.
- OLAP / Columnar Strategy Mindset: Demonstrated experience optimizing complex, highly nested arrays and hierarchical data structures for analytical query engines.
- Identity & Governance: Direct experience productizing Master Data Management (MDM), entity resolution, or identity matching tools.
- Fault-Tolerant Systems: Experience with event-driven architectures, Dead-Letter Queues (DLQs), and data replay/remediation workflows.
- Technical Proficiency: Ability to comfortably write and execute SQL and Python to analyze raw datasets, validate engine output, and define logic rules
- Experience with healthcare data standards, including HL7 FHIR, C-CDA, or HL7 v2.
- Exposure to healthcare terminologies (SNOMED CT, LOINC, RxNorm, ICD-10).
- Familiarity with SQL-on-FHIR analytics standards.
Skills Required
- 5+ years of Technical Product Management experience leading complex backend data platforms, distributed data engineering products, or big-data ingestion systems
- Proven experience productizing solutions powered by Apache Spark (batch or streaming) and distributed data processing frameworks
- Databricks / Lakehouse expertise including Medallion Architecture (Bronze/Silver/Gold) and Delta Lake/Parquet performance patterns
- Experience optimizing nested/hierarchical data structures for OLAP and columnar analytics
- Direct experience productizing Master Data Management (MDM), entity resolution, or identity matching tools at scale
- Experience with event-driven architectures, Dead-Letter Queues (DLQs), fault-tolerant systems, and data replay/remediation workflows
- Ability to write and execute SQL and Python to analyze datasets, validate outputs, and define logic rules
- Experience with Kafka and streaming ingestion patterns (e.g., Kafka/Spark Structured Streaming)
- Experience with healthcare data standards (HL7 FHIR, C-CDA, HL7 v2)
- Exposure to healthcare terminologies (SNOMED CT, LOINC, RxNorm, ICD-10)
- Familiarity with SQL-on-FHIR analytics standards
What We Do
Smile Digital Health (doing business as Smile CDR Inc.) specializes in delivering fast, secure, compliant data infrastructures as a service to enable and empower interconnectivity for data-intensive sectors such as healthcare. We are a solutions platform helping organizations like governments, researchers, health systems, healthcare providers, and app developers build connected health solutions and products by leveraging our core expertise in health data and HL7 FHIR. Our flagship product, Smile CDR, is the world’s first FHIR-based clinical data repository (CDR) as a service. Smile CDR is a high performance and secure solution built on the principles of Privacy by Design. Flexibility is a hallmark of the service as it can be hosted in the cloud or on-premises depending on your needs. The design was strongly influenced by real-world experience in both the jurisdictional and organizational setting. Smile CDR provides a rich set of features and capabilities including: -Multiple FHIR versions -FHIR Profiles -Full text indexing of clinical records -Type ahead search functionality -FHIR, HL7v2 and custom ETL for data input -Federated identity and identity provider functionality -International character locale support -Terminology services -Auditing -Smart on FHIR support -Rapid deployment -Extensive administrative tooling Leveraging more than 40 years of experience in building enterprise-class systems, our team includes experts in building integrated healthcare systems including FHIR-based solutions such as HAPI, e-prescriptions and other e-health solutions. We were motivated by the need to improve upon existing options for sharing health data within and across organizations. Smile CDR is the maintainer of HAPI FHIR, the prevailing open source reference implementation of FHIR worldwide








