The role partners with product owners, data modelers, domain stakeholders, and platform teams to translate business requirements into scalable pipelines, reusable frameworks, and clear standards for storing, processing, and moving data. The Senior Data Engineer practices DataOps — CI/CD/CT, automated validation, observability, and Agile delivery — and provides technical leadership through design reviews, mentoring, and engineering best practices.
They own reliable datasets that enable analytics and reporting — not ad-hoc analysis—and ensure pipelines support lifecycle management, resiliency, and governed access across the Lakehouse (raw → refine → publish).
Key Responsibilities
- Design and implement ELT/ETL solutions for batch and streaming ingestion, integration, refinement, and publish patterns on the Lakehouse.
- Develop reusable data processing frameworks and configuration-driven pipelines using Python and PySpark (EMR, Glue, or comparable Spark runtimes).
- Build and maintain scalable orchestration workflows (e.g., Airflow) for production data delivery, including retries, historical loads, and operational runbooks.
- Implement data quality checks, validation frameworks, and monitoring so data products meet defined contracts and SLAs.
- Apply DataOps practices: Git-based development, CI/CD/CT for data pipelines, automated testing, and controlled promotion across environments.
- Contribute to data lifecycle practices (retention, archival, disaster recovery / resiliency considerations) in partnership with platform and governance teams.
- Support platform modernization and cloud migration of legacy data flows into Lakehouse patterns (Iceberg on S3, governed catalog access).
- Collaborate with stakeholders to map technical designs to business processes, non-functional requirements, and consumption needs (Athena, Redshift, APIs, exports, streams).
- Establish and document standards, naming/conventions, and engineering practices; participate in Agile ceremonies and cross-team delivery.
- Provide technical leadership: mentor engineers, conduct design and code reviews, and continuously improve reliability, performance, and cost efficiency.
Minimum Qualifications:
- Bachelor's degree in Computer Science, Information Systems, Engineering, or a related technical field, or equivalent work experience.
- 8+ years of overall IT experience.
- 5+ years of hands-on experience designing and developing enterprise-scale data engineering solutions.
- Strong experience developing scalable data pipelines and reusable frameworks using Python and PySpark.
- Experience implementing enterprise data ingestion, integration, and transformation solutions using AWS Glue, dbt, Apache Spark, or comparable technologies.
- Strong understanding of data warehousing concepts, dimensional modeling, and modern data lake / Lakehouse architectures.
- Experience with cloud platforms such as AWS, Microsoft Azure, or Google Cloud Platform (AWS preferred).
- Strong SQL expertise with relational databases; familiarity with NoSQL databases is a plus.
- Experience with distributed data processing technologies such as Apache Spark, Amazon EMR, or Hadoop-based platforms.
- Experience using Git-based source control and Agile software development methodologies.
- Strong analytical, problem-solving, and communication skills, with the ability to collaborate across technical and business teams.
Preferred Qualifications:
- Experience designing cloud-native data platforms using AWS services such as S3, Glue, EMR, Athena, Redshift, Lambda, Lake Formation, IAM, and CloudWatch.
- Hands-on experience with Apache Iceberg (or similar open table formats) and governed Lakehouse patterns.
- Experience implementing CI/CD pipelines, DataOps practices, continuous testing (CT), and infrastructure automation (e.g., Terraform).
- Experience with workflow orchestration tools such as Apache Airflow, MWAA, AWS Step Functions, or similar platforms.
- Experience developing RESTful APIs or other access layers and integrating with enterprise applications for data-product consumption.
- Experience building and supporting real-time or streaming platforms using Kafka, Kinesis, or Spark Structured Streaming.
- Strong understanding of data quality, observability, monitoring, and automated validation frameworks.
- Knowledge of data governance, metadata management, lineage, and enterprise data catalog solutions.
- Strong understanding of data security, encryption, access controls, and healthcare regulatory compliance (HIPAA/PHI).
- Experience optimizing distributed workloads for scalability, reliability, and cloud cost efficiency.
- Experience mentoring engineers, conducting design and code reviews, and establishing engineering best practices.
Compliance and Regulatory Responsibilities: N/A
License/Certification: N/A
WE ARE AN EQUAL OPPORTUNITY EMPLOYER. HF Management Services, LLC complies with all applicable laws and regulations. Applicants and employees are considered for positions and are evaluated without regard to race, color, creed, religion, sex, national origin, sexual orientation, pregnancy, age, disability, genetic information, domestic violence victim status, gender and/or gender identity or expression, military status, veteran status, citizenship or immigration status, height and weight, familial status, marital status, or unemployment status, as well as any other legally protected basis. HF Management Services, LLC shall not discriminate against any disabled employee or applicant in regard to any position for which the employee or applicant is otherwise qualified.
If you have a disability under the Americans with Disability Act or a similar law and want a reasonable accommodation to assist with your job search or application for employment, please contact us by sending an email to [email protected] or calling 212-519-1798 . In your email please include a description of the accommodation you are requesting and a description of the position for which you are applying. Only reasonable accommodation requests related to applying for a position within HF Management Services, LLC will be reviewed at the e-mail address and phone number supplied. Thank you for considering a career with HF Management Services, LLC.
Know Your Rights
All hiring and recruitment at Healthfirst is transacted with a valid “@healthfirst.org” email address only or from a recruitment firm representing our Company. Any recruitment firm representing Healthfirst will readily provide you with the name and contact information of the recruiting professional representing the opportunity you are inquiring about. If you receive a communication from a sender whose domain is not @healthfirst.org, or not one of our recruitment partners, please be aware that those communications are not coming from or authorized by Healthfirst. Healthfirst will never ask you for money during the recruitment or onboarding process.
Hiring Range*:
Greater New York City Area (NY, NJ, CT residents): $134,600 - $194,480
All Other Locations (within approved locations): $119,600 - $177,905
As a candidate for this position, your salary and related elements of compensation will be contingent upon your work experience, education, licenses and certifications, and any other factors Healthfirst deems pertinent to the hiring decision.
In addition to your salary, Healthfirst offers employees a full range of benefits such as, medical, dental and vision coverage, incentive and recognition programs, life insurance, and 401k contributions (all benefits are subject to eligibility requirements). Healthfirst believes in providing a competitive compensation and benefits package wherever its employees work and live.
*The hiring range is defined as the lowest and highest salaries that Healthfirst in “good faith” would pay to a new hire, or for a job promotion, or transfer into this role.
Skills Required
- Bachelor's degree in Computer Science, Information Systems, Engineering, or related field (or equivalent experience)
- 8+ years of overall IT experience
- 5+ years hands-on experience designing and developing enterprise-scale data engineering solutions
- Strong experience developing scalable data pipelines and reusable frameworks using Python and PySpark
- Experience implementing enterprise data ingestion, integration, and transformation using AWS Glue, dbt, Apache Spark, or comparable technologies
- Strong understanding of data warehousing concepts, dimensional modeling, and modern data lake/Lakehouse architectures
- Experience with cloud platforms such as AWS (preferred), Azure, or GCP
- Strong SQL expertise with relational databases; familiarity with NoSQL is a plus
- Experience with distributed data processing technologies such as Apache Spark, Amazon EMR, or Hadoop-based platforms
- Experience using Git-based source control and Agile software development methodologies
- Strong analytical, problem-solving, and communication skills with cross-team collaboration ability
- Experience designing cloud-native data platforms using AWS services (S3, Glue, EMR, Athena, Redshift, Lambda, Lake Formation, IAM, CloudWatch)
- Hands-on experience with Apache Iceberg or similar open table formats and governed Lakehouse patterns
- Experience implementing CI/CD pipelines, DataOps practices, continuous testing, and infrastructure automation (e.g., Terraform)
- Experience with workflow orchestration tools such as Apache Airflow, MWAA, or AWS Step Functions
- Experience building/supporting real-time or streaming platforms using Kafka, Kinesis, or Spark Structured Streaming
- Knowledge of data governance, metadata management, lineage, and enterprise data catalog solutions
- Understanding of data security, encryption, access controls, and healthcare regulatory compliance (HIPAA/PHI)
- Experience mentoring engineers, conducting design and code reviews, and establishing engineering best practices
Healthfirst, Inc Compensation & Benefits Highlights
The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Healthfirst, Inc and has not been reviewed or approved by Healthfirst, Inc.
-
Healthcare Strength — Healthcare coverage is positioned as a comprehensive offering, including medical/dental/vision options and additional mental-health support through Spring Health. Wellness credits and disability coverage further strengthen the perceived breadth of health-related support.
-
Leave & Time Off Breadth — Time off appears robust, with paid time off plus a set of paid holidays and additional early office-closure days. A designated Social Justice day and noted parental/bereavement leave options add to the overall leave breadth.
-
Retirement Support — Retirement support is framed around a 401(k) program with a company match that can reach a meaningful level after a tenure milestone. This match is repeatedly highlighted as a notable component of total rewards value.
Healthfirst, Inc Insights
What We Do
Healthfirst is a provider-sponsored health insurance company that serves 1.8 million members in downstate New York. Healthfirst offers top-quality Medicaid, Medicare Advantage, Child Health Plus, and Managed Long Term Care plans. Healthfirst Leaf Qualified Health Plans and the Healthfirst Essential Plan are offered on NY State of Health, The Official Health Plan Marketplace. Healthfirst offers Healthfirst Pro and Pro Plus, Exclusive Provider Organization (EPO) plans for small-business owners and their employees, and Healthfirst Total, an EPO for individuals. For more information on Healthfirst, visit www.healthfirst.org







