Location: Remote (US-based candidates only)
Manager, Site Reliability Engineering (SRE)Position Overview
We are seeking a Manager, Site Reliability Engineering (SRE) to lead our US-based SRE team and drive operational excellence across our production platforms.
This is a player-coach leadership role that combines people management with hands-on technical leadership. You will mentor and grow a team of Site Reliability Engineers while actively participating in major incident response, reliability initiatives, and operational reviews. The role is a key part of our global follow-the-sun support model and requires close collaboration with SRE leadership in India.
Key ResponsibilitiesTeam Leadership & Development- Lead, mentor, and develop a US-based team of Site Reliability Engineers.
- Conduct regular 1:1s, performance reviews, and career development discussions.
- Own hiring, onboarding, and retention efforts as the team scales.
- Foster a culture of ownership, blameless postmortems, and continuous improvement.
- Lead day-to-day production operations and ensure timely incident triage, resolution, and escalation.
- Serve as an escalation point and incident commander for major production incidents.
- Drive problem management and root cause analysis processes.
- Carry PagerDuty on-call escalation responsibilities for critical issues.
- Track and report operational KPIs, SLAs, and SLOs, including availability, MTTR, and incident trends.
- Improve system reliability, observability, and resilience using Datadog and related tooling.
- Drive automation, self-healing capabilities, and runbook maturity.
- Partner with Development, DevOps, DevSecOps, and Engineering teams to embed reliability into the SDLC.
- Contribute hands-on to tooling, automation, and technical reviews as needed.
- Coordinate closely with SRE leadership in India to ensure seamless follow-the-sun coverage.
- Represent the US SRE organization in cross-functional planning and operational reviews.
- Communicate effectively with both technical and non-technical stakeholders.
- Maintain high-quality documentation for incidents, postmortems, runbooks, and operational procedures.
- Ensure adherence to healthcare and fintech compliance standards, including HIPAA, PCI DSS, SOC 2, ISO 27001, and HITRUST.
- 5–8 years of experience in Site Reliability Engineering, DevOps, Production Support, or Platform Engineering.
- 1–2+ years of experience leading, mentoring, or managing engineers.
- Demonstrated success operating in a player-coach leadership model.
- Strong hands-on experience with production incident management and escalation processes.
- Proficiency with Datadog or similar observability platforms.
- Hands-on experience with Kubernetes and Docker in production environments.
- Strong scripting or programming skills in PowerShell, Bash, Python, Java, or C#.
- Experience with Helm, CI/CD pipelines, and deployment automation.
- Working knowledge of ITIL processes and Agile methodologies.
- Experience working with SQL, MySQL, or NoSQL databases.
- Excellent communication and stakeholder management skills.
- Willingness to participate in PagerDuty on-call escalation and work within a global follow-the-sun operating model.
- Experience with cloud platforms such as AWS, Azure, or GCP.
- Experience building or scaling SRE teams and on-call programs.
- Experience defining and managing SLOs, SLIs, and error budgets.
- Prior experience in the healthcare or fintech industry.
- Knowledge of security and compliance frameworks relevant to regulated environments.
- Competitive compensation and comprehensive benefits.
- Unlimited PTO.
- Fully remote work environment (US-based).
- Opportunity to lead and grow a high-impact SRE organization.
- Exposure to modern cloud-native technologies and large-scale reliability challenges.
- Collaborative culture focused on innovation, learning, and continuous improvement.
- Meaningful work that directly impacts healthcare technology and millions of members.
We are looking for a technically strong SRE leader who enjoys building teams, improving operational maturity, and remaining hands-on during critical production events. The ideal candidate combines leadership, systems thinking, and automation expertise to help scale reliability practices across a fast-growing Healthcare FinTech organization.
NationsBenefits is an Equal Opportunity Employer.
Skills Required
- 5-8 years of experience in Site Reliability Engineering, DevOps, Production Support, or Platform Engineering
- 1-2+ years of experience leading, mentoring, or managing engineers
- Demonstrated success operating in a player-coach leadership model
- Strong hands-on experience with production incident management and escalation processes
- Proficiency with Datadog or similar observability platforms
- Hands-on experience with Kubernetes and Docker in production environments
- Strong scripting or programming skills in PowerShell, Bash, Python, Java, or C#
- Experience with Helm, CI/CD pipelines, and deployment automation
- Working knowledge of ITIL processes and Agile methodologies
- Experience working with SQL, MySQL, or NoSQL databases
- Excellent communication and stakeholder management skills
- Willingness to participate in PagerDuty on-call escalation and work within a global follow-the-sun operating model
- Experience with cloud platforms such as AWS, Azure, or GCP
- Experience building or scaling SRE teams and on-call programs
- Experience defining and managing SLOs, SLIs, and error budgets
- Prior experience in the healthcare or fintech industry
- Knowledge of security and compliance frameworks relevant to regulated environments (HIPAA, PCI DSS, SOC 2, ISO 27001, HITRUST)
NationsBenefits Compensation & Benefits Highlights
The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about NationsBenefits and has not been reviewed or approved by NationsBenefits.
-
Fair & Transparent Compensation — Pay is frequently characterized as decent or good for certain entry-level and frontline roles, with timely pay also highlighted. Compensation is sometimes positioned as competitive relative to the work performed in those positions.
-
Leave & Time Off Breadth — Unlimited PTO is described as available for some salaried roles, which can increase perceived flexibility. Paid holidays and paid time off are presented as part of the standard package for eligible employees.
-
Wellbeing & Lifestyle Benefits — A fitness stipend and occasional company-sponsored outings or training-related perks are included among the extra benefits. These additions can modestly strengthen the overall rewards experience beyond core insurance.
NationsBenefits Insights
What We Do
NationsBenefits® is a leading supplemental benefits company providing managed care organizations with innovative healthcare solutions helping to promote independence, health, and well-being for more than 20 million members across the U.S. When the company was founded in 2015 by Glenn Parker, M.D., we set out to disrupt the healthcare industry. In 2020, we rebranded to NationsBenefits to expand the company’s core offering and broaden the scope of our clinically focused services. Today, we surpass traditional benefit management programs by helping our health plan partners drive growth, improve outcomes, reduce costs, and delight members. Our best-in-class service model engages members in meaningful and measurable ways with technology-based solutions tailored to the unique needs of each population.








