Senior Site Reliability Engineer

Posted 5 Days Ago
Be an Early Applicant
New York, NY, USA
Hybrid
200K-240K Annually
Senior level
Cloud • Healthtech • Internet of Things • Machine Learning • Software
Intelligent Orchestration for Healthcare Operations
The Role
Designs, builds, and operates resilient AWS infrastructure for a healthcare platform. Responsibilities include incident response, disaster recovery, observability, CI/CD, Terraform, GitOps, Kubernetes, infrastructure security, and HIPAA/SOC 2 compliance. The role shares production ownership and on-call responsibilities while improving reliability, failover readiness, and system behavior.
Summary Generated by Built In

About Kontakt.io

Inside health systems, where every second can matter, operations are still spread across dozens of disconnected tools and platforms. Kontakt.io is changing that.

We combine proprietary hardware, AI-powered intelligence, and deep integrations with the technology health systems already have in place to build real-time understanding of what's happening across their operations. That intelligence becomes the execution layer care teams have been missing, helping them make smarter decisions and deliver better patient care.

Backed by Goldman Sachs and trusted by leading health systems including HCA Healthcare, Sutter Health, AdventHealth, Trinity Health, Northwell Health, Cleveland Clinic, and the U.S. Department of Veterans Affairs, we’ve more than doubled our revenue and are rapidly scaling with a clear path toward $100M in annual recurring revenue.

If you're excited to solve hard problems and help health systems deliver better care, we'd love to meet you!

About the role

We're looking for a Senior Site Reliability Engineer to join our Infrastructure Engineering team and get their hands directly into the systems that keep our healthcare platform running for hospitals and care teams who can't afford downtime. This is a builder's seat — you'll carry real operational weight and have direct influence over how our infrastructure evolves.

What you'll do

  • Personally design, build, and operate resilient, self-healing infrastructure across our AWS-based platform

  • Own incident response end-to-end: detection, mitigation, root-cause investigation, and postmortems that actually change how the system behaves next time

  • Design and run disaster-recovery and failover exercises with real RTO/RPO targets — you'll be the one who knows exactly what happens when things break

  • Build out observability that's genuinely tuned — SLIs, SLOs, and alerting people trust, not noise

  • Build and maintain CI/CD pipelines and infrastructure as code (Terraform, GitOps)

  • Work hands-on in Kubernetes, below the abstraction layer — you'll know the system, not just the dashboard

  • Partner day-to-day with our platform lead, sharing real production ownership and on-call

  • Shape our security and compliance posture (HIPAA, SOC 2 Type 2) as it relates to infrastructure handling protected health data

What you bring

  • 4+ years in Site Reliability Engineering or Cloud Infrastructure

  • Deep, current expertise in AWS, Kubernetes, and distributed systems, with the depth to go past the vocabulary

  • Real experience running disaster recovery or failover exercises

  • A track record of driving incident response and postmortems yourself

  • Solid grounding in CI/CD automation, GitOps, and infrastructure as code

  • An appetite for staying close to the system rather than one step removed from it

  • Bonus: healthcare IT, EHR data, or HIPAA/SOC 2-governed environments

  • Bonus: experience with high-traffic, mission-critical SaaS or IoT platforms

Logistics, Perks & Benefits

  • Built for collaboration - our team a hybrid schedule of 3 days/week minimum from our New York City office

  • Equity in a high-growth company scaling toward $400M+ ARR and backed by leading investors

  • Full health, dental, and vision coverage, a 401k, paid time off, paid parental leave and all the tools you need to do your best work

  • Autonomy to solve meaningful problems with work that ships quickly and makes a difference

Compensation

The expected salary range for this role is $200,000 – $240,000 for New York-based candidates. Actual compensation within this range will be determined based on relevant experience, skills, and qualifications. In exceptional cases, where a candidate’s experience or qualifications significantly exceed those anticipated for this role, we may consider the candidate for a more senior level. This role may also be eligible for equity and bonus compensation.

Skills Required

  • 4+ years of experience in Site Reliability Engineering or Cloud Infrastructure
  • Deep, current expertise in AWS, Kubernetes, and distributed systems
  • Experience running disaster recovery or failover exercises
  • Experience driving incident response and postmortems
  • Experience with CI/CD automation, GitOps, and infrastructure as code
  • Healthcare IT, EHR data, or HIPAA/SOC 2-governed environments
  • Experience with high-traffic, mission-critical SaaS or IoT platforms
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: New York, New York
135 Employees
Year Founded: 2013

What We Do

Kontakt is the only AI platform in healthcare that delivers operational predictability — maximizing capacity, safety, and decision making to enhance care delivery for patients and caregivers.

Why Work With Us

Collaborative, mission-driven team solving real healthcare challenges. Build impactful tech, grow fast, and make a difference—together.

Gallery

Gallery

Similar Jobs

Formation Bio Logo Formation Bio

Senior Site Reliability Engineer

Artificial Intelligence • Big Data • Healthtech • Biotech • Pharmaceutical
Easy Apply
Hybrid
3 Locations
150 Employees
186K-232K Annually
Remote or Hybrid
United States
1750 Employees

Zocdoc Logo Zocdoc

Senior Site Reliability Engineer

Healthtech • Information Technology • Software • Telehealth
Easy Apply
Remote or Hybrid
USA
900 Employees
180K-220K Annually

PwC Logo PwC

Site Reliability Engineer

Artificial Intelligence • Professional Services • Business Intelligence • Consulting • Cybersecurity • Generative AI
Hybrid
58 Locations
370000 Employees
151K-187K Annually

Similar Companies Hiring

Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Revel Thumbnail
Aerospace • Hardware • Robotics • Software
Marina Del Rey, California
60 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software
New York, New York
30 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account