The Role
Own production EKS cluster availability, performance, observability, and incident response. Define SLOs, SLIs, and error budgets; lead post-mortems; and maintain blameless reliability practices. Build infrastructure using Terraform, Helm, and GitOps, while partnering with engineering on capacity planning, load testing, and chaos engineering. Optimize AWS autoscaling, networking, and costs, participate in on-call rotations, and resolve underlying causes of incidents.
Summary Generated by Built In
RapidAI is the trusted leader in deep clinical AI, helping hospitals deliver faster, more informed care through intelligent imaging and integrated workflows. The Rapid Enterprise™ Platform supports disease states across the care spectrum, but it’s our clinical depth that drives the most meaningful impact — improving decision-making, patient outcomes, and health-system performance. Used by more than 2,500 hospitals in over 100 countries and backed by 700+ clinical studies, including research that helped expand national stroke-treatment guidelines, RapidAI is the most clinically validated AI platform in healthcare.
- Own the availability, performance, and incident response for Rapid's production EKS clusters
- Design and operate the full observability stack — metrics, logs, traces — with
Open Telemetry as the foundation - Define and track SLOs/SLIs/error budgets; lead post-mortems and drive blameless culture
- Build and maintain infrastructure-as-code using Terraform, Helm, and GitOps patterns
- Partner with engineering to bake reliability in early — capacity planning, load testing, chaos engineering
- Tune autoscaling, networking, and cost efficiency across AWS workloads
- On-call rotation with the expectation you'll also fix the underlying cause, not just the alert
- 10+ years in SRE, DevOps, or infrastructure engineering roles
- Deep AWS expertise — EKS, EC2, VPC, IAM, RDS, S3, CloudWatch, and the
surrounding ecosystem - Production Kubernetes experience at scale: multi-cluster, multi-tenant, real traffic
- Hands-on Open Telemetry instrumentation and pipeline ownership (collectors, exporters, backends)
- Strong foundation in Linux, networking, and distributed systems fundamentals
- Experience with observability platforms (Prometheus, Grafana, Jaeger, or equivalents)
Comfortable writing automation in Go, Python, or Bash — you reach for code when the GUI runs out - Startup mindset: you make decisions with incomplete information and iterate quickly
What You Do:
What We Looking For:
RapidAI is committed to creating an inclusive and diverse workplace. We provide equal employment opportunities to all employees and applicants and prohibit discrimination and harassment of any type in regard to race, color, religion, age, sex, national origin, disability status, genetics, protected veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by federal, state or local laws.
Skills Required
- 10+ years of experience in SRE, DevOps, or infrastructure engineering roles
- Deep AWS expertise, including EKS, EC2, VPC, IAM, RDS, S3, and CloudWatch
- Production Kubernetes experience at scale, including multi-cluster, multi-tenant environments with real traffic
- Hands-on OpenTelemetry instrumentation and ownership of collectors, exporters, and backends
- Strong Linux, networking, and distributed systems fundamentals
- Experience with observability platforms such as Prometheus, Grafana, or Jaeger
- Ability to write automation in Go, Python, or Bash
- Startup mindset and ability to make decisions with incomplete information
Am I A Good Fit?
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.
Success! Refresh the page to see how your skills align with this role.
The Company
What We Do
RapidAI is the global leader in using AI to combat life-threatening vascular and neurovascular conditions. Leading the next evolution of clinical decision-making and patient workflow, RapidAI is empowering physicians to make faster decisions for better patient outcomes. Based on intelligence gained from over 5 million scans in more than 2,000 hospitals in over 60 countries, the Rapid® platform transforms care coordination, offering care teams a level of patient visibility never before possible. RapidAI — where AI meets patient care.









