Job Summary
We are seeking a highly skilled Senior Site Reliability Engineer to ensure the reliability, scalability, and performance of our cloud-native systems and applications. This role drives automation, monitoring, and incident management practices while operating Kubernetes platforms (Amazon EKS) and leveraging observability tools such as Prometheus, Grafana, Dynatrace, and OpenSearch to maintain high availability and operational excellence.
Key Responsibilities
• Design, build, and maintain highly available, scalable, and reliable production systems.
• Define and manage SLIs, SLOs, and SLAs to drive system reliability.
• Automate infrastructure provisioning and operations using Terraform (IaC).
• Operate and manage cloud-native platforms, including Amazon EKS.
• Implement and maintain monitoring, logging, and alerting using Prometheus, Grafana, Dynatrace, and OpenSearch.
• Lead incident management — on-call rotation, production troubleshooting, and root cause analysis (RCA).
• Drive AI-assisted investigations as a core part of incident response, and build and maintain the prompts, integrations, and guardrails that make AI-driven triage and RCA reliable.
• Improve system reliability through automation, self-healing mechanisms, and performance tuning.
• Collaborate with development teams to improve application reliability, scalability, and deployment processes.
• Build and maintain CI/CD pipelines (GitLab CI, Jenkins, or GitHub Actions) for fast, reliable software delivery.
• Perform capacity planning and cost optimization for infrastructure and services.
• Ensure security, compliance, and best practices across infrastructure and applications.
Required Qualifications
• Bachelor’s degree in computer science or a related field, or equivalent practical experience.
• 7+ years in Site Reliability Engineering or Platform Engineering.
• Proven ownership of incident management, on-call support, and root cause analysis (RCA) for production systems.
• Strong expertise in Terraform and Infrastructure as Code.
• Hands-on experience with AWS and EKS.
• Strong understanding of monitoring, logging, and observability (Prometheus, Grafana, Dynatrace, OpenSearch).
• Proficiency in Python, Java, Go, or Bash.
• Experience with Agile development and CI/CD pipelines (GitLab CI, Jenkins, or GitHub Actions).
• Strong problem-solving, documentation, and communication skills.
• Proven ability to troubleshoot effectively in high-pressure production environments.
• Experience with autoscaling, performance tuning, and cost optimization.
• Familiarity with AI-assisted automation tools and a track record of using them to reduce toil and improve reliability.
Preferred Skills
• Docker and Linux administration.
• Build systems and dependency management (Maven, Gradle, npm).
• Additional AWS services: Cognito, WAF, Elasticsearch, SNS, SQS, S3, Systems Manager.
• Database infrastructure knowledge (RDS, MySQL, SQL Server).
• Cloud or Kubernetes certifications.
Skills Required
- Bachelor's degree in computer science or a related field, or equivalent practical experience
- 7+ years of experience in Site Reliability Engineering or Platform Engineering
- Proven ownership of incident management, on-call support, and root cause analysis for production systems
- Strong expertise in Terraform and Infrastructure as Code
- Hands-on experience with AWS and Amazon EKS
- Strong understanding of monitoring, logging, and observability using Prometheus, Grafana, Dynatrace, and OpenSearch
- Proficiency in Python, Java, Go, or Bash
- Experience with Agile development and CI/CD pipelines using GitLab CI, Jenkins, or GitHub Actions
- Strong problem-solving, documentation, and communication skills
- Ability to troubleshoot effectively in high-pressure production environments
- Experience with autoscaling, performance tuning, and cost optimization
- Familiarity with AI-assisted automation tools and experience using them to reduce toil and improve reliability
- Experience with Docker and Linux administration
- Experience with build systems and dependency management, including Maven, Gradle, or npm
- Knowledge of additional AWS services, including Cognito, WAF, Elasticsearch, SNS, SQS, S3, or Systems Manager
- Knowledge of database infrastructure, including RDS, MySQL, or SQL Server
- Cloud or Kubernetes certifications
Clearwater Analytics (CWAN) Compensation & Benefits Highlights
-
Healthcare Strength — Employer-provided medical, dental, and vision coverage is consistently listed in current job postings, with disability insurance also referenced. Feedback suggests these core health benefits are solid even if plan richness is not portrayed as top-tier across sources.
-
Retirement Support — A 401(k) plan with employer matching is consistently cited in company and employer-verified materials. This reliable match supports long-term savings and is presented as a standard component of total rewards.
-
Leave & Time Off Breadth — Immediate eligibility for paid time off, holidays, and volunteer time, along with parental leave, appears across job postings. This breadth of leave options provides practical flexibility from day one.
Clearwater Analytics (CWAN) Insights
What We Do
CWAN was founded on a simple belief: investment professionals deserve modern technology that actually works for them. Not legacy systems that slow them down. Not fragmented data that creates confusion. But one comprehensive platform that gives you complete visibility and crystal-clear insights. The result? Investment management that works as seamlessly as your investment strategy. Since our founding in 2004, CWAN has been the trusted technology partner powering the world’s leading institutional investors — from insurance companies, asset managers, and hedge funds to asset owners like corporations, endowments, and pension funds managing over $10 trillion in assets.
Why Work With Us
We continue to grow, fueled by a strong foundation, an ambitious vision, and a commitment to delivering exceptional value to our clients, partners, and team members around the world. What started as a bold idea in Boise, Idaho has rapidly transformed into a global presence. We’ve expanded our footprint significantly—now operating out of 24 offices
Gallery
Clearwater Analytics (CWAN) Offices
Hybrid Workspace
Employees engage in a combination of remote and on-site work.


_1.jpg)








_1.jpg)





