Senior Site Reliability Engineer (US)

Posted Yesterday
Hiring Remotely in United States
Remote
135K-170K Annually
Senior level
Big Data • Analytics
The Role
Own production reliability across Azure, colocation, and edge Kubernetes environments. Build observability, alerting, automated recovery, high-availability architectures, deployment automation, and cost optimization. Lead incident response, troubleshoot distributed systems, improve disaster recovery and operational maturity, support platform lifecycle management, and mentor engineers. Participate in separate weekday and weekend on-call rotations for customer-facing systems.
Summary Generated by Built In
Senior Site Reliability Engineer

Remote | US


About Climavision  

At Climavision, we’re rebuilding climate technology from the ground up and changing the way we see weather. We merge the power of our proprietary, high-resolution weather radar and satellite network with advanced weather prediction modelling and decades of industry expertise to reduce existing coverage gaps and drastically improve forecasting ability. Our revolutionary new approach to climate technology weather solutions is poised to help reduce the economic risks of climate change on companies, governments, and societies alike. We are backed by The Rise Fund, the world’s largest global impact platform committed to achieving measurable, positive social and environmental outcomes alongside competitive financial returns. Climavision is headquartered in Louisville, KY, with research and development operations in Raleigh, NC.   


The Work

Are you an experienced Site Reliability Engineer who thrives at the intersection of software engineering and production operations? Do you take pride in keeping mission-critical customer systems reliable under real-world operational pressure? Are you looking for an opportunity to own production reliability for a modern hybrid infrastructure platform spanning cloud, colocation, and edge environments?

If so, we have an exceptional opportunity for you.

Climavision is seeking a Senior Site Reliability Engineer to contribute towards reliability, operational excellence, and production resilience across the company's platform and data services. This role sits on a shared SRE team that supports the full business rather than a single product line, covering both the radar network and the weather intelligence sides of the company as priorities shift. A central focus of this role is building the observability layer that puts the health of the full fleet in one place, and then automating recovery so that systems heal themselves. Multi-cluster and multi-replica high availability across our distributed edge fleet remains a core part of the work.

This is a hands-on engineering role for someone who is equally comfortable troubleshooting Kubernetes clusters, leading incident response, and improving operational maturity across the organization. The successful candidate will combine deep production operations expertise with a disciplined approach to reliability engineering and strong automation skills.

Climavision operates a hybrid infrastructure footprint spanning Microsoft Azure, colocation data centers, and edge Kubernetes clusters, deployed alongside weather radar systems. This role will drive production reliability across Azure, colocation, and edge environments. Right-sizing cluster resources and migrating workloads off Azure to reduce spend are active priorities for the team.

35% Kubernetes Platform Reliability and Operations

30% Production Reliability Engineering and Incident Response

20% Observability, Monitoring, and Alerting

15% Automation, Recovery, and Cost Optimization


Primary Responsibilities:

• Own production reliability for Climavision's customer-facing platform and data services across Azure, colocation, and edge Kubernetes environments.

• Work as part of a shared SRE function supporting the whole company rather than a single product line, taking on work across both the radar network and the weather intelligence sides of the business as priorities shift.

• Contribute to the definition and improvement of SLIs, SLOs, alerting standards, and operational metrics used to measure platform reliability.

• Build and own the observability layer for the fleet. Today the underlying metrics exist but are only reachable from the command line inside each cluster. This role is responsible for surfacing that data in shared dashboards and building the alerting that tells the team something is going wrong before a customer does.

• Design and build automated recovery and self-healing for production systems, so that common failure modes are detected and remediated without human intervention.

• Optimize cluster resourcing and cost, including right-sizing workloads and nodes and supporting the migration of workloads off Azure to reduce spend.

• Support and coordinate production incident response efforts, including troubleshooting, mitigation, communication, and postmortem analysis.

• Diagnose and resolve complex production issues across application services, Kubernetes infrastructure, storage, and distributed systems.

• Drive multi-replica and multi-cluster high availability across Climavision's services, including workload placement, scheduling, and deployment patterns that allow services to run safely as multiple replicas across multiple clusters.

• Contribute to the multi-cluster high-availability strategy across Climavision's hybrid fleet, including active-active and active-passive failover behavior, traffic routing, data replication considerations, and graceful degradation when a cluster becomes unavailable.

• Operate and improve Climavision's self-managed Kubernetes platform spanning cloud-hosted, colocation, and edge clusters, with a focus on availability, resiliency, recovery, and operational performance.

• Ensure Kubernetes platform lifecycle activities including upgrades, patching, cluster health, node management, and production change management are executed in a manner that preserves service availability and minimizes customer-facing risk.

• Improve reliability and operational maturity of production platform services, including observability, autoscaling, ingress, and distributed storage. Partner with the teams responsible for the underlying networking and security primitives rather than owning those areas directly.

• Design and validate Kubernetes workloads for resiliency, scalability, and operational efficiency, including autoscaling behavior, workload placement, resource management, and graceful degradation strategies.

• Partner with software engineering teams across the company to improve production readiness, resiliency patterns, deployment safety, and operational visibility before services reach production.

• Maintain and improve deployment pipelines, Helm charts, Kubernetes manifests, and infrastructure automation supporting safe and repeatable production releases.

• Support and evolve Climavision's observability platform, including metrics, logging, distributed tracing, dashboarding, and alerting.

• Conduct performance engineering and capacity-planning efforts for customer-facing services during peak weather-event demand.

• Help facilitate blameless postmortem reviews and drive operational follow-up items through completion.

• Improve disaster recovery, failover, and business continuity capabilities across cloud, colocation, and edge environments.

• Drive operational excellence initiatives, including automation, reduction of operational toil, game days, production readiness reviews, and reliability best practices.

• Contribute as a senior technical resource and mentor on reliability engineering and production operations practices.


On-Call Expectation:

Climavision operates customer-facing production systems under contractual SLAs that do not pause outside business hours. The Senior Site Reliability Engineer will participate in a rotating on-call schedule made up of two separate rotations:

• A weekday rotation. The engineer on a weekday shift is the first point of contact for production incidents and for engineering teams needing support during the business week.

• A separate weekend rotation, so that the engineer carrying weekday support is not also carrying the weekend.

At current and planned team size, engineers can expect a weekday shift roughly every five weeks and a weekend shift roughly every five weeks. The two are scheduled as far apart from each other as the rotation allows, so that a weekday shift and a weekend shift do not fall close together.


Qualifications

• A bachelor's degree in computer science, software engineering, or a related field; equivalent professional experience considered.

• Minimum of 7 years of experience in Site Reliability Engineering, DevOps, Production Engineering, Platform Engineering, or a related infrastructure-focused role, with at least 4 years in a role formally titled Site Reliability Engineer or carrying explicit SLO / error-budget accountability.

• Deep, hands-on experience operating native Kubernetes. Managed distributions such as AKS and EKS are acceptable, but experience running native or self-managed Kubernetes is strongly preferred and is the primary technical requirement for this role.

• Demonstrated experience optimizing Kubernetes clusters, including right-sizing workloads and node pools, resource management, and reducing infrastructure cost without sacrificing reliability. Be prepared to walk through a specific cluster optimization project you led.

• Demonstrated experience increasing operational visibility, including building dashboards, metrics pipelines, and alerting in an environment where little or none existed before.

• Experience designing and operating workloads for safe horizontal scaling across multiple replicas, including idempotency, concurrency, and state handling considerations.

• Experience designing or operating multi-cluster high-availability architectures, including failover behavior, traffic routing, and cross-cluster service deployment.

• Experience supporting customer-facing production systems with uptime, reliability, and incident-response responsibilities.

• Experience diagnosing and resolving production incidents across application, platform and Kubernetes infrastructure layers, including workload scheduling, storage, ingress, and cluster-level failures.

• Experience operating Kubernetes outside of strictly managed cloud environments, including bare-metal, colocation, edge, or hybrid infrastructure.

• Experience with Kubernetes operational tooling and ecosystem technologies such as Rancher, Helm, autoscaling frameworks, observability stacks, or distributed storage systems.

• Strong understanding of infrastructure automation and Infrastructure as Code concepts using tools such as Terraform and Ansible.

• Experience supporting CI/CD and production deployment pipelines. GitHub Actions is used for CI/CD at Climavision.

• Experience with monitoring, logging, and observability platforms such as DataDog, Prometheus, Grafana, Loki, OpenTelemetry, or comparable technologies.

• Experience operating distributed systems and microservice-based architectures in production environments.

• Working knowledge of Microsoft Azure infrastructure.

• Strong troubleshooting skills across infrastructure, application, and platform layers.

• Demonstrated experience participating in a structured production on-call rotation supporting business-critical systems.

• Working familiarity with Jira, Confluence, and Microsoft Entra, which the team uses day to day for ticketing, documentation, and authentication.

• Strong written and verbal communication skills, including incident documentation and postmortem authoring.

• Experience working in start-up, scale-up, or other fast-moving engineering environments, and comfort with the pace and ambiguity that comes with them.


Nice to have, but not required:

• Experience operating Kubernetes platforms using RKE2 and Rancher, which is how Climavision manages its clusters.

• Experience with Octopus Deploy.

• Experience with SOC 2 or comparable security auditing and compliance work.

• Experience supporting hybrid cloud and colocation infrastructure environments.

• Experience with service mesh technologies such as Istio.

• Experience with Kubernetes-native storage platforms such as Longhorn.

• Experience operating PostgreSQL or PostGIS in Kubernetes environments.

• Experience with distributed messaging systems such as RabbitMQ or NATS.

• Experience supporting GPU-enabled workloads in Kubernetes.

• Familiarity with reliability engineering practices, including SLIs, SLOs, error budgets, and operational maturity metrics.


Physical Demands & Work Environment:  

• This is a full-time, exempt position

• Fully Remote - United States   

• This job requires frequent use of a computer to complete tasks, attend meetings, and communicate via Microsoft Teams. 


Once you land this position, you’ll get to enjoy:  

• Benefits of a dynamic and growing organization  

• A challenging, hands-on role that will have real impact on the business  

• Competitive compensation 

• Comprehensive benefits package

• 401(k) Savings Plan 

• Medical/Dental/Vision Benefits 

• Health Savings Account (HSA) and Flexible Spending Account (FSA)

• Unlimited Paid Time-off 

• 11 Paid Holidays

• Paid Parental Leave 

• Company Paid Short-term Disability (STD)

• Company Paid Long-term Disability (LTD)

• Company Paid Life Insurance  

The salary range for this position is $130,000-170,000 annually, however Climavision considers several factors when extending an offer of employment including but not limited to, the applicant’s education, experience, the responsibilities of the role, training, knowledge, skills, and abilities, as well as internal equity and alignment with market data. Any offer of employment is contingent on completion of a background check to company standard. Please note this job description is not designed to cover or contain a comprehensive listing of activities, duties or responsibilities that are required of the employee for this job. Duties, responsibilities, and activities may change at any time with or without notice.

Climavision is an equal opportunity employer. All aspects of employment including the decision to hire, promote, discipline, or discharge, will be based on merit, competence, performance, and business needs. We do not discriminate on the basis of race, color, religion, marital status, age, national origin, ancestry, physical or mental disability, medical condition, pregnancy, genetic information, gender, sexual orientation, gender identity or expression, veteran status, or any other status protected under federal, state, or local law.

Skills Required

  • Bachelor’s degree in computer science, software engineering, or a related field; equivalent professional experience considered
  • At least 7 years of experience in Site Reliability Engineering, DevOps, Production Engineering, Platform Engineering, or a related infrastructure-focused role
  • At least 4 years in a formally titled Site Reliability Engineer role or a role with explicit SLO and error-budget accountability
  • Deep hands-on experience operating native or self-managed Kubernetes
  • Experience optimizing Kubernetes clusters through workload and node-pool right-sizing, resource management, and infrastructure cost reduction
  • Experience building dashboards, metrics pipelines, and alerting in environments with limited prior observability
  • Experience designing and operating horizontally scaled workloads across multiple replicas, including idempotency, concurrency, and state handling
  • Experience designing or operating multi-cluster high-availability architectures, including failover, traffic routing, and cross-cluster deployment
  • Experience supporting customer-facing production systems with uptime, reliability, and incident-response responsibilities
  • Experience diagnosing and resolving incidents across application, platform, and Kubernetes infrastructure layers
  • Experience operating Kubernetes in bare-metal, colocation, edge, or hybrid infrastructure environments
  • Experience with Kubernetes operational tooling and ecosystem technologies such as Rancher, Helm, autoscaling frameworks, observability stacks, or distributed storage
  • Strong understanding of infrastructure automation and Infrastructure as Code using tools such as Terraform and Ansible
  • Experience supporting CI/CD and production deployment pipelines
  • Experience with monitoring, logging, and observability platforms such as Datadog, Prometheus, Grafana, Loki, or OpenTelemetry
  • Experience operating distributed systems and microservice architectures in production
  • Working knowledge of Microsoft Azure infrastructure
  • Strong troubleshooting skills across infrastructure, application, and platform layers
  • Experience participating in a structured production on-call rotation for business-critical systems
  • Working familiarity with Jira, Confluence, and Microsoft Entra
  • Strong written and verbal communication skills, including incident documentation and postmortem authoring
  • Experience in startup, scale-up, or other fast-moving engineering environments
  • Experience operating Kubernetes with RKE2 and Rancher
  • Experience with Octopus Deploy
  • Experience with SOC 2 or comparable security auditing and compliance work
  • Experience supporting hybrid cloud and colocation infrastructure environments
  • Experience with service mesh technologies such as Istio
  • Experience with Kubernetes-native storage platforms such as Longhorn
  • Experience operating PostgreSQL or PostGIS in Kubernetes environments
  • Experience with distributed messaging systems such as RabbitMQ or NATS
  • Experience supporting GPU-enabled workloads in Kubernetes
  • Familiarity with reliability engineering practices including SLIs, SLOs, error budgets, and operational maturity metrics
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Louisville, KY
34 Employees
Year Founded: 2020

What We Do

Climavision brings together the power of a proprietary, high-resolution weather radar and satellite network combined with advanced weather prediction modelling and decades of industry expertise to reduce existing coverage gaps and drastically improve forecast ability. Climavision’s revolutionary new approach to climate technology weather solutions is poised to help reduce the economic risks of climate change on companies, governments, and societies alike. Climavision is backed by The Rise Fund, the world’s largest global impact platform committed to achieving measurable, positive social and environmental outcomes alongside competitive financial returns. The company is headquartered in Louisville, KY, with research and development operations in Raleigh, NC. To learn more, visit www.Climavision.com

Similar Jobs

In-Office or Remote
5 Locations
35858 Employees
104K-143K Annually

SitusAMC Logo SitusAMC

Senior Site Reliability Engineer

Real Estate • Financial Services • PropTech
Remote
USA
4472 Employees
95K-135K Annually

Hilton Logo Hilton

Guest Services Agent F&B

Software • Hospitality
Remote
Georgia, USA
121228 Employees
In-Office or Remote
3 Locations
121228 Employees
110K-145K Annually

Similar Companies Hiring

Northslope Thumbnail
Artificial Intelligence • Information Technology • Software • Analytics • Consulting • Generative AI
London, GB
100 Employees
Scotch Thumbnail
Artificial Intelligence • eCommerce • Fintech • Payments • Retail • Software • Analytics
US
35 Employees
Milestone Systems Thumbnail
Artificial Intelligence • Security • Software • Analytics • Big Data Analytics
Lake Oswego, OR
1500 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account