Sr Site Reliability Engineer

Posted Yesterday
Hiring Remotely in United States
Remote
140K-170K Annually
Senior level
eCommerce • Manufacturing
The Role
Manage Azure infrastructure and AKS clusters, build GitHub Actions CI/CD pipelines, and improve Grafana-based observability and incident response. Define SLIs, SLOs, and error budgets; maintain infrastructure as code with Pulumi; troubleshoot reliability issues; perform capacity planning and performance tuning; participate in on-call support; and document operational procedures. Collaborate with development teams to deliver scalable, reliable production systems.
Summary Generated by Built In

Who we are:
Through a partnership-based approach, Coterie helps insurance professionals unlock untapped revenue in the small commercial space. With an innovative quoting platform that delivers accurate pricing and bindable quotes in less than one minute, Coterie makes small business insurance effortless.  
We are on a mission to build and foster a world-class team to bring speed, simplicity, and service to commercial insurance. We value integrity, humility, passion, and intelligence. If you want to push yourself and reshape a $200B+ market, we’re excited to talk to you!


What will the Site Reliability Engineer do?

We're looking for a Senior Site Reliability Engineer who's passionate about building and maintaining reliable, scalable infrastructure and who thrives on making systems better every day. In this role, you'll join our SRE team to help keep our platforms running smoothly, improve our observability and incident response capabilities, and partner with development teams to deliver infrastructure that supports high-quality, reliable software.

You'll play a key role in managing our cloud infrastructure, strengthening our CI/CD pipelines, and helping us get the most out of our monitoring and alerting tools, particularly Grafana. This is a great opportunity for a mid-level engineer ready to take ownership of meaningful infrastructure challenges.
Key Responsibilities:

  • Manage and maintain cloud infrastructure on Azure, including Azure Kubernetes Service (AKS) clusters and supporting resources
  • Build, improve, and maintain CI/CD pipelines using GitHub Actions to support reliable and repeatable deployments
  • Own and enhance our Grafana implementation; designing dashboards, configuring alerts, and supporting incident management workflows
  • Monitor system health, triage incidents, and drive root cause analysis to prevent recurrence
  • Collaborate with development teams to define and track SLIs, SLOs, and error budgets that align with business goals
  • Contribute to infrastructure-as-code practices using Pulumi
  • Identify and resolve reliability risks through capacity planning, performance tuning, and proactive system improvements
  • Participate in an on-call rotation to support production systems and respond to incidents
  • Document runbooks, operational procedures, and architectural decisions to support team knowledge sharing

What we are looking for:  

  • 5+ years of experience in a Site Reliability Engineering, DevOps, or Infrastructure role
  • 3+ years experience working with infrastructure as code
  • 2+ years of experience architecting CI/CD pipelines and cloud-based infrastructure
  • Strong hands-on experience with:
  • Azure Cloud services and resource management
  • Kubernetes and AKS administration, including deployments, networking, and troubleshooting
  • GitHub Actions for CI/CD pipeline development and maintenance
  • 3+ experience with Grafana or similar tooling, including dashboard creation, alerting configuration, and incident management
  • Hands-on experience with Prometheus, Loki, or other observability tools in the Grafana ecosystem
  • Proficiency in at least one scripting or programming language such as Python or Bash
  • Understanding of networking fundamentals, DNS, load balancing, and container orchestration concepts
  • Strong analytical and communication skills; able to diagnose complex system issues and clearly communicate findings
  • Demonstrated ability to collaborate across teams and contribute to a culture of reliability
  • Experience working in an agile environment with modern DevOps practices

What will make you stand out:

  • Experience working at a startup or in a fast-paced, cross-functional environment
  • Familiarity with the insurance industry or other regulated sectors
  • Familiarity with service mesh technologies (e.g., Istio)


Our interview process:

Our hiring process generally consists of 4 phases. The goal is to provide an opportunity for us to learn more about our candidates while allowing them to get to know us as well!

  • Phase 1: Qualified candidates will first meet with a member of our People Operations team for a phone interview.  This discussion is a high-level conversation to understand more about your background and interests and for us to share more about Coterie and the position.
  • Phase 2: Selected candidates will be invited to meet with our Hiring Manager for a 2nd interview via Teams video. This interview is designed to be more detail oriented and allows you to learn more about the role and expected to be 30 minutes in length.
  • Phase 3: Top candidates will be invited to participate in an experiential assessment phase, which includes a take-home coding exercise. Candidates who successfully complete the project will be invited to a 1-hour technical interview with our hiring manager and members of the engineering team.
  • Phase 4: Final candidates will receive an invite to our final interview series. This series will include 1:1 interview with senior leadership. The final series is roughly 1 hour total.


What's in it for you:

Coterie has excellent benefits for all full-time employees. We offer the following:

  • 100% remote
  • Health insurance through Aetna (we pay 100% of premiums)
  • Dental and vision insurance through Guardian (we pay 100% of premiums)
  • Basic life insurance (we pay 100% of premiums)
  • Access to flexible spending account (FSA) or health savings account (HSA) (for those using HSA eligible plans)
  • 401K plan (up 4% match with immediate vest). Must be 21 years of age or older to participate
  • Flexible PTO policy offering employees up to 4 weeks of PTO in their first 12 months. Thereafter, PTO usage aligns with company standards and typically does not exceed 5 weeks per calendar year.
  • 12 company-paid holidays each year
  • Continuing education annual stipend
  • Annual salary estimated between $140,000-$170,000 based on national data. Candidates who meet all the minimum requirements and possess additional relevant experience, as outlined in the job description, may be considered for a salary above the midpoint of the above range. Salary is based on internal equity; internal salary ranges; market data/ranges; applicant’s skills; prior relevant experience; degrees or certifications, etc. 

Work Authorization:
At this time, Coterie Insurance is unable to consider candidates who require current or future visa sponsorship. Applicants must have authorization to work in the United States without the need for sponsorship now or in the future. Falsification of an application, including work authorization status, is immediate grounds for dismissal from consideration.

Skills Required

  • 5+ years of experience in Site Reliability Engineering, DevOps, or infrastructure
  • 3+ years of experience working with infrastructure as code
  • 2+ years of experience architecting CI/CD pipelines and cloud-based infrastructure
  • Strong hands-on experience with Azure cloud services and resource management
  • Strong hands-on experience administering Kubernetes and AKS, including deployments, networking, and troubleshooting
  • Strong hands-on experience developing and maintaining GitHub Actions CI/CD pipelines
  • 3+ years of experience with Grafana or similar tools, including dashboard creation, alerting, and incident management
  • Hands-on experience with Prometheus, Loki, or similar Grafana ecosystem observability tools
  • Proficiency in at least one scripting or programming language, such as Python or Bash
  • Understanding of networking fundamentals, DNS, load balancing, and container orchestration
  • Strong analytical and communication skills, including the ability to diagnose complex system issues
  • Ability to collaborate across teams and contribute to a culture of reliability
  • Experience working in an agile environment with modern DevOps practices
  • Experience working at a startup or in a fast-paced, cross-functional environment
  • Familiarity with the insurance industry or other regulated sectors
  • Familiarity with service mesh technologies such as Istio

Coterie Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Coterie and has not been reviewed or approved by Coterie.

  • Parental & Family Support Parental leave around 16 weeks and access to Maven family-planning/parenting support are highlighted in role and benefits descriptions. This points to strong caregiver support within the package.
  • Healthcare Strength Benefits materials highlight medical, dental, and vision coverage as core offerings. Coverage is consistently listed across multiple role descriptions for the baby-care brand.
  • Leave & Time Off Breadth Unlimited PTO is described across job and benefits pages. This signals broad flexibility around time away from work.

Coterie Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Sevenoaks
90 Employees
Year Founded: 2018

What We Do

We make the most absorbent, soft-as-cashmere diaper, the ergonomically designed and versatile pant, and 100% plant-based wipes, with only the safest materials

Similar Jobs

Remote or Hybrid
United States
1750 Employees

Zocdoc Logo Zocdoc

Senior Site Reliability Engineer

Healthtech • Information Technology • Software • Telehealth
Easy Apply
Remote or Hybrid
USA
900 Employees
180K-220K Annually
Remote
United States
350 Employees
180K-220K Annually

GitLab Logo GitLab

Site Reliability Engineer

Cloud • Security • Software • Cybersecurity • Automation
Easy Apply
Remote
United States
2500 Employees

Similar Companies Hiring

Rosendin Thumbnail
Other • Manufacturing
San Jose, CA
6219 Employees
Amalgamated Sugar Thumbnail
Food • Greentech • Agriculture • Industrial • Manufacturing
Boise, Idaho
768 Employees
Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account