Senior Site Reliability Engineer

Posted Yesterday
Hiring Remotely in United States
Remote
Senior level
Big Data • Software
Empowering Data Driven Contracting through best-in-class technology solutions, purpose-built for the MEP industry.
The Role
Owns production reliability for a B2B SaaS platform, including SLOs, error budgets, observability, alerting, incident response, recovery exercises, performance and capacity engineering, load testing, disaster recovery, and production readiness. The role writes automation and production code, leads incidents, coaches engineering teams, and improves reliability across Kubernetes, Azure, databases, messaging, and service infrastructure.
Summary Generated by Built In

Stratus, deriving from the Latin term meaning 'layer', offers an advanced set of MEP specific solutions that seamlessly layer across a contractor's entire workflow from design to fabrication to installation. Our team of seasoned industry experts, skilled technology leaders, innovators, and entrepreneurs understands that fabrication does not occur in isolation, and increasingly, it may not happen within your own fabrication shop. Through close relationships with our customers—who include some of the most innovative and largest MEP contractors—we have developed a suite of Stratus tools to digitize, automate, and optimize piping, plumbing, sheet metal, and electrical contracting. Stratus provides the software layer an MEP Contractor needs to optimize profits with true "Data Driven Contracting."

GENERAL DESCRIPTION:

The Senior Site Reliability Engineer is accountable for how Stratus behaves in production. Reporting to the Senior Director, Platform Engineering, this role brings genuine SRE discipline to a platform that MEP contractors run their fabrication shops on — where downtime does not mean a degraded experience, it means work stops on a job site. Stratus is a ~50-person, primarily remote, Series B software company growing quickly.

Unrelenting Reliability is one of our company values, and this is the role that operationalizes it. You will define what reliable means in numbers, instrument the system so we can see it, and close the loop from production signal back into engineering priority. This is an engineering role with a production mandate: you will write code, tune queries, build alerting, run load tests, and lead incidents — and you will be measured on customer-visible availability and latency, not on tickets closed.

The load-bearing problems on your plate are: (1) establishing service level objectives and an error budget the whole engineering organization operates against, with the measurement infrastructure to back them; (2) building the detection and response capability — high-signal alerting, clear on-call and escalation paths, per-system incident ownership, and well-exercised recovery paths — so that production problems are caught and resolved fast; and (3) building the performance and capacity engineering practice for our data and messaging layers, with headroom measured rather than assumed.

You will work across every engineering pod, with our platform and security functions, and with the customer-facing teams who see reliability problems first. The right candidate is comfortable being the person who says a number out loud and defends it, and is drawn to a place where the reliability practice is being built rather than maintained.

KEY RESPONSIBILITIES:

  • Define and own service level indicators, objectives, and error budgets for customer-facing services; build the measurement pipeline that makes them trustworthy and publish availability and latency against target on a regular cadence.
  • Build and own production observability: instrumentation standards, dashboards, and actionable alerting on our stack — AKS on Azure, Prometheus, Loki, Tempo, and Grafana, with Istio as the service mesh and Flux for GitOps — extending coverage across every environment.
  • Establish and run incident response practice — on-call rotation and paging paths, severity definitions, per-system incident ownership, escalation into and out of customer support, and blameless post-mortems in incident.io.
  • Set the recovery standard: define what rollback and recovery must demonstrate, and verify through regular exercise that every service meets it.
  • Lead performance and capacity engineering: query and index tuning, connection-pool sizing, capacity modeling against real load, and designing for graceful degradation under partial failure.
  • Work with the existing team that runs our k6 load, stress, spike, and soak testing against a production-like environment; help set per-endpoint latency thresholds tied to our SLOs and make the results a gate on delivery.
  • Drive production readiness: review new services and significant changes against readiness criteria covering instrumentation, alerting, failure modes, resource limits, and rollback before they ship.
  • Close the loop on remediation: track post-mortem action items through to verified production change.
  • Contribute to business continuity and disaster recovery planning, including backup and restore validation, failover design, and recovery objectives.
  • Partner with engineering pods to push reliability ownership outward — coach teams on instrumenting and operating their own services.
  • Write code: tooling, automation, instrumentation libraries, and fixes in the production codebase.

QUALIFICATIONS:

  • 6+ years of professional engineering experience, with at least 3+ years in a dedicated SRE or production engineering role at a B2B SaaS company.
  • Demonstrated ownership of SLOs and error budgets in production — you have defined them, measured them, argued about them with product leadership, and changed engineering behavior with them.
  • Deep, practical observability skills with Prometheus, Loki, Tempo, and Grafana or comparable: you have instrumented real systems and built alerting that pages on customer impact rather than on CPU.
  • Hands-on incident command experience at meaningful severity, including building on-call and escalation practice from the ground up.
  • Strong database performance skills — query profiling, index design, connection pooling, and diagnosing saturation under load. MongoDB experience strongly preferred; comparable document or relational depth acceptable.
  • Production Kubernetes experience (AKS preferred) sufficient to debug a live problem — pod scheduling, resource limits, networking, and service mesh behavior (Istio preferred).
  • Solid coding ability in at least one general-purpose language (Go, Python, C#, or TypeScript); willingness to work in a C#/.NET codebase.
  • Experience with load and performance testing tooling (k6, JMeter, Gatling, or comparable) and with turning results into engineering priority.
  • Production experience on Azure or AWS, with real understanding of the failure modes of managed services.
  • Fluency with AI-assisted engineering tooling and a track record of designing AI-leveraged workflows for your team — this is a graded expectation at every level at Stratus.
  • Excellent written communication: you write post-mortems, runbooks, and reliability reports that executives and engineers both read and act on.
  • Judgment and steadiness under pressure, and the credibility to tell engineering and product leadership something they do not want to hear.

NICE TO HAVE:

  • Experience establishing or maturing an SRE practice — defining the discipline, not inheriting it.
  • Experience with Sentry or comparable application error-monitoring platforms.
  • Experience operating event-driven and real-time systems — message brokers (Azure Service Bus, Kafka) and websocket or push layers (SignalR or comparable).
  • Experience operating MongoDB Atlas at production scale, including replica set topology and Atlas performance tooling.
  • Experience with durable workflow orchestration (Temporal or comparable).
  • Background in multi-region or multi-zone architecture and DR design against stated recovery objectives.
  • Experience with incident.io, PagerDuty, or comparable incident management platforms.
  • Familiarity with DORA metrics and with reliability work inside SOC 2 or NIST 800-171 scope.
  • Experience with legacy monolith reliability — improving the operational behavior of an ASP.NET or comparable application you cannot rewrite.
  • Domain interest in MEP, BIM, AEC, or construction technology.
  • Prior experience in a Series B / growth-stage company navigating the transition from product-market fit to scale.


BENEFITS:

  • Comprehensive and competitive health benefits plan
  • Matching 401k contributions
  • 20 days annual PTO
  • Primarily remote work with occasional annual team onsites.


E-VERIFY STATEMENT 
Stratus participates in E-Verify. After you join the team, we'll verify your eligibility to work in the U.S. by submitting information from your Form I-9 to the Social Security Administration and, if needed, the Department of Homeland Security. This process happens post-hire only - we never use E-Verify to pre-screen applicants. 
E-Verify Notice 
Right to Work Notice 

Skills Required

  • 6+ years of professional engineering experience
  • At least 3+ years in a dedicated SRE or production engineering role at a B2B SaaS company
  • Demonstrated ownership of production SLOs and error budgets
  • Practical observability experience with Prometheus, Loki, Tempo, Grafana, or comparable tools
  • Hands-on incident command experience and experience building on-call and escalation practices
  • Strong database performance skills, including query profiling, index design, connection pooling, and saturation diagnosis
  • Production Kubernetes experience; AKS preferred
  • Solid coding ability in Go, Python, C#, or TypeScript, with willingness to work in a C#/.NET codebase
  • Experience with load and performance testing tools such as k6, JMeter, or Gatling
  • Production experience with Azure or AWS and understanding of managed-service failure modes
  • Fluency with AI-assisted engineering tooling and experience designing AI-leveraged workflows
  • Excellent written communication for post-mortems, runbooks, and reliability reports
  • Judgment, steadiness under pressure, and credibility with engineering and product leadership
  • MongoDB experience
  • Experience establishing or maturing an SRE practice
  • Experience with Sentry or comparable application error-monitoring platforms
  • Experience operating event-driven and real-time systems, including message brokers and websocket or push layers
  • Experience operating MongoDB Atlas at production scale
  • Experience with durable workflow orchestration such as Temporal
  • Multi-region or multi-zone architecture and disaster recovery design experience
  • Experience with incident.io, PagerDuty, or comparable incident management platforms
  • Familiarity with DORA metrics and reliability work within SOC 2 or NIST 800-171 scope
  • Experience improving legacy ASP.NET or comparable monolith reliability
  • Domain interest in MEP, BIM, AEC, or construction technology
  • Experience at a Series B or growth-stage company
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Mentor, OH
87 Employees
Year Founded: 2010

What We Do

The Leading Tool for MEP Fabrication Workflows. Stratus is a cloud-based software platform that revolutionizes MEP fabrication and construction management by seamlessly integrating CAD software like Revit and AutoCAD with manufacturing tools to reduce errors and boost efficiency. By leveraging digital models for precision fabrication and enabling real-time collaboration, Stratus enhances communication among teams and ensures accurate project tracking. This platform empowers specialty contractors to streamline their workflows from BIM to installation, optimizing planning, resource allocation, and project execution. Stratus... - Optimizes Project Management and Decision Making with a modern, cloud-based platform that connects your VDC, Shop and Field teams - Eliminates Paper and PDF Workflows in the shop and in the field with direct access to the model - Reduces Waste with smart cut lists, material management and smart labels - Automates Fabrication with direct output to various cutting equipment - Eliminates Manual Conversion Steps with direct integration from CAD software like Revit and AutoCAD to manufacturing tools, automating the fabrication process and reducing errors -Leverages Historical Data to forecast future project timelines and productivity, enabling better planning and resource allocation -Assigns custom tracking statuses to work packages, offering unparalleled visibility into project progress and logistics

Similar Jobs

In-Office or Remote
12 Locations
125 Employees
75K-195K Annually

Zocdoc Logo Zocdoc

Senior Site Reliability Engineer

Healthtech • Information Technology • Software • Telehealth
Easy Apply
Remote or Hybrid
USA
900 Employees
180K-220K Annually
Remote
United States
350 Employees
180K-220K Annually
Remote
USA
164 Employees

Similar Companies Hiring

Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Revel Thumbnail
Aerospace • Hardware • Robotics • Software
Marina Del Rey, California
60 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software
New York, New York
30 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account