SRE

Posted 2 Days Ago
Be an Early Applicant
27 Locations
Remote
53-53 Hourly
Senior level
HR Tech • Other • Professional Services
The Role
Build and operate production-grade observability infrastructure across GCP and Kubernetes. Design OpenTelemetry pipelines, improve metrics, traces, logs, dashboards, SLIs, SLOs, and burn-rate alerting. Manage infrastructure with Terraform, Helm, and Kubernetes; participate in on-call and incident response; conduct root-cause analyses; and collaborate with engineers on Go or Node.js instrumentation and reliability improvements.
Summary Generated by Built In
Senior Site Reliability Engineer — Observability

Contract: Approximately 3.5 months, through December 31, with strong potential to convert
Location: Remote within Europe
Working hours: Must consistently overlap with US Eastern Time mornings
Rate: Approximately $53 USD per hour, depending on experience
Start date: Mid-to-late September

About our Client

Our client is transforming the automotive care industry through an all-in-one shop management platform. Its software helps automotive businesses manage workflows, communicate with customers, process estimates and payments, and operate more efficiently across one or multiple locations.

We are looking for a Senior Site Reliability Engineer with deep observability expertise to strengthen the reliability and visibility of Shopmonkey’s production infrastructure.

The Role

This is a hands-on contract role for an experienced SRE who has built production-grade observability systems—not simply operated tools configured by others.

You will take ownership of OpenTelemetry-based telemetry pipelines within a Google Cloud and Kubernetes environment. You will help establish meaningful service-level objectives, improve alerting, and give engineering teams better visibility into the health and performance of production services.

The right person can work independently, make pragmatic infrastructure decisions, and collaborate effectively with application engineers during incidents and reliability improvements.

What You’ll Do
  • Design, build, and improve OpenTelemetry pipelines for metrics, traces, and logs.

  • Strengthen observability across services running on GCP and GKE.

  • Integrate telemetry with tools such as Google Cloud Monitoring, Cloud Trace, Managed Service for Prometheus, and Grafana.

  • Define and implement SLIs and SLOs based on customer and service reliability goals.

  • Build actionable burn-rate alerts that identify reliability risks without generating unnecessary noise.

  • Manage observability infrastructure through Terraform, Helm, and Kubernetes.

  • Improve dashboards, alerting, tracing, and service-level visibility.

  • Participate in on-call and incident-response activities.

  • Lead or contribute to root-cause analyses and follow incidents through to preventative improvements.

  • Collaborate with software engineers and open pull requests against Go or Node.js services when instrumentation or application-level changes are needed.

  • Document operational standards and help engineering teams adopt strong observability practices.

What We’re Looking For
  • 5+ years of professional experience in Site Reliability Engineering, Platform Engineering, Production Engineering, or Observability.

  • Proven, hands-on experience building an OpenTelemetry pipeline in a production GCP environment.

  • Strong experience with Google Kubernetes Engine (GKE).

  • Experience with Google Cloud Monitoring, Cloud Trace, and Managed Service for Prometheus.

  • Strong knowledge of Grafana, Prometheus, distributed tracing, metrics, logging, and telemetry architecture.

  • Practical experience defining SLIs and SLOs and implementing multi-window, multi-burn-rate alerting.

  • Strong hands-on experience managing infrastructure with Terraform and Helm.

  • Production-level Kubernetes experience, including deploying and operating infrastructure components.

  • Experience participating in on-call rotations, responding to incidents, and conducting thorough root-cause analyses.

  • Ability to read, troubleshoot, and submit production code changes in Go or Node.js.

  • Strong written and verbal English communication skills.

  • Availability to work remotely from Europe with reliable overlap during US Eastern Time mornings.

This Role May Be a Great Fit If You
  • Have personally designed observability infrastructure rather than only maintaining dashboards.

  • Understand how telemetry moves from application instrumentation through collectors, processors, exporters, and storage or visualization systems.

  • Can explain the tradeoffs behind sampling, cardinality, telemetry cost, alert thresholds, and signal quality.

  • Build alerts around user impact and service reliability rather than individual infrastructure symptoms.

  • Are comfortable moving between Kubernetes infrastructure, cloud services, observability tooling, and application code.

  • Work effectively in an autonomous, fast-moving environment.

Interview Process
  • Initial recorded screening focused on the core technical requirements.

  • Technical interview with members of the engineering team.

  • No take-home assessment.

  • Fast-moving process for a single hire.

This engagement is expected to run through December 31 and offers strong potential for conversion based on performance and ongoing business needs.

Skills Required

  • 5+ years of professional experience in Site Reliability Engineering, Platform Engineering, Production Engineering, or Observability
  • Hands-on experience building an OpenTelemetry pipeline in a production GCP environment
  • Strong experience with Google Kubernetes Engine
  • Experience with Google Cloud Monitoring, Cloud Trace, and Managed Service for Prometheus
  • Strong knowledge of Grafana, Prometheus, distributed tracing, metrics, logging, and telemetry architecture
  • Experience defining SLIs and SLOs and implementing multi-window, multi-burn-rate alerting
  • Strong hands-on experience managing infrastructure with Terraform and Helm
  • Production-level Kubernetes experience, including deploying and operating infrastructure components
  • Experience participating in on-call rotations, responding to incidents, and conducting thorough root-cause analyses
  • Ability to read, troubleshoot, and submit production code changes in Go or Node.js
  • Strong written and verbal English communication skills
  • Availability to work remotely from Europe with reliable overlap during US Eastern Time mornings
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Delray Beach, Florida
201 Employees
Year Founded: 2016

What We Do

G2i is a hiring community connecting remote developers with world-class engineering teams. Our unique approach combines rigorous technical assessments with a solid commitment to developer health, ensuring companies get skilled developers who are supported, valued, and ready to execute from day one. Our transparent vetting process includes in-depth, performance-ranked developer profiles, recorded technical interviews, and soft-skills assessments. Whether you're working on a short-term project or burning down a backlog, G2i connects you with a community of pre-vetted developers. Planning to hire ten or more engineers? We create a Custom Talent Pipeline, allowing for specific customizations in sourcing, assessment criteria, technical interview questions, and integration with your existing HR systems and processes. G2i partners with clients who support the developer health mission—matching developers with environments that improve their health, support recovery from burnout, and enable professional growth through restful work. Is your team overworked or understaffed? Contact us today to learn how G2i can help you. More information about our mission and commitment to developers and clients can be found at https://g2i.co or follow us on X @g2i_co

Similar Jobs

Dash0 Logo Dash0

Solutions Engineer

Artificial Intelligence
Remote
27 Locations

Elastic Logo Elastic

Senior Site Reliability Engineer

Cloud • Security • Software • Generative AI
Remote
Greece
3222 Employees
71K-94K Annually

ScorePlay Logo ScorePlay

Senior Platform Engineer

Artificial Intelligence • Digital Media • Software • Sports
In-Office or Remote
27 Locations
70 Employees

P2P.org Logo P2P.org

Site Reliability Engineer

Information Technology
Remote
30 Locations
179 Employees

Similar Companies Hiring

Compa Thumbnail
Artificial Intelligence • HR Tech • Software • Business Intelligence
Irvine, California
75 Employees
Rosendin Thumbnail
Other • Manufacturing
San Jose, CA
6219 Employees
OmniCable Thumbnail
Other
Houston, Texas
815 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account