Site Reliability Engineer

Posted 27 Days Ago
Scottsdale, AZ, USA
In-Office
Mid level
Information Technology • Professional Services • Software • Consulting
The Role
Responsible for operating and improving reliability for large-scale hybrid (on‑prem/cloud) applications: build automation, dashboards, observability (OTEL), maintain containerized workloads (GKE/RKE/AKE), support cloud migrations (GCP/Rancher), troubleshoot networking and databases, and work with GraphQL frameworks.
Summary Generated by Built In
We are seeking an experienced Site Reliability Engineer (SRE) to support large-scale, high-performance applications running in a hybrid environment (on-premises and cloud). The ideal candidate will have strong experience in cloud infrastructure, Kubernetes, observability, automation, and production operations.

Requirements
  • Service reliability/operation experience running large-scale, high-performance applications in a hybrid environment (on-prem and cloud).
  • Experience in writing automation scripts and building dashboards for Application Performance management to manage Transaction journeys.
  • Experience working with Programming languages such as Go, Python, Java, Rust etc.
  • Working knowledge on with one or more databases- Oracle, SQL Server, Redis, Clickhouse, postgres, Mongo or any  time-series databases
  • Experience in transitioning platforms to the cloud and Containerization – GCPand Rancher
  • Experience maintaining containerized app in GKE/RKE/AKE environments.
  • Experience Implementing Cloud observability using OTEL to enable real-time monitoring, distributed tracing and incident resolution.
  • Experience working with specific GraphQL Framework (Apollo, Prisma, Hasura etc...).
  • Experience using knowledge of networking protocols such as TCP/IP, HTTP, DNS, Load balancing and service mesh to troubleshoot issues in high pressure situations.Service reliability/operation experience running large-scale, high-performance applications in a hybrid environment (on-prem and cloud).
  • Experience in writing automation scripts and building dashboards for Application Performance management to manage Transaction journeys.
  • Experience working with Programming languages such as Go, Python, Java, Rust etc.
  • Working knowledge on with one or more databases- Oracle, SQL Server, Redis, Clickhouse, postgres, Mongo or any  time-series databases
  • Experience in transitioning platforms to the cloud and Containerization – GCPand Rancher
  • Experience maintaining containerized app in GKE/RKE/AKE environments.
  • Experience Implementing Cloud observability using OTEL to enable real-time monitoring, distributed tracing and incident resolution.
  • Experience working with specific GraphQL Framework (Apollo, Prisma, Hasura etc...).
  • Experience using knowledge of networking protocols such as TCP/IP, HTTP, DNS, Load balancing and service mesh to troubleshoot issues in high pressure situations.


Skills Required

  • Service reliability/operations experience running large-scale, high-performance applications in hybrid (on-prem and cloud) environments.
  • Experience writing automation scripts and building dashboards for Application Performance Management and transaction journeys.
  • Proficiency with programming languages such as Go, Python, Java, Rust.
  • Working knowledge of one or more databases: Oracle, SQL Server, Redis, ClickHouse, Postgres, MongoDB, or time-series databases.
  • Experience transitioning platforms to the cloud and containerization (GCP and Rancher).
  • Experience maintaining containerized applications in GKE, RKE, AKE environments.
  • Experience implementing cloud observability using OpenTelemetry (OTEL) for monitoring and distributed tracing.
  • Experience with GraphQL frameworks (Apollo, Prisma, Hasura).
  • Knowledge of networking protocols (TCP/IP, HTTP, DNS), load balancing, and service mesh for troubleshooting.
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
45 Employees
Year Founded: 2008

What We Do

Workiy is a global company with more than 20 years of experience that provides end-to-end digital solutions, consulting and implementation services to its clients, including digital solutions and staffing services.

Similar Jobs

CrowdStrike Logo CrowdStrike

Senior Engineer

Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity
Remote or Hybrid
USA
11000 Employees
140K-215K Annually

The Aerospace Corporation Logo The Aerospace Corporation

Senior Site Reliability Engineer

Aerospace • Artificial Intelligence • Cloud • Machine Learning • Software • Cybersecurity • Defense
Hybrid
Phoenix, AZ, USA
4600 Employees

Akamai Technologies Logo Akamai Technologies

Site Reliability Engineer

Cloud • Security • Software • Cybersecurity
In-Office or Remote
2 Locations
10285 Employees
95K-171K Annually

Akamai Technologies Logo Akamai Technologies

Site Reliability Engineer

Cloud • Security • Software • Cybersecurity
In-Office or Remote
2 Locations
10285 Employees
76K-136K Annually

Similar Companies Hiring

Kepler  Thumbnail
Artificial Intelligence • Fintech • Software
New York, New York
9 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Revel.io Thumbnail
Aerospace • Hardware • Robotics • Software
US
50 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account