The Role
Responsible for operating and improving reliability for large-scale hybrid (on‑prem/cloud) applications: build automation, dashboards, observability (OTEL), maintain containerized workloads (GKE/RKE/AKE), support cloud migrations (GCP/Rancher), troubleshoot networking and databases, and work with GraphQL frameworks.
Summary Generated by Built In
We are seeking an experienced Site Reliability Engineer (SRE) to support large-scale, high-performance applications running in a hybrid environment (on-premises and cloud). The ideal candidate will have strong experience in cloud infrastructure, Kubernetes, observability, automation, and production operations.
Requirements
- Service reliability/operation experience running large-scale, high-performance applications in a hybrid environment (on-prem and cloud).
- Experience in writing automation scripts and building dashboards for Application Performance management to manage Transaction journeys.
- Experience working with Programming languages such as Go, Python, Java, Rust etc.
- Working knowledge on with one or more databases- Oracle, SQL Server, Redis, Clickhouse, postgres, Mongo or any time-series databases
- Experience in transitioning platforms to the cloud and Containerization – GCPand Rancher
- Experience maintaining containerized app in GKE/RKE/AKE environments.
- Experience Implementing Cloud observability using OTEL to enable real-time monitoring, distributed tracing and incident resolution.
- Experience working with specific GraphQL Framework (Apollo, Prisma, Hasura etc...).
- Experience using knowledge of networking protocols such as TCP/IP, HTTP, DNS, Load balancing and service mesh to troubleshoot issues in high pressure situations.Service reliability/operation experience running large-scale, high-performance applications in a hybrid environment (on-prem and cloud).
- Experience in writing automation scripts and building dashboards for Application Performance management to manage Transaction journeys.
- Experience working with Programming languages such as Go, Python, Java, Rust etc.
- Working knowledge on with one or more databases- Oracle, SQL Server, Redis, Clickhouse, postgres, Mongo or any time-series databases
- Experience in transitioning platforms to the cloud and Containerization – GCPand Rancher
- Experience maintaining containerized app in GKE/RKE/AKE environments.
- Experience Implementing Cloud observability using OTEL to enable real-time monitoring, distributed tracing and incident resolution.
- Experience working with specific GraphQL Framework (Apollo, Prisma, Hasura etc...).
- Experience using knowledge of networking protocols such as TCP/IP, HTTP, DNS, Load balancing and service mesh to troubleshoot issues in high pressure situations.
Skills Required
- Service reliability/operations experience running large-scale, high-performance applications in hybrid (on-prem and cloud) environments.
- Experience writing automation scripts and building dashboards for Application Performance Management and transaction journeys.
- Proficiency with programming languages such as Go, Python, Java, Rust.
- Working knowledge of one or more databases: Oracle, SQL Server, Redis, ClickHouse, Postgres, MongoDB, or time-series databases.
- Experience transitioning platforms to the cloud and containerization (GCP and Rancher).
- Experience maintaining containerized applications in GKE, RKE, AKE environments.
- Experience implementing cloud observability using OpenTelemetry (OTEL) for monitoring and distributed tracing.
- Experience with GraphQL frameworks (Apollo, Prisma, Hasura).
- Knowledge of networking protocols (TCP/IP, HTTP, DNS), load balancing, and service mesh for troubleshooting.
Am I A Good Fit?
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.
Success! Refresh the page to see how your skills align with this role.
The Company
What We Do
Workiy is a global company with more than 20 years of experience that provides end-to-end digital solutions, consulting and implementation services to its clients, including digital solutions and staffing services.








