Site Reliability Engineer

Posted 5 Days Ago
Be an Early Applicant
Barrington, RI, USA
In-Office
Entry level
Information Technology • Professional Services • Software • Consulting
The Role
Design, deploy, and maintain highly available production systems across cloud environments. Build automation and reliability tooling, manage Kubernetes and Docker workloads, implement Infrastructure as Code and CI/CD pipelines, and develop monitoring and observability using OpenTelemetry, Prometheus, and Grafana. Troubleshoot complex infrastructure and application issues, participate in incident response and root-cause analysis, improve system performance and capacity, and establish SRE reliability practices.
Summary Generated by Built In

We are seeking an experienced Site Reliability Engineer (SRE) to design, implement, and maintain highly available, scalable, secure, and reliable production systems. The ideal candidate will have strong software engineering skills combined with hands-on experience in cloud infrastructure, Kubernetes, automation, monitoring, and observability.

The SRE will work closely with development, infrastructure, and operations teams to improve system reliability, automate operational processes, and resolve complex production issues.



Requirements
  • Design, deploy, and maintain highly reliable and scalable production systems.
  • Develop automation and reliability tooling using Go, Python, Java, or Rust.
  • Manage and support cloud environments across AWS, Azure, or GCP.
  • Deploy and manage containerized applications using Docker and Kubernetes.
  • Implement and maintain monitoring, logging, metrics, and distributed tracing solutions.
  • Build and enhance observability using OpenTelemetry (OTel) and related technologies.
  • Troubleshoot complex infrastructure, application, and production issues.
  • Participate in incident response, root-cause analysis, and post-incident reviews.
  • Automate repetitive operational tasks and improve engineering efficiency.
  • Implement Infrastructure as Code using Terraform and/or Ansible.
  • Develop and maintain CI/CD pipelines for reliable and automated deployments.
  • Monitor system performance, availability, capacity, and overall reliability.
  • Identify reliability risks and implement proactive solutions.
  • Collaborate with software developers to improve application reliability and performance.
  • Establish and improve SRE best practices, operational procedures, and reliability standards.
Required Skills & Experience
  • Proven experience as a Site Reliability Engineer, Production Engineer, DevOps Engineer, or similar role.
  • Strong programming experience with at least one of:
    • Go/Golang
    • Python
    • Java
    • Rust
  • Hands-on experience with AWS, Azure, or GCP.
  • Strong experience with Kubernetes and Docker.
  • Strong Linux/Unix administration and troubleshooting skills.
  • Experience with OpenTelemetry and observability.
  • Knowledge of monitoring and visualization tools such as Prometheus and Grafana.
  • Experience with Terraform, Ansible, or similar Infrastructure-as-Code tools.
  • Strong understanding of CI/CD pipelines and DevOps practices.
  • Experience with production incident management and Root Cause Analysis (RCA).
  • Strong knowledge of automation, scripting, networking, and distributed systems.


Skills Required

  • Proven experience as a Site Reliability Engineer, Production Engineer, DevOps Engineer, or similar role
  • Strong programming experience with at least one of Go, Python, Java, or Rust
  • Hands-on experience with AWS, Azure, or GCP
  • Strong experience with Kubernetes and Docker
  • Strong Linux/Unix administration and troubleshooting skills
  • Experience with OpenTelemetry and observability
  • Knowledge of monitoring and visualization tools such as Prometheus and Grafana
  • Experience with Terraform, Ansible, or similar Infrastructure as Code tools
  • Strong understanding of CI/CD pipelines and DevOps practices
  • Experience with production incident management and root-cause analysis
  • Strong knowledge of automation, scripting, networking, and distributed systems
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
45 Employees
Year Founded: 2008

What We Do

Workiy is a global company with more than 20 years of experience that provides end-to-end digital solutions, consulting and implementation services to its clients, including digital solutions and staffing services.

Similar Jobs

CrowdStrike Logo CrowdStrike

Senior Engineer

Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity
Remote or Hybrid
USA
11000 Employees
140K-215K Annually

Akamai Technologies Logo Akamai Technologies

Site Reliability Engineer

Cloud • Security • Software • Cybersecurity
In-Office or Remote
2 Locations
10285 Employees
76K-136K Annually

Fabric Health Logo Fabric Health

Site Reliability Engineer

Artificial Intelligence • Healthtech • Software • Telehealth
In-Office or Remote
2 Locations
304 Employees
135K-160K Annually

Ping Identity Logo Ping Identity

Site Reliability Engineer

Cloud • Security • Software
Remote or Hybrid
USA
2300 Employees
136K-181K Annually

Similar Companies Hiring

Kepler  Thumbnail
Artificial Intelligence • Fintech • Software
New York, New York
9 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Revel.io Thumbnail
Aerospace • Hardware • Robotics • Software
US
50 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account