Site Reliability Engineer

Posted 8 Days Ago
Be an Early Applicant
11 Locations
Remote
Senior level
Information Technology • Professional Services • Software
The Role
Lead incident response and root-cause analysis, automate repetitive tasks with scripting, manage SLOs and observability, run post-incident reviews, and plan capacity to ensure reliable cloud infrastructure.
Summary Generated by Built In

We are looking for a Site Reliability Engineer based in Latin America to work on a long-term project for one of our clients, a software company based in San Ramon, California.

Our client provides cloud solutions trusted by governments worldwide to accelerate their digital transformation, deliver vital services, and build stronger communities.

Responsibilities

  • Lead deep technical analysis during critical incidents, coordinating cross-functional teams to identify root causes, restore services quickly, and drive long-term reliability improvements.
  • Automation and Toil Reduction: Writing code and scripts (Python, Go, or Bash) to eliminate repetitive manual work and handle provisioning.
  • Incident Management: Responding to system alerts, troubleshooting production issues, joining on-call rotations, and restoring services quickly.
  • SLO and Error Budget Management: Establishing Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets to balance feature delivery speed with system stability.
  • Monitoring and Observability: Building dashboards and setting up tools (such as Prometheus, Grafana, or Datadog) to track system health and latency.
  • Post-Incident Reviews: Conducting blameless post-mortems to find root causes and prevent repeat failures.
  • Capacity Planning: Analyzing resource usage trends to forecast future infrastructure and scaling needs.

Requirements

  • Advanced Level of English.
  • 5+ years of experience working as a Site Reliability Engineer
  • 1+ years of experience within Microsoft Azure.
  • Strong experience with Python, Bash or Go for automation purposes.
  • Experience working with Monitoring and Observability tools such as Prometheus, Grafana, or Datadog.
  • Ability to keep a focused mind during high-severity production outages to lead teams effectively.
  • Proven experience explaining complex infrastructure failures to non-technical business stakeholders in clear, simple terms.
  • Strong prioritization instincts, with the ability to quickly determine whether to resolve a live issue manually or engineer a permanent automated solution.
  • Experience working with AI apps or agents.

Bonus Points

  • Bachelor’s Degree in Computer Science, Systems Engineering or related fields.
  • SaaS experience

What we offer

  • Long term positions
  • Compensation in USD
  • Paid time off
  • Cool clients and products
  • Work with great engineers

4tech

Skills Required

  • Advanced level of English
  • 5+ years working as a Site Reliability Engineer
  • 1+ years experience with Microsoft Azure
  • Strong experience with Python, Bash, or Go for automation
  • Experience with monitoring/observability tools such as Prometheus, Grafana, or Datadog
  • Ability to remain focused and lead teams during high-severity production outages
  • Proven ability to explain complex infrastructure failures to non-technical stakeholders
  • Strong prioritization instincts to choose manual fixes versus automated solutions
  • Experience working with AI apps or agents
  • Bachelor's degree in Computer Science, Systems Engineering or related field
  • SaaS experience
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
119 Employees
Year Founded: 2017

What We Do

Prediktive is a technology business partner and engineering services company that helps tech-enabled companies build and scale digital products. It provides software product development execution, engineering talent, and global remote support for startups, midsize businesses, and enterprises. Established in Silicon Valley, the company works across fields and time zones, connecting qualified professionals with client projects and helping organizations grow with confidence.

Similar Jobs

Kraken Digital Asset Exchange Logo Kraken Digital Asset Exchange

Site Reliability Engineer

Blockchain • Financial Services • Cryptocurrency • Web3
Remote
12 Locations
2900 Employees
Remote
11 Locations
359 Employees
84K-144K Annually

Thoughtworks Logo Thoughtworks

Senior Site Reliability Engineer

Cloud • Information Technology • Software • Analytics
Remote
Chile
7674 Employees

Alpaca Logo Alpaca

Senior Site Reliability Engineer

Fintech • Information Technology
Remote
13 Locations
132 Employees

Similar Companies Hiring

Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees
Kepler  Thumbnail
Fintech • Software
New York, New York
6 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account