We are looking for a Site Reliability Engineer based in Latin America to work on a long-term project for one of our clients, a software company based in San Ramon, California.
Our client provides cloud solutions trusted by governments worldwide to accelerate their digital transformation, deliver vital services, and build stronger communities.
Responsibilities
- Lead deep technical analysis during critical incidents, coordinating cross-functional teams to identify root causes, restore services quickly, and drive long-term reliability improvements.
- Automation and Toil Reduction: Writing code and scripts (Python, Go, or Bash) to eliminate repetitive manual work and handle provisioning.
- Incident Management: Responding to system alerts, troubleshooting production issues, joining on-call rotations, and restoring services quickly.
- SLO and Error Budget Management: Establishing Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets to balance feature delivery speed with system stability.
- Monitoring and Observability: Building dashboards and setting up tools (such as Prometheus, Grafana, or Datadog) to track system health and latency.
- Post-Incident Reviews: Conducting blameless post-mortems to find root causes and prevent repeat failures.
- Capacity Planning: Analyzing resource usage trends to forecast future infrastructure and scaling needs.
Requirements
- Advanced Level of English.
- 5+ years of experience working as a Site Reliability Engineer
- 1+ years of experience within Microsoft Azure.
- Strong experience with Python, Bash or Go for automation purposes.
- Experience working with Monitoring and Observability tools such as Prometheus, Grafana, or Datadog.
- Ability to keep a focused mind during high-severity production outages to lead teams effectively.
- Proven experience explaining complex infrastructure failures to non-technical business stakeholders in clear, simple terms.
- Strong prioritization instincts, with the ability to quickly determine whether to resolve a live issue manually or engineer a permanent automated solution.
- Experience working with AI apps or agents.
Bonus Points
- Bachelor’s Degree in Computer Science, Systems Engineering or related fields.
- SaaS experience
What we offer
- Long term positions
- Compensation in USD
- Paid time off
- Cool clients and products
- Work with great engineers
4tech
Skills Required
- Advanced level of English
- 5+ years working as a Site Reliability Engineer
- 1+ years experience with Microsoft Azure
- Strong experience with Python, Bash, or Go for automation
- Experience with monitoring/observability tools such as Prometheus, Grafana, or Datadog
- Ability to remain focused and lead teams during high-severity production outages
- Proven ability to explain complex infrastructure failures to non-technical stakeholders
- Strong prioritization instincts to choose manual fixes versus automated solutions
- Experience working with AI apps or agents
- Bachelor's degree in Computer Science, Systems Engineering or related field
- SaaS experience
What We Do
Prediktive is a technology business partner and engineering services company that helps tech-enabled companies build and scale digital products. It provides software product development execution, engineering talent, and global remote support for startups, midsize businesses, and enterprises. Established in Silicon Valley, the company works across fields and time zones, connecting qualified professionals with client projects and helping organizations grow with confidence.









