We are looking for an Incident Manager based in Latin America to work on a long-term project for one of our clients, a software company based in San Ramon, California.
Our client provides cloud solutions trusted by governments across the globe to accelerate their digital transformation, deliver vital services, and build stronger communities.
Responsibilities
- Oversee and manage high-priority incidents for a SaaS service, ensuring timely detection, diagnosis, and resolution to minimize service disruptions and deliver a quality customer experience.
- Collaborate with engineering, operations, and product teams during incidents, providing technical leadership and ensuring clear communication across all stakeholders.
- Proactively monitor system health using tools like Datadog, responding to critical alerts and ensuring infrastructure performance aligns with aggressive service level agreements (SLAs).
- Conduct detailed post-incident reviews, driving root cause analysis and recommending preventive measures to enhance system resilience and minimize future incidents.
- Continuously refine alert thresholds, escalation processes, and monitoring tools (e.g., Datadog) to ensure effective detection of potential issues.
- Maintain up-to-date incident response procedures, runbooks, and system documentation to ensure team preparedness and streamlined troubleshooting during future incidents.
- Identify gaps in current incident management processes, collaborating with stakeholders to implement strategies for improving response times, communication, and issue resolution.
- Serve as the primary escalation point during on-call shifts, offering technical guidance and decision-making in high-pressure situations.
- Mentor junior engineers and support teams in incident response best practices, ensuring they are equipped to handle on-call duties effectively.
- Drive onboarding and training of incident team members, in collaboration with management, to ensure gaps in technical skills and domain expertise are filled where needed.
Requirements
- Advanced Level of English
- 7+ years of experience in SRE and incident management, managing and resolving high-priority incidents in a SaaS environment, ensuring system reliability and uptime.
- Hands-on experience with implementing strategies that optimize availability, performance, and reliability of large-scale SaaS infrastructure.
- Advanced knowledge in monitoring, alerting, and incident management tools such as Datadog and Atlassian, with experience setting up effective alert thresholds and processes.
- Ability to quickly diagnose and resolve system issues under pressure, ensuring minimal service disruption during incidents.
- Strong ability to lead incident response efforts, coordinate cross-functional teams, and ensure effective communication during high-pressure situations.
- Demonstrative experience with conducting thorough post-incident reviews and driving process improvements to enhance system reliability and incident response.
- Experience working with AI apps or agents.
Bonus Points
- Bachelor’s Degree in Computer Science, Systems Engineering or related fields
- In-depth Site Reliabilty Engineering experience with a SaaS, multi-tenant service delivery model in Microsoft Azure.
- Proficiency in scripting languages such as Python for automation and tooling.
- Advanced knowledge of Datadog for effective monitoring.
- Expertise in diverse tech stacks, adapting to evolving technologies.
- Strong working knowledge of Atlassian collaboration software (JIRA, Confluence, etc.), Slack.
- Familiarity with Amazon AWS.
What we offer
- Long term positions
- Compensation in USD
- Paid time off
- Cool clients and products
- Work with great engineers
4tech
Skills Required
- Advanced level of English
- 7+ years experience in SRE and incident management for SaaS environments
- Hands-on experience implementing strategies to optimize availability, performance, and reliability of large-scale SaaS infrastructure
- Advanced knowledge of monitoring, alerting, and incident management tools (Datadog, Atlassian) and setting alert thresholds and processes
- Ability to quickly diagnose and resolve system issues under pressure
- Proven ability to lead incident response, coordinate cross-functional teams, and communicate during high-pressure incidents
- Experience conducting post-incident reviews, root cause analysis, and driving process improvements
- Experience working with AI apps or agents
- Bachelor's degree in Computer Science, Systems Engineering or related field
- In-depth SRE experience with multi-tenant SaaS on Microsoft Azure
- Proficiency in scripting languages such as Python for automation and tooling
- Familiarity with Atlassian collaboration software (JIRA, Confluence) and Slack
- Familiarity with Amazon AWS
What We Do
Prediktive is a technology business partner and engineering services company that helps tech-enabled companies build and scale digital products. It provides software product development execution, engineering talent, and global remote support for startups, midsize businesses, and enterprises. Established in Silicon Valley, the company works across fields and time zones, connecting qualified professionals with client projects and helping organizations grow with confidence.







