Lead SRE Engineer

Posted 9 Days Ago
Be an Early Applicant
2 Locations
Remote or Hybrid
Senior level
Information Technology • Logistics • Professional Services • Consulting
The Role
Lead and scale SRE practices for Salesforce and Azure environments, ensuring high availability, observability, incident management, automation (IaC, CI/CD), integrations, compliance (SOX), and mentoring distributed teams to improve reliability and operational excellence.
Summary Generated by Built In

Hybrid: 3 days on-site: Ave. Eugenio Garza Lagüera, San Pedro Garza García, Nuevo León, Mexico.

Role Overview

We are seeking a highly experienced Senior SRE Lead to lead reliability engineering and observability initiatives for critical platforms supporting GM Financials’ ecosystem, with a primary focus on Salesforce and Microsoft Azure environments. This role will be responsible for establishing and scaling SRE practices, driving operational excellence, and ensuring high availability, performance, and resilience of business-critical applications. The ideal candidate will bring deep expertise in cloud-native architecture, observability frameworks, and enterprise-scale production support, along with strong leadership capabilities in a global delivery model.

Key Responsibilities - SRE Leadership & Strategy

  • Lead the SRE function for Salesforce and Azure platforms, defining the roadmap and maturity model

  • Establish and drive SRE best practices, including SLIs, SLOs, and error budgets

  • Build and mentor a high-performing SRE team across onshore and offshore locations

  • Collaborate with GM Financial stakeholders, product teams, and engineering leadership Platform Reliability (Salesforce & Azure)

  • Ensure high availability, performance, and scalability of Salesforce applications and Azure-hosted services

  • Lead major incident management (P1/P2), including triage, stakeholder communication, and resolution

  • Drive root cause analysis (RCA) and implement preventive measures

  • Manage production stability across integrations between Salesforce and Azure services Observability & Monitoring

  • Design and implement end-to-end observability across Salesforce and Azure ecosystems

  • Establish unified monitoring across logs, metrics, and traces

  • Implement and optimize tools such as Azure Monitor, Application Insights, Splunk, Datadog, or similar

  • Define dashboards, alerting strategies, and actionable insights for proactive issue detection Automation & DevOps

  • Drive automation across incident response, remediation, and operational workflows

  • Implement Infrastructure as Code (IaC) practices using tools such as Terraform, ARM templates, or similar

  • Enhance CI/CD pipelines for Salesforce and Azure deployments

  • Enable self-healing systems and reduce manual intervention Cloud & Integration Engineering

  • Optimize Azure infrastructure for performance, resilience, and cost efficiency

  • Support Salesforce platform stability, including integrations, APIs, and middleware components

  • Work closely with integration teams to ensure reliable data flows and system interactions Governance, Risk & Compliance

  • Ensure adherence to GM Financial’s security, compliance, and regulatory requirements (including SOX)

  • Maintain audit-ready processes, documentation, and operational controls

  • Participate in governance forums, audits, and compliance reviews

Required Skills & Qualifications Technical Expertise

  • Strong experience in SRE, DevOps, or Production Engineering roles

  • Hands-on experience with Microsoft Azure (mandatory)

  • Experience supporting Salesforce platforms (Sales Cloud, Service Cloud, integrations)

  • Expertise in observability tools (Azure Monitor, Application Insights, Splunk, Datadog, etc.)

  • Strong scripting/programming skills (Python, PowerShell, or similar)

  • Experience with CI/CD tools (Azure DevOps, GitHub Actions, Jenkins)

  • Familiarity with Infrastructure as Code (Terraform, ARM, Bicep) Operational Excellence

  • Proven experience in managing high-availability production environments

  • Strong understanding of incident management, RCA, and problem management

  • Experience defining and managing SLIs, SLOs, and error budgets Leadership & Stakeholder Management

  • Experience leading distributed/global teams

  • Strong communication and stakeholder management skills

  • Ability to operate in a fast-paced, high-impact environment

  • Strong decision-making and problem-solving capabilities

Skills Required

  • Strong experience in SRE, DevOps, or Production Engineering roles
  • Hands-on experience with Microsoft Azure
  • Experience supporting Salesforce platforms (Sales Cloud, Service Cloud, integrations)
  • Expertise in observability tools (Azure Monitor, Application Insights, Splunk, Datadog, etc.)
  • Strong scripting/programming skills (Python, PowerShell, or similar)
  • Experience with CI/CD tools (Azure DevOps, GitHub Actions, Jenkins)
  • Familiarity with Infrastructure as Code (Terraform, ARM, Bicep)
  • Proven experience managing high-availability production environments
  • Experience defining and managing SLIs, SLOs, and error budgets
  • Experience leading distributed/global teams
  • Strong incident management, RCA, and problem management experience
  • Knowledge of security, compliance, and regulatory requirements including SOX
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
21 Employees
Year Founded: 2016

What We Do

VALCE Talent Solutions is a recruitment agency and consulting firm specializing in IT talent acquisition and nearshoring, primarily in Mexico. They design customized solutions in IT talent and process optimization to help businesses scale intelligently and profitably. Their expertise focuses on specialized industries including Information Technology, Operations Management, and Supply Chain, providing end-to-end talent offerings to attract and retain top professional profiles.

Similar Jobs

Capital One Logo Capital One

Site Reliability Engineer

Fintech • Machine Learning • Payments • Software • Financial Services
Remote or Hybrid
Mexico City, Ciudad De México, MEX
55000 Employees

Mastercard Logo Mastercard

Site Reliability Engineer

Blockchain • Fintech • Payments • Consulting • Cryptocurrency • Cybersecurity • Quantum Computing
Remote or Hybrid
Mexico City, Ciudad De México, MEX
38800 Employees

Photon Logo Photon

Site Reliability Engineer

Agency • Information Technology
In-Office or Remote
2 Locations
5017 Employees

Similar Companies Hiring

Amplify Platform Thumbnail
Fintech • Financial Services • Consulting • Cloud • Business Intelligence • Big Data Analytics
Scottsdale, AZ
62 Employees
Standard Template Labs Thumbnail
Artificial Intelligence • Information Technology • Software
New York, NY
25 Employees
Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account