SRE Lead | Guadalajara,Mexico

Reposted 3 Days Ago
Be an Early Applicant
2 Locations
In-Office or Remote
Senior level
Agency • Information Technology
The Role
Lead SRE for customer-facing microservices: define and own SLIs/SLOs, drive reliability via error budgets, build observability, automate toil reduction, lead incident response and postmortems, mentor SREs, and partner with engineering/product teams to embed operability and resilience.
Summary Generated by Built In

Job Description

Site Reliability Engineers are responsible for ensuring the availability, reliability, scalability, and performance of the firm’s most critical, customer-facing microservices that power all eCommerce channels. This role applies Google-inspired SRE principles to balance feature velocity and system reliability using Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets.
The role combines software engineering, cloud engineering, automation, and production operations, with a strong emphasis on building systems that are observable, resilient, and operable by default.
Primary Responsibilities
Define, implement, and own SLIs, SLOs, and error budgets for critical microservices in collaboration with product and engineering teams.
Use error budgets to influence release decisions, prioritize reliability work, and manage operational risk.
Design and maintain observability platforms including metrics, logs, traces, and real-time telemetry.
Track, manage, and reduce operational toil by converting repetitive operational work into Jira stories and epics with clear ownership and measurable outcomes.
Design, implement, and validate resiliency mechanisms such as graceful degradation, redundancy, automated failover, and disaster recovery.
Lead incident response, act as an escalation point for high-severity incidents, and drive blameless postmortems.
Capture incident action items and reliability improvements in Jira, ensuring closure, accountability, and continuous improvement.
Partner with scrum teams to improve reliability through release readiness reviews, production change validation, and testing strategies.
Perform deep root cause analysis, debugging, and performance tuning across distributed systems.
Promote shift-left reliability by embedding operability, monitoring, and failure testing early in the SDLC.
Drive continuous improvement through automation, self-healing systems, chaos engineering, and capacity planning.
Maintain runbooks, playbooks, and knowledge repositories, linking documentation to Jira tasks to reduce MTTR.
Provide technical leadership and mentoring to junior SREs and engineers.
Collaborate with global, distributed teams, leveraging Jira for transparent planning, dependency tracking, and execution.
Core Competencies & Accomplishments
6+ years of experience in SRE, software engineering, or production operations supporting large-scale eCommerce platforms.
Hands-on experience with Java/J2EE-based distributed systems. React experience is a plus.
Proven ability to design and operate systems using SLO-driven reliability models.
Experience defining and measuring SLIs (availability, latency, error rates, throughput, saturation).
Good understanding with NoSQL technologies and RDBMS. Should be able to write queries to fetch results from database.
Experience deploying and operating services on cloud platforms (AWS, Azure, or Google Cloud).
Expertise with observability, APM, and caching tools (Dynatrace, Splunk, ELK, Akamai, QuantumMetric/Tealeaf, etc.).
Strong experience using Jira for backlog management, incident follow-ups, toil reduction tracking, and cross-team coordination.
Ability to independently own services and drive reliability initiatives end-to-end.
Strong communication skills and ability to influence engineering and product teams.

Experience being on On-Call rotation and handling critical/high incidents. 

Desired Skills
Experience building and operating microservices architectures using Spring Boot, Groovy, React, or similar.
Strong understanding of CI/CD pipelines, release automation, and progressive delivery.
Experience with eCommerce domains such as Catalog, Customer Data, and Order Management.
Familiarity with search platforms (Endeca, Solr, Lucene, Elasticsearch).
Proficiency in scripting and automation (Python, Bash, Ruby, Perl, PowerShell).
Experience with ITSM tools integrated with Jira workflows.
Exposure to capacity planning, load testing, and chaos engineering.


Skills Required

  • 6+ years experience in SRE, software engineering, or production operations supporting large-scale eCommerce platforms
  • Hands-on experience with Java/J2EE-based distributed systems
  • Proven ability to design and operate systems using SLO-driven reliability models
  • Experience defining and measuring SLIs (availability, latency, error rates, throughput, saturation)
  • Good understanding of NoSQL technologies and RDBMS; able to write queries
  • Experience deploying and operating services on cloud platforms (AWS, Azure, or Google Cloud)
  • Expertise with observability, APM, and caching tools (Dynatrace, Splunk, ELK, Akamai, QuantumMetric/Tealeaf, etc.)
  • Strong experience using Jira for backlog management, incident follow-ups, and coordination
  • Experience on on-call rotation and handling critical/high incidents
  • Ability to independently own services and drive reliability initiatives end-to-end
  • Provide technical leadership and mentoring to junior SREs and engineers
  • Strong communication skills and ability to influence engineering and product teams
  • Experience building and operating microservices architectures using Spring Boot, Groovy, React (desired)
  • Strong understanding of CI/CD pipelines, release automation, and progressive delivery (desired)
  • Familiarity with search platforms (Endeca, Solr, Lucene, Elasticsearch) (desired)
  • Proficiency in scripting and automation (Python, Bash, Ruby, Perl, PowerShell) (desired)
  • Experience with ITSM tools integrated with Jira workflows (desired)
  • Exposure to capacity planning, load testing, and chaos engineering (desired)
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: London
5,017 Employees
Year Founded: 2007

What We Do

Photon.com has emerged as one of the world’s largest and fastest-growing Digital Agencies. We work with 40% of the Fortune 100 on their Digital initiatives and are known for our ability to integrate Strategy Consulting, Creative Design, and Technology at scale. Please visit www.photon.com to learn more about us, how we work, and our customer case studies. Digital Transformation Starts Here.

Similar Jobs

Remote
6 Locations
77 Employees

Magna International Logo Magna International

Program Manager

Automotive • Hardware • Robotics • Software • Transportation • Manufacturing
Remote or Hybrid
COA, Coahuila De Zaragoza, MEX
171000 Employees

Samsara Logo Samsara

Technical Support

Artificial Intelligence • Cloud • Computer Vision • Hardware • Internet of Things • Software
Easy Apply
Remote or Hybrid
México
4000 Employees

Capital One Logo Capital One

Sr. Manager, Associate Relations

Fintech • Machine Learning • Payments • Software • Financial Services
Remote or Hybrid
Mexico City, Ciudad De México, MEX
55000 Employees

Similar Companies Hiring

Scrunch  Thumbnail
Artificial Intelligence • Information Technology • Marketing Tech • Software • SEO
Salt Lake City, Utah
Standard Template Labs Thumbnail
Artificial Intelligence • Information Technology • Software
New York, NY
25 Employees
Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account