Staff Site Reliability Engineer

Posted 5 Hours Ago
Be an Early Applicant
Hiring Remotely in Spain
Remote
Expert/Leader
eCommerce • Fashion • Retail
The Role
Owns reliability, scalability, and operability for enterprise data and AI platforms across GCP and Azure. Defines SLOs, SLIs, error budgets, observability, incident response, and toil-reduction programs. Builds self-service infrastructure using Terraform, Helm, and GitOps; architects GCP workloads, multi-cloud platforms, self-healing systems, and security controls. Applies SRE practices to agentic AI systems, mentors engineers, leads reliability improvements, and communicates platform health and roadmaps to technical and executive stakeholders.
Summary Generated by Built In

Job Location: Spain

 

Calling all originals: At Levi Strauss & Co., you can be yourself — and be part of something bigger. We’re a company of people who like to forge our own path and leave the world better than we found it. Who believe that what makes us different makes us stronger. So add your voice. Make an impact. Find your fit — and your future. 

We're seeking an exceptional Staff Site Reliability Engineer to join our Data & AI Platform Engineering team. In this role, you'll own and elevate the reliability, scalability, and operability of our enterprise data and AI platforms — the platforms that power everything from the design of our iconic jeans to the optimization of our global retail and supply chain.

As a hands-on technical leader, you'll embody the principles of Google's SRE discipline: eliminating toil, engineering for reliability, and building a culture of shared ownership between development and operations. This is a unique opportunity to shape how a legendary brand runs production at scale on Google Cloud Platform, with a growing multi-cloud footprint across GCP and Azure.

 

About the Job 

Reliability & Incident Management

  • Define, instrument, and enforce SLOs, SLIs, and error budgets across all platform services, ensuring alignment with business and product commitments

  • Drive continuous reduction in MTTD and MTTR through improved observability, automated alerting, and runbook-driven incident response

  • Lead blameless post-mortems and translate findings into durable reliability improvements, ensuring systemic issues are eliminated rather than patched

 

Toil Reduction & Automation

  • Systematically identify, measure, and eliminate operational toil; track toil percentage per sprint and enforce guardrails to keep it below 50% of engineering capacity

  • Build and maintain self-serve infrastructure capabilities — enabling product and data engineering teams to provision, scale, and operate their own resources safely and consistently

  • Automate deployment pipelines, configuration management, and operational workflows using Infrastructure-as-Code principles (Terraform, Helm, GitOps)

Platform Engineering & Architecture

  • Serve as the primary GCP subject matter expert — architecting and optimizing workloads across GKE, Cloud Run, BigQuery, Pub/Sub, GCS, Composer, Dataflow, and Vertex AI

  • Lead multi-cloud architecture decisions across GCP and Azure, ensuring consistent security posture, cost efficiency, and operational practices across environments

  • Design and implement self-healing infrastructure patterns, auto-scaling strategies, and capacity planning models to support high-availability data and AI platforms

  • Champion data security and governance best practices — including encryption at rest and in transit, IAM least-privilege, secrets management, and audit logging

AI, Agentic Systems & Modern Observability

  • Apply SRE principles to agentic AI workloads — defining reliability expectations for LLM-based and multi-agent systems, including latency SLOs, fallback patterns, and model observability

  • Partner with AI Platform teams to productionize agentic pipelines with robust monitoring, drift detection, and rollback capabilities

  • Drive adoption of AI-assisted operations tooling to enhance observability, anomaly detection, and predictive incident management

Leadership & Culture

  • Guide and mentor junior and mid-level SREs — conducting code reviews, running reliability reviews, and elevating the team's engineering craft

  • Collaborate cross-functionally with Data Engineering, Software Engineering, Security, and Product teams to embed reliability as a shared value from design through deployment

  • Champion a culture of psychological safety, continuous learning, and reliability excellence

  • Communicate platform health, risk posture, and reliability roadmaps clearly to both technical and executive audiences

About You 

Required Qualifications

  • Master's degree in Computer Science, Engineering, or related field (or equivalent practical experience)

  • 10+ years of experience in Site Reliability Engineering, DevOps, or Platform Engineering, with a strong track record in large-scale production environments

  • Deep, hands-on expertise in GCP — including GKE, Cloud Run, BigQuery, Pub/Sub, GCS, Composer, Dataflow, and Vertex AI

  • Proficiency with Infrastructure-as-Code tools (Terraform, Helm) and GitOps workflows (ArgoCD, Flux)

  • Strong command of observability tooling — distributed tracing, structured logging, metrics pipelines, and alerting platforms (e.g., Cloud Monitoring, Datadog, Prometheus/Grafana)

  • Proven experience defining and operating against SLOs, SLIs, and error budgets in production environments

  • Solid understanding of data security principles: IAM, encryption, secrets management, network policies, and compliance frameworks

  • Experience with multi-cloud environments (GCP + Azure), including cross-cloud networking, identity federation, and cost governance

  • Demonstrated ability to lead without authority — influencing engineers across teams and driving reliability improvements at the organizational level

  • Excellent written and verbal communication skills; ability to translate complex reliability concepts for non-technical stakeholders

Technical Depth

  • Fluency in at least one systems or scripting language (Python, Go, or Bash) for automation and tooling

  • Experience with container orchestration (Kubernetes/GKE), service mesh, and traffic management patterns

  • Familiarity with data engineering patterns: batch and streaming pipelines, data warehouses, and the operational challenges of large-scale data platforms

  • Understanding of agentic AI architectures and the unique reliability challenges of LLM-based, event-driven, and multi-agent systems

  • Working knowledge of data governance frameworks, data lineage tooling, and platform-level data quality enforcement

Desirable Experience

  • Experience operating data platforms in retail or e-commerce environments

  • Familiarity with SRE principles in practice — error budget policies, CRE engagements, production readiness reviews

  • Exposure to FinOps practices — cloud cost attribution, commitment optimization, and unit economics for data workloads

  • Experience with Azure-native services (AKS, Azure Data Factory, Event Hubs) and cross-cloud identity management

  • Prior experience in a Staff or Principal-level SRE role with organization-wide scope

LOCATIONSpain - RemoteFULL TIME/PART TIMEFull timeCurrent LS&Co Employees, apply via your Workday account.

Skills Required

  • Master's degree in Computer Science, Engineering, or a related field, or equivalent practical experience
  • 10+ years of experience in Site Reliability Engineering, DevOps, or Platform Engineering
  • Hands-on expertise with GCP, including GKE, Cloud Run, BigQuery, Pub/Sub, GCS, Composer, Dataflow, and Vertex AI
  • Proficiency with Terraform, Helm, and GitOps workflows using ArgoCD or Flux
  • Strong command of observability tooling, including distributed tracing, structured logging, metrics pipelines, and alerting platforms
  • Experience defining and operating against SLOs, SLIs, and error budgets in production
  • Understanding of IAM, encryption, secrets management, network policies, and compliance frameworks
  • Experience with multi-cloud environments using GCP and Azure, including cross-cloud networking, identity federation, and cost governance
  • Ability to lead without authority and drive organization-wide reliability improvements
  • Excellent written and verbal communication skills
  • Fluency in at least one systems or scripting language: Python, Go, or Bash
  • Experience with Kubernetes or GKE, container orchestration, service mesh, and traffic management
  • Familiarity with batch and streaming data pipelines, data warehouses, and large-scale data platform operations
  • Understanding of agentic AI architectures and reliability challenges for LLM-based, event-driven, and multi-agent systems
  • Working knowledge of data governance frameworks, data lineage tooling, and platform-level data quality enforcement
  • Experience operating data platforms in retail or e-commerce environments
  • Familiarity with SRE practices such as error budget policies, CRE engagements, and production readiness reviews
  • Exposure to FinOps practices, including cloud cost attribution, commitment optimization, and unit economics
  • Experience with Azure-native services such as AKS, Azure Data Factory, and Event Hubs
  • Prior Staff- or Principal-level SRE experience with organization-wide scope
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: San Francisco, CA

What We Do

We’re a company of people who like to forge our own path. We invented the blue jean in 1873, and we reinvented khaki pants in 1986. We pioneered labor and environmental guidelines in manufacturing. And we work to build sustainability into everything we do. We just might be the original startup.

Similar Jobs

Remote
Spain
1300 Employees
94K-113K Annually

OneSignal Logo OneSignal

Staff Software Engineer

Mobile • Other • Software • Analytics
Remote
2 Locations
102 Employees
100K-145K Annually

Stellar Cyber Logo Stellar Cyber

Site Reliability Engineer

Software • Cybersecurity
Remote
Spain
93 Employees
85K-115K Annually

Similar Companies Hiring

Tastewise Thumbnail
Artificial Intelligence • Big Data • Food • Retail • Software • Generative AI • Big Data Analytics
NYC, NYC
120 Employees
Scotch Thumbnail
Artificial Intelligence • eCommerce • Fintech • Payments • Retail • Software • Analytics
US
35 Employees
Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account