Principal SRE

Posted 3 Days Ago
Be an Early Applicant
Bengaluru, Bengaluru Urban, Karnataka, IND
In-Office
Expert/Leader
Marketing Tech • Software
The Role
Lead Azira’s cloud infrastructure and site reliability strategy across high-volume data and AI workloads. Design resilient, secure, highly available systems; improve observability, automation, disaster recovery, and cloud cost efficiency. Establish SLI, SLO, and error-budget practices, lead complex incident response, and drive long-term reliability improvements. Partner across Engineering, Data, AI, Product, and Security while mentoring engineers and communicating technical trade-offs to stakeholders.
Summary Generated by Built In

About Azira

Azira is a data-first media and insights company on a mission to reinvent how brands use data to make smarter decisions, from where to open their next location to how they connect with customers in the real world. We blend marketing, location analytics, and strategy into a single platform, helping leading brands take action with confidence. We move fast, think boldly, and care deeply about building things that matter.

Why This Role Matters

This role will contribute to the Development and Improvement of Azira’s Cloud infrastructure, addressing key areas such as scalability, observability, security, and cost efficiency. The role will collaborate with teams across Engineering, AI, Product, Data, and Security to improve infrastructure standards, solve complex reliability challenges, and provide technical guidance and mentorship across the organization.

What you’ll do

  • Define and evolve Azira’s cloud infrastructure and reliability strategy, ensuring it supports global products, data platforms, and AI initiatives.
  • Design and implement scalable, resilient, secure, and highly available systems, including architecture standards, best practices, and disaster-recovery capabilities.
  • Improve the reliability and performance of distributed, high-volume production systems by identifying and addressing single points of failure, capacity constraints, and recurring sources of instability.
  • Develop and mature service-level indicators (SLIs), service-level objectives (SLOs), and error-budget practices to establish measurable reliability standards.
  • Strengthen observability across applications and infrastructure through effective monitoring, logging, tracing, and alerting.
  • Lead the technical response to complex production incidents, contribute as a senior escalation point when needed, and ensure incident learnings result in lasting improvements.
  • Advance infrastructure-as-code, automation, and environment management practices to improve consistency, repeatability, scalability, and operational efficiency.
  • Partner with Engineering, Data, AI, Product, and Security teams to embed reliability, security, and operational readiness throughout the development lifecycle, including supporting the infrastructure needs of AI workloads.
  • Improve cloud cost visibility and efficiency while balancing performance, capacity, scalability, and spend; strengthen business continuity, backup, and disaster-recovery practices.
  • Mentor SREs and engineers, raise technical standards, and communicate infrastructure risks, trade-offs, and recommendations clearly to technical leaders and business stakeholders.

What you Bring

  • Bachelor’s or Master’s degree in Computer Science, Engineering, or a related field.
  • 12–15 years of experience in Site Reliability Engineering, Cloud Infrastructure, Platform Architecture, or related roles, with experience operating at a Principal, Staff-plus, Architect, or equivalent level.
  • Significant experience in Site Reliability Engineering, cloud infrastructure, platform engineering, DevOps, or a related discipline, with experience operating at a Principal, Staff-plus, Architect, or equivalent level.
  • Deep expertise in AWS, including experience architecting and operating infrastructure for products involving big data, high-volume workloads, or distributed systems. Experience with GCP in similar environments is also valuable.
  • Strong understanding of cloud-native architecture, including Kubernetes, containers, networking, Linux, storage, databases, Amazon EMR, and cloud security.
  • Advanced experience with Infrastructure-as-Code tools such as Terraform or OpenTofu, with a focus on building consistent, repeatable, and scalable infrastructure.
  • Strong scripting and programming skills in Python, Bash, Go, or comparable languages, along with experience working with CI/CD and source-control platforms such as GitHub or Bitbucket.
  • Experience with centralized authentication and authorization systems, including concepts such as OIDC and RBAC, as well as a solid understanding of infrastructure-level security and compliance requirements.
  • Experience designing and operating distributed, data-intensive, or highly available systems at scale, with a strong understanding of observability across metrics, logs, traces, and alerting.
  • Proven experience leading complex incident response, root-cause analysis, and reliability improvement initiatives, with the ability to balance immediate operational needs with long-term architectural improvements.
  • Strong technical judgment and communication skills, with the ability to evaluate trade-offs across reliability, performance, security, speed, and cost, and communicate recommendations clearly across teams and regions.
  • A collaborative leadership approach with a track record of mentoring engineers, influencing without formal authority, constructively challenging existing approaches, and taking ownership of problems through to resolution.

Skills Required

  • Bachelor's or Master's degree in Computer Science, Engineering, or a related field
  • 12-15 years of experience in Site Reliability Engineering, Cloud Infrastructure, Platform Architecture, or related roles
  • Experience operating at a Principal, Staff-plus, Architect, or equivalent level
  • Significant experience in Site Reliability Engineering, cloud infrastructure, platform engineering, DevOps, or a related discipline
  • Deep expertise in AWS
  • Experience architecting and operating infrastructure for big data, high-volume workloads, or distributed systems
  • Experience with GCP in similar environments
  • Strong understanding of cloud-native architecture, Kubernetes, containers, networking, Linux, storage, databases, Amazon EMR, and cloud security
  • Advanced experience with Terraform or OpenTofu
  • Strong scripting and programming skills in Python, Bash, Go, or comparable languages
  • Experience with CI/CD and source-control platforms such as GitHub or Bitbucket
  • Experience with centralized authentication and authorization systems, including OIDC and RBAC
  • Understanding of infrastructure-level security and compliance requirements
  • Experience designing and operating distributed, data-intensive, or highly available systems at scale
  • Strong understanding of observability across metrics, logs, traces, and alerting
  • Experience leading complex incident response, root-cause analysis, and reliability improvement initiatives
  • Strong technical judgment and communication skills
  • Experience mentoring engineers and influencing without formal authority
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Pasadena, California
420 Employees
Year Founded: 2012

What We Do

Azira LLC, a global Consumer Insights platform, helps marketing and operational leaders improve their effectiveness with actionable intelligence to drive business results. Azira delivers innovative marketing solutions to curate audiences, activate omnichannel campaigns, and understand footfall attribution. It also provides operational insights for use cases such as site selection, trade area analysis, competitive intelligence and more. Azira serves enterprises in retail, hospitality, travel, real estate, financial services and media. A global company, Azira is headquartered in Los Angeles with offices in Paris, Bangalore, Singapore, Sydney, and Tokyo. To learn more, please visit https://azira.com.

Similar Jobs

Hybrid
Bangalore, Bengaluru Urban, Karnataka, IND
150 Employees
71K-95K Annually

ServiceTitan Logo ServiceTitan

Site Reliability Engineer

Artificial Intelligence • Cloud • Fintech • Machine Learning • Mobile • Software
Hybrid
Bengaluru, Bengaluru Urban, Karnataka, IND
2760 Employees
In-Office
Bengaluru, Bengaluru Urban, Karnataka, IND
6000 Employees

Similar Companies Hiring

Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Revel Thumbnail
Aerospace • Hardware • Robotics • Software
Marina Del Rey, California
60 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software
New York, New York
30 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account