Associate Principal Site Reliability Engineer

Reposted 27 Days Ago
Be an Early Applicant
Bengaluru, Bengaluru Urban, Karnataka, IND
Hybrid
8-8 Annually
Senior level
Software
The Role
As a Staff Site Reliability Engineer, you will oversee reliability for Saviynt's AI-driven platform, manage AWS and Kubernetes systems, automate infrastructure, and enhance incident management and observability.
Summary Generated by Built In
Saviynt's AI-powered identity platform manages and governs human and non-human access to all of an organization's applications, data, and business processes. Customers trust Saviynt to safeguard their digital assets, drive operational efficiency, and reduce compliance costs. Built for the AI age, Saviynt is today helping organizations safely accelerate their deployment and usage of AI. Saviynt is recognized as the leader in identity security, with solutions that protect and empower the world’s leading brands, Fortune 500 companies and government institutions. For more information, please visit www.saviynt.com.

We’re a fast-moving AI Security Company building AI-native infrastructure and
applications powered by LLMs and autonomous agents. Our stack is deeply integrated with AWS, Kubernetes, and OpenAI-based systems, and we’re rethinking reliability in a world where software can reason, adapt, and self-heal.
 
We’re hiring a Staff SRE Engineer to own reliability across our cloud-native and AI-driven platform. You’ll work at the intersection of distributed systems, Kubernetes operations, and LLM-powered automation, building systems that don’t just scale—but think and fix themselves.

WHAT YOU WILL BE DOING

    • Own uptime, reliability, and performance of services running on AWS + Kubernetes (EKS).
    • Design and implement self-healing infrastructure using automation and AI agents.
    • Build LLM-powered operational tooling using APIs such as the OpenAI API for:
      • Intelligent alert triage
      • Incident summarization
      • Root cause analysis
      • Runbook automation
      • Manage and scale Kubernetes workloads:
        • Deployments, autoscaling, resource optimization
        • Cluster reliability and cost efficiency
        • Build and evolve observability systems:
          • Metrics (Prometheus), dashboards (Grafana)
          • Logs (ELK / OpenSearch)
          • Tracing (OpenTelemetry)
          • Define and enforce SLOs, SLAs, and error budgets tied to business metrics.
          • Automate infrastructure using Terraform and CI/CD pipelines.
          • Lead incident response, postmortems, and continuous reliability improvements.
          • Introduce chaos engineering practices to proactively test system resilience.

WHAT YOU BRING

    • 8+ years in SRE / DevOps / Platform Engineering.
    • Strong hands-on experience with:
      • AWS infrastructure at scale
      • Kubernetes (production-grade clusters)
      • Proven ability to debug complex distributed systems under pressure.
      • Strong coding skills (Python or Go)—you build internal platforms and tools.
      • Experience implementing monitoring, alerting, and incident management systems.
      • Bonus (AI / LLM Focus)

        • Experience working with LLM APIs such as the OpenAI API.
        • Familiarity with agent frameworks like:
          • LangChain
          • AutoGen
          • Built or experimented with:
            • AI agents for DevOps / SRE workflows
            • Retrieval-Augmented Generation (RAG) systems
            • Vector databases (Pinecone, Weaviate, etc.)
            • Exposure to AIOps or intelligent automation systems.
            •  
               

Saviynt is an amazing place to work. We are a high-growth, Platform as a Service company focused on Identity Authority to power and protect the world at work. You will experience tremendous growth and learning opportunities through challenging yet rewarding work which directly impacts our customers, all within a welcoming and positive work environment. If you're resilient and enjoy working in a dynamic environment you belong with us!

Security & Compliance
This role requires adherence to Saviynt’s information security and privacy policies and procedures, including annual security training.

Saviynt is an equal opportunity employer and we welcome everyone to our team.  All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, or veteran status.

Skills Required

  • 8+ years in SRE, DevOps, or Platform Engineering
  • Strong experience with AWS infrastructure
  • Hands-on experience with production-grade Kubernetes clusters
  • Strong coding skills in Python or Go
  • Experience implementing monitoring and incident management systems

Saviynt Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Saviynt and has not been reviewed or approved by Saviynt.

  • Leave & Time Off Breadth Flexible or unlimited PTO, generous parental leave, paid holidays, and periodic mental health days point to a wide range of time‑off options. These policies are positioned to support rest and balance when teams can make use of them.
  • Healthcare Strength Medical, dental, and vision coverage are provided alongside options like FSAs, reflecting comprehensive core health benefits. Affordability and plan quality are often highlighted as positives.
  • Retirement Support A 401(k) program with employer contributions is available, reinforcing long‑term financial security. This complements equity and ESPP elements within total rewards.

Saviynt Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: El Segundo, CA
Year Founded: 2010

What We Do

Saviynt’s Enterprise Identity Cloud helps modern enterprises scale cloud initiatives and solve the toughest security and compliance challenges in record time. The company brings together identity governance (IGA), granular application access, cloud security, and privileged access to secure the entire business ecosystem and provide a frictionless user experience.

Similar Jobs

Cloudflare Logo Cloudflare

Technical Accounting Analyst (Infrastructure)

Cloud • Information Technology • Security • Software • Cybersecurity
Hybrid
Bengaluru, Bengaluru Urban, Karnataka, IND
4400 Employees

Cloudflare Logo Cloudflare

Software Engineer

Cloud • Information Technology • Security • Software • Cybersecurity
Hybrid
Bengaluru, Bengaluru Urban, Karnataka, IND
4400 Employees

Expedia Group Logo Expedia Group

Data Engineer

AdTech • eCommerce • Information Technology • Software • Travel • Generative AI
Hybrid
Bangalore, Bengaluru Urban, Karnataka, IND
16000 Employees

Capco Logo Capco

Tester (RWA/Basel)

Fintech • Professional Services • Consulting • Energy • Financial Services • Cybersecurity • Generative AI
Remote or Hybrid
India
6000 Employees

Similar Companies Hiring

Kepler  Thumbnail
Artificial Intelligence • Fintech • Software
New York, New York
9 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Revel.io Thumbnail
Aerospace • Hardware • Robotics • Software
US
50 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account