Senior Site Reliability Engineer

Posted 7 Days Ago
Be an Early Applicant
Hyderabad, Telangana, IND
In-Office
Senior level
Cloud • Sales • Software
We empower thousands of teams to grow and win.
The Role
Designs and maintains cloud automation and reliability tooling across multi-cloud SaaS environments. The role improves observability, incident response, SLOs, capacity planning, disaster recovery, and production readiness; participates in a global on-call rotation; leads high-severity incident response and postmortems; partners with engineering, product, security, vendors, and customer-facing teams; and responsibly applies AI-assisted workflows to reliability operations.
Summary Generated by Built In
About Us

Seismic is the Go-To-Market Performance company, driving how organizations turn their strategy into revenue. Seismic brings together signals from customer conversations, buyer engagement, governed content, coaching and seller behavior to turn intelligence into winning action. Through its AI revenue execution platform, the company equips customer-facing teams with the insights, governance, and contextual guidance they need to execute with confidence and deliver measurable business outcomes. Trusted by 2,500 organizations and over 3.5 million users around the globe, Seismic is headquartered in San Diego with offices across North America, Europe and Asia-Pacific. Learn more at seismic.com.


Seismic is committed to building an inclusive workplace that ignites growth for our employees and creates a culture of belonging that allows all employees to be seen and valued for who they are. Learn more about DEI at Seismic here.

Overview

About the Role 

Seismic is seeking a Senior Cloud Engineer to advance reliability across our AWS, Azure, IBM Cloud, and OCI environments, working hands-on to build automation, improve observability, and strengthen incident response as part of a globally distributed SRE org.  We prioritize a cloud agnostic approach to architecture, automation, and engineering standards to support velocity and scale. You will contribute to our reliability roadmap, provide technical guidance on reliability best practices, participate in major incident response, and work collaboratively across Product & Engineering, Security, and Customer-facing teams. 

Who You Are: 

  • Experience in a production facing SRE role supporting a complex SaaS environment. 
  • You approach discussions about existing solutions, processes, and proposals with curiosity, respect, and a collaborative mindset.   
  • Strong technical judgment across distributed systems, multi-cloud environments, Kubernetes, networking, various infrastructure technologies, GitOps, and CI/CD. 
  • Experience establishing and maturing SRE principles and practices, including SLOs, error budgets, observability, capacity planning, incident response, and toil elimination. 
  • Proficient in using observability data to resolve high-severity incidents and dig deeper into root cause during postmortems. 
  • Experience leading through high-pressure incidents and communicating clearly with technical teams, executives, customer-facing stakeholders, and third-party vendors. 
Who you are:
  • Bachelor's or Master's degree in Computer Science or related field.
  • 6+ years of experience in software engineering
  • 4+ years of experience in Devops roles, with a strong focus on building CI/CD pipelines
  • Expertise in Kubernetes, containerization (Docker), and orchestration tools.
  • Strong experience with cloud platforms like AWS, Azure, or GCP.
  • Proficiency in Infrastructure as Code (IaC) tools such as Terraform, Chef, or Ansible.
  • Hands-on experience with observability tools (e.g., New Relic, Prometheus, Grafana).
  • Proficient in scripting and programming languages (e.g., Python, Go, Bash).
  • Knowledge of event-driven autoscaling and advanced Kubernetes configurations.
  • Familiarity with CI/CD pipelines and tools like Buildkite, Spinnaker or GitHub Actions.
  • Experience with modern software development practices such as microservices architecture, containerization, and DevOps.
  • Strong understanding of distributed systems, scalability, high availability and performance optimization.
  • Excellent problem-solving skills and ability to work in a fast-paced environment.
What you'll be doing:

Operating Model:

  • Design, build, and maintain automation and tooling that reduces operational toil. 
  • Contribute to maturing reliability practices in partnership with Product & Engineering leaders. 
  • Contribute to an inclusive, high-accountability culture that encourages curiosity, collaboration, and blameless improvement. 

Incident Management and Operational Excellence:

  • Actively participate in the health and continuous improvement of the incident-management lifecycle, including detection, engagement, escalation, mitigation, stakeholder communication, post-incident review, and corrective-action follow-through. 
  • Participate in a 12-hour follow-the-sun on-call rotation within the Global SRE team.   
  • Ensure incident practices are customer-centered, data-driven, blameless, and consistent across teams while preserving accurate severity and escalation decisions. 
  • Use alert, incident, support, and SLO trends to move the organization from reactive response toward proactive risk reduction. 
  • Build measurable feedback loops that connect incident learning to engineering standards, service maturity, product priorities, and vendor actions. 
  • Work closely with application engineering teams and incorporate their feedback to improve developer experience and reduce toil.   

Reliability Strategy and Service Maturity:

  • Partner with service owners to ensure production readiness standards are met before each release stage.   
  • Provide an SRE point of view on capacity planning, resilience testing, game days, disaster-recovery readiness, and modernization of fragile or legacy workloads. 
  • Partner with Product and Engineering leaders to document critical customer workflows, define health expectations, surface dependencies early, and align reliability investment with business priorities. 
  • Participate in cross-team reliability engagements, influencing outcomes without relying on direct authority.  
  • Build strategic relationships with vendors in the observability, incident response, and cloud infrastructure domains. 

AI-First Reliability Engineering:

  • Responsibly adopt AI-assisted and agentic workflows for alert triage, incident mitigation, postmortems, trend analysis, capacity planning, SLO analysis, and self-service knowledge. 
  • Keep qualified humans in the decision loop for production-impacting actions. 
  • Improve the context available to reliability workflows by strengthening service metadata, observability data, incident records, runbooks, architecture documentation, and corrective-action quality. 
Job Posting Footer 

Please note this job description is not designed to cover or contain a comprehensive listing of activities, duties or responsibilities that are required of the employee for this job. Duties, responsibilities and activities may change at any time with or without notice.   

  

Please be aware we have noticed an increase in hiring scams potentially targeting Seismic candidates. Read our full statement on our Careers page. 

Skills Required

  • Bachelor's or Master's degree in Computer Science or a related field
  • 6+ years of software engineering experience
  • 4+ years of DevOps experience focused on building CI/CD pipelines
  • Production-facing SRE experience in a complex SaaS environment
  • Expertise with Kubernetes, Docker, and container orchestration
  • Strong experience with AWS, Azure, or GCP
  • Proficiency with Infrastructure as Code tools such as Terraform, Chef, or Ansible
  • Hands-on experience with observability tools such as New Relic, Prometheus, or Grafana
  • Proficiency in Python, Go, Bash, or similar scripting and programming languages
  • Knowledge of event-driven autoscaling and advanced Kubernetes configurations
  • Familiarity with Buildkite, Spinnaker, GitHub Actions, or similar CI/CD tools
  • Understanding of microservices, distributed systems, scalability, high availability, and performance optimization
  • Experience with SLOs, error budgets, observability, capacity planning, incident response, and toil elimination
  • Ability to lead high-pressure incidents and communicate with technical teams, executives, customers, and vendors
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: San Diego, CA
1,400 Employees
Year Founded: 2010

What We Do

Here at Seismic, we ignite growth for our company, industry, and people. We’re enablement innovators seeking the best, brightest teammates who are mission-driven and empowered by our values. We are the global leader in enablement with employees and offices around the world.. We know it’s easy to find sales clouds and marketing clouds, but enablement clouds? You don’t see those every day — or ever. That’s why we’ve introduced the first-ever Seismic Enablement Cloud. And it won’t end there, we #NeverStopGrowing.

Gallery

Gallery

Similar Jobs

Optum Logo Optum

Senior Site Reliability Engineer

Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
In-Office
Hyderabad, Telangana, IND
160000 Employees
Hybrid
Hyderabad, Telangana, IND
289097 Employees

MetLife Logo MetLife

Senior Site Reliability Engineer

Fintech • Information Technology • Insurance • Financial Services • Big Data Analytics
Hybrid
Hyderabad, Telangana, IND
43000 Employees
Hybrid
Hyderabad, Telangana, IND
289097 Employees

Similar Companies Hiring

Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Revel Thumbnail
Aerospace • Hardware • Robotics • Software
Marina Del Rey, California
70 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software
New York, New York
30 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account