Senior Software Engineer - SRE & AIOps

Posted Yesterday
Be an Early Applicant
Santa Clara, CA, USA
Hybrid
143K-243K Annually
Senior level
Artificial Intelligence • Cloud • HR Tech • Information Technology • Productivity • Software • Automation
We're putting AI to work for people.
The Role
Deploy and operate production Kubernetes across hybrid and multi-cloud environments; build auto-remediation, observability, SLO, alerting, runbook, Infrastructure-as-Code, and GitOps systems. Support incident response, on-call operations, infrastructure optimization, and continuous reliability improvements. Mentor junior engineers, reduce operational toil, strengthen security and compliance controls, and promote resilient DevOps practices across global engineering teams.
Summary Generated by Built In
Company Description

It all started when engineer Fred Luddy wrote code that automated a tedious task for his coworker, Phyllis. She cried tears of joy. That moment inspired Fred to build a company that could do that for everyone—freeing people from busywork so they could focus on meaningful work. Today, ServiceNow is the AI control tower for business reinvention. Our ServiceNow AI platform brings together any AI, any data, and any workflow— helping 85% of the Fortune 500® work smarter, faster, and better. We're building an AI-native culture where technology and talent are unstoppable together. And we're just getting started.

Join us to put AI to work for people.

 

Job Description

About the role:

ServiceNow is seeking a Senior Software Engineer - SRE & AIOps to contribute to infrastructure automation, operational resilience, and toil elimination across our hybrid cloud and data center operations. Embedded within the Site Reliability & Database Engineering organization, you will implement automation-first systems that reduce manual intervention, accelerate incident remediation, and enable our global engineering teams to operate reliably at scale.

This role combines solid hands-on technical expertise in Kubernetes, cloud platforms, and DevOps practices with growing technical leadership capabilities. You will contribute to SRE tooling design, develop auto-remediation capabilities, and help establish patterns that maintain ServiceNow's cloud platform reliability while minimizing operational toil across follow-the-sun global teams.

What you get to do in this role:

  • Deploy, operate, and troubleshoot production Kubernetes clusters across hybrid and multi-cloud environments, maintaining operational standards and supporting high-velocity application deployments.
  • Implement and maintain closed-loop auto-remediation systems that detect, classify, and resolve transient infrastructure failures, leveraging automation frameworks and machine learning insights to reduce MTTR and on-call burden.
  • Contribute to the design and evolution of SRE tooling stack, including monitoring platforms, incident management systems, log aggregation, and observability integrations that support global on-call operations.
  • Develop and maintain SLO frameworks, alerting policies, and automated runbooks that empower on-call engineers to resolve issues autonomously while managing alert fatigue.
  • Build and maintain Infrastructure-as-Code frameworks and GitOps pipelines that enable reproducible infrastructure deployments across hybrid and multi-cloud environments with security and compliance guardrails.
  • Support hybrid cloud and data center operations, including on-premises infrastructure, public cloud environments, and workload optimization across multi-region deployments.
  • Contribute to adoption of containerization, microservices, and DevOps patterns across engineering teams, establishing CI/CD best practices and network security controls.
  • Support on-call rotation operations and incident response processes across different time zones, helping develop runbooks and contributing to post-incident reviews that drive continuous improvement.
  • Share knowledge and mentor junior SRE engineers on reliability patterns, incident investigation techniques, and automation best practices.
  • Champion a culture of blameless incident analysis, data-driven decision-making, and continuous improvement through knowledge sharing and documentation.
  • Identify and systematically automate repetitive operational tasks, from infrastructure provisioning to incident response, improving team efficiency and capacity.

Qualifications

To be successful in this role you have:

  • Kubernetes Proficiency: Solid hands-on experience operating production Kubernetes clusters, including deployment models, pod orchestration, resource management, network policies, and troubleshooting runtime issues.
  • Incident Remediation Experience: Demonstrated experience designing and implementing automated remediation systems, including alert automation, runbook development, and self-healing mechanisms.
  • Cloud Platform Knowledge: Strong hands-on experience with AWS (EKS, EC2, RDS) and/or Azure (AKS, VMs) or GCP (GKE), with understanding of core SRE-related services.
  • DevOps & IaC Skills: Solid experience with Infrastructure-as-Code tools (Terraform, CloudFormation) and GitOps practices.
  • SRE Tooling Familiarity: Working knowledge of observability platforms, incident management systems, and log aggregation tools.
  • Distributed Systems Understanding: Understanding of distributed system challenges, fault tolerance, and resilience patterns.
  • On-Call Operations: Experience participating in on-call rotations and understanding 24/7 operational models, runbook development, and escalation procedures.
  • Cloud & Hybrid Operations: Hands-on experience working with cloud infrastructure and understanding hybrid cloud concepts.
  • Systems Administration: Strong foundation in Linux system administration, performance troubleshooting, and scripting (Python, Go, or Bash).
  • Collaborative Mindset: Ability to work effectively with infrastructure and application teams, contribute to technical discussions, and help drive reliability improvements.

Qualifications

  • Experience in leveraging or critically thinking about how to integrate AI into work processes, decision-making, or problem-solving. This may include using AI-powered tools, automating workflows, analyzing AI-driven insights, or exploring AI's potential impact on the function or industry.
  • 5+ years in software engineering or infrastructure operations, with 3+ years in SRE, DevOps, or cloud platform engineering roles with a Bachelor's degree; or 3 years and a Master's degree; or a PhD without experience; or equivalent work experience.
  • 2+ years of hands-on experience working with production Kubernetes clusters.
  • Proficiency in at least one Infrastructure-as-Code tool: Terraform, CloudFormation, or equivalent.
  • Demonstrable hands-on experience with at least one major cloud platform: AWS, Azure, or GCP.
  • Experience operating in on-call environments and participating in incident response.
  • Experience implementing or improving automated remediation and alert systems.
  • Strong foundation in Linux system administration, performance troubleshooting, and scripting (Python, Go, Bash).
  • Demonstrated commitment to reliability engineering and continuous improvement through hands-on contributions.
  • Bachelor's degree in computer science, Computer Engineering, or related field (or equivalent professional experience).

Preferred:

  • Kubernetes certification (CKA, CKAD, or equivalent).
  • Experience with service mesh technologies or advanced Kubernetes networking.
  • Background in cloud migration or infrastructure modernization projects.
  • Experience with cost optimization in cloud environments.
  • Track record of implementing automation solutions that significantly reduced operational toil.

Why This Role?

This role offers the opportunity to work with infrastructure automation and reliability engineering at scale. You will implement systems and practices that directly reduce operational burden across ServiceNow's global engineering teams. Your contributions will help establish reliable, automated infrastructure operations and provide a clear career path toward senior technical leadership. This is a role for an engineer who enjoys solving complex operational challenges, continuous learning, and working collaboratively to improve how systems operate.

 

 

 

For positions in this location, we offer a base pay of $143,200 - $243,400, plus equity (when applicable), variable/incentive compensation and benefits. Sales positions generally offer a competitive On Target Earnings (OTE) incentive compensation structure. Please note that the base pay shown is a guideline, and individual total compensation will vary based on factors such as qualifications, skill level, competencies, and work location. We also offer health plans, including flexible spending accounts, a 401(k) Plan with company match, ESPP, matching donations, a flexible time away plan and family leave programs. Compensation is based on the geographic location in which the role is located and is subject to change based on work location.

Additional Information

Work Personas

We approach our distributed world of work with flexibility and trust. Work personas (flexible, remote, or required in office) are categories that are assigned to ServiceNow employees depending on the nature of their work and their assigned work location. Learn more here. To determine eligibility for a work persona, ServiceNow may confirm the distance between your primary residence and the closest ServiceNow office using a third-party service.

Equal Opportunity Employer

ServiceNow is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, national origin, age, disability, gender identity,  veteran status, or any other category protected by law. In addition, all qualified applicants with arrest or conviction records will be considered for employment in accordance with legal requirements.  

Accommodations

We strive to create an accessible and inclusive experience for all candidates. If you require a reasonable accommodation to complete any part of the application process, or are unable to use this online application and need an alternative method to apply, please contact [email protected] for assistance. 

Export Control Regulations

For positions requiring access to controlled technology subject to export control regulations, including the U.S. Export Administration Regulations (EAR), ServiceNow may be required to obtain export control approval from government authorities for certain individuals. All employment is contingent upon ServiceNow obtaining any export license or other approval that may be required by relevant export control authorities. 

From Fortune. ©2026 Fortune Media IP Limited. All rights reserved. Used under license.

Skills Required

  • Experience integrating AI into work processes, decision-making, or problem-solving
  • 5+ years in software engineering or infrastructure operations, including 3+ years in SRE, DevOps, or cloud platform engineering with a Bachelor's degree; or 3 years with a Master's degree; or a PhD without experience; or equivalent work experience
  • 2+ years of hands-on experience working with production Kubernetes clusters
  • Proficiency in at least one Infrastructure-as-Code tool, such as Terraform or CloudFormation
  • Hands-on experience with at least one major cloud platform: AWS, Azure, or GCP
  • Experience operating in on-call environments and participating in incident response
  • Experience implementing or improving automated remediation and alert systems
  • Strong foundation in Linux system administration, performance troubleshooting, and scripting with Python, Go, or Bash
  • Demonstrated commitment to reliability engineering and continuous improvement through hands-on contributions
  • Bachelor's degree in computer science, Computer Engineering, or related field, or equivalent professional experience
  • Kubernetes certification such as CKA or CKAD
  • Experience with service mesh technologies or advanced Kubernetes networking
  • Background in cloud migration or infrastructure modernization projects
  • Experience with cost optimization in cloud environments
  • Track record of implementing automation solutions that significantly reduced operational toil

What the Team is Saying

Shanequa
Katya
Suzanne
Alexander
Jaime
Pat
Brady
Hasan
Jamil
Viviana

ServiceNow Compensation & Benefits Highlights

  • Healthcare Strength Health coverage is described as comprehensive with multiple plan choices and strong perceived coverage, alongside mental-health resources and wellbeing support. Company materials and employer-verified summaries also note inclusive care options and supportive programs.
  • Parental & Family Support Parental leave and family-planning support are characterized as generous, with fully paid leave and resources such as fertility, caregiving, and adoption assistance. Backup care and other family-focused programs are also highlighted as part of the package.
  • Equity Value & Accessibility Equity components like RSUs and an employee stock purchase plan are presented as meaningful, widely available parts of total rewards. Many role and benefits overviews emphasize equity’s role in boosting overall compensation alongside bonuses.

ServiceNow Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Santa Clara, CA
29,000 Employees
Year Founded: 2004

What We Do

As the AI platform for business transformation, we're putting AI to work across organizations — freeing people for work that matters. Making old tech work with new tech. Reaching across departments, from the front office to the back office and every office in between. Our ambition? To become the AI defining enterprise software company of the 21st century (or "AI DESCO21C," as we like to call it). With more than 8,400+ customers, we serve approximately 90% of the Fortune 500®, and we're proud to be a Fortune 100 Best Companies to Work For® and World's Most Admired Companies™. Explore your future career with us, visit www.careers.servicenow.com From Fortune. ©2026 Fortune Media IP Limited. All rights reserved. Used under license.

Why Work With Us

By joining ServiceNow, you are part of an ambitious team of change-makers who have a restless curiosity and a drive for ingenuity. We're committed to helping our people do their best work and live their best lives so we can fulfill our purpose together. At the fastest-growing enterprise software company, you can grow your career faster.

Gallery

Gallery
Gallery
Gallery
Gallery
Gallery
Gallery
Gallery
Gallery
Gallery
Gallery

ServiceNow Offices

Hybrid Workspace

Employees engage in a combination of remote and on-site work.

At ServiceNow, we lead with flexibility and trust. For some, home is the primary workplace. For those who come into a ServiceNow workplace, you are empowered to make team-guided and individual-led decisions on how and when you use the workplace.

Typical time on-site: Flexible
Company Office Image
HQSanta Clara, CA
Heredia
Ciudad de México
District of Columbia
Osaka
Aarhus, DK
Aarhus, DK
Company Office Image
Addison, TX
Amsterdam, North Holland
Atlanta, Georgia
Auckland, Auckland
Bangkok, Bangkok
Bengaluru, Karnataka
Bengaluru, Karnataka
Berlin, Berlin
Brasília, Federal District
Brisbane, Queensland
Brussels, BE
Cairo, Cairo Governorate
Canberra, Australian Capital Territory
Charlottesville, Virginia
Company Office Image
Chicago, IL
Deerfield, Illinois
Company Office Image
Denver, CO
Dubai, Dubai
Dublin, Leinster
Düsseldorf, Nordrhein-Westfalen
Frankfurt am Main, Hesse
Goteborg, Västra Götaland County
Gurugram, Haryana
Hamburg, Hamburg
Hanyang, Seoul
Helsinki, Uusimaa
Hong Kong, Hong Kong
Houston, TX
Hyderabad, Telangana
Issy-les-Moulineaux, Île-de-France
Johannesburg, Gauteng
Kirkland, WA
Lausanne, Vaud
Lille, Hauts de France
London, England
London, England
Madrid, Community of Madrid
Melbourne, Victoria
Milano, Lombardia
Milwaukee, WI
Company Office Image
Montréal, QC
Mumbai, Maharashtra
Munich, Bavaria
Company Office Image
New York, NY
Opfikon, Zürich
Orlando, FL
Oslo, Oslo
Perth, Western Australia
Petah Tikva, Central District
Company Office Image
Pleasanton, CA
Riyadh, Riyadh Province
Rome, Lazio
Company Office Image
San Diego, CA
San Francisco, Heredia
Company Office Image
San Francisco, CA
São Paulo, SP
Singapore, SG
Solna, Stockholm County
Sydney, New South Wales
Tokyo, Tokyo
Toronto, Ontario
Vancouver, British Columbia
Company Office Image
Vienna, VA
Vienna, AT
Company Office Image
Waltham, MA
Washington, DC
Wellington, Wellington
West Palm Beach, Florida
Learn more

Similar Jobs

ServiceNow Logo ServiceNow

Staff Software Engineer

Artificial Intelligence • Cloud • HR Tech • Information Technology • Productivity • Software • Automation
Hybrid
Santa Clara, CA, USA
29000 Employees
191K-334K Annually

ServiceNow Logo ServiceNow

Senior Manager - Software Engineering Management - AI Engineering

Artificial Intelligence • Cloud • HR Tech • Information Technology • Productivity • Software • Automation
Hybrid
Santa Clara, CA, USA
29000 Employees
201K-352K Annually

ServiceNow Logo ServiceNow

Principal Product Designer

Artificial Intelligence • Cloud • HR Tech • Information Technology • Productivity • Software • Automation
Hybrid
Santa Clara, CA, USA
29000 Employees
221K-387K Annually

ServiceNow Logo ServiceNow

Marketing Associate

Artificial Intelligence • Cloud • HR Tech • Information Technology • Productivity • Software • Automation
Hybrid
Santa Clara, CA, USA
29000 Employees
82K-118K Annually

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account