Senior Site Reliability Specialist II

Posted 3 Hours Ago
Be an Early Applicant
Hiring Remotely in United States
Remote
145K-177K Annually
Senior level
Information Technology • Software • Consulting
The Role
Build and operate reliable, scalable cloud platforms; lead initiatives across Kubernetes, infrastructure, observability, automation, networking, and developer platforms. Improve availability, resilience, monitoring, incident response, disaster recovery, and operational readiness. Coach engineering teams, establish reliability standards, implement SLOs and error budgets, automate repetitive work, lead high-severity incident response, and drive post-incident corrective actions.
Summary Generated by Built In

At Everbridge, reliability isn’t just about uptime—it’s about ensuring that critical systems are available when they matter most. Every improvement you make helps organizations deliver life-saving communications and maintain operations during emergencies.

As a Senior Site Reliability Engineer II, you’ll do more than operate infrastructure. You’ll improve the resiliency of our engineering organization by building reliable platforms, eliminating operational toil, mentoring engineers, and helping teams design systems that are secure, scalable, and resilient by
default.

This is a highly collaborative technical leadership role for an engineer who enjoys solving systemic problems, influencing architecture, and enabling others to build reliable software.

What you'll do:

  • Build platform capabilities that enable engineering teams to deliver reliable software safely and efficiently.
  • Lead complex technical initiatives spanning cloud infrastructure, Kubernetes, observability, automation, networking, and developer platforms.
  • Influence engineering decisions through technical expertise, collaboration, and data.
  • Help engineering teams become increasingly self-sufficient through coaching, automation, and well-designed platform capabilities.
  • Continuously improve operational excellence by reducing complexity, eliminating manual work, and strengthening engineering practices.
  • Design and implement solutions that improve the availability, scalability, performance, and resilience of our platform.
  • Build automation that eliminates repetitive operational work and reduces engineering toil.
  • Improve observability, monitoring, alerting, and operational readiness across the organization.
  • Use production data, reliability metrics, and engineering judgment to identify systemic improvements.
  • Lead complex cross-functional engineering initiatives from design through production.
  • Partner with architects, software engineers, security, product, and platform teams to build resilient systems from the beginning.
  • Review architectures and designs with a focus on reliability, scalability, recoverability, and operational excellence.
  • Establish and evolve engineering standards, best practices, and operational readiness guidance.
  • Partner directly with engineering teams to improve the reliability of the services they own.
  • Coach teams on observability, incident response, disaster recovery, capacity planning, and production readiness.
  • Help teams adopt SLOs, error budgets, meaningful alerting, and engineering practices that improve customer outcomes.
  • Make the right engineering decisions easier through automation, paved roads, and self-service capabilities.
  • Participate in an on-call rotation supporting critical production systems.
  • Lead the technical response during high-severity incidents.
  • Facilitate blameless post-incident reviews that focus on learning and systemic improvement.
  • Drive corrective actions through completion and measure their effectiveness over time.
  • Share knowledge through documentation, technical design reviews, and collaborative problem solving.
  • Foster a culture of ownership, continuous improvement, operational excellence, and customer focus.

What you'll bring:

  • Experience with:
  • Designing and operating complex production systems.
  • Cloud infrastructure and cloud-native architectures.
  • Distributed systems and container platforms.
  • Infrastructure as Code and automation.
  • CI/CD and software delivery practices.
  • Observability, monitoring, logging, and telemetry.
  • Incident response and operational excellence.
  • Reliability engineering principles including SLOs, SLIs, capacity planning, and performance optimization.
  • Writing software or automation using one or more modern programming languages.
  • Linux and networking fundamentals.
  • Experience working within regulated environments such as FedRAMP, DoD, IL4/IL5, SOC 2, or ISO 27001 is a plus.

The reasonably estimated salary for this role at Everbridge ranges from $145,000 - $177,000 and may also include variable compensation. Actual compensation is based on factors such as the candidate's skills, qualifications, and experience. In addition, Everbridge offers a wide range of best in class, comprehensive and inclusive employee benefits for this role including healthcare, dental, parental planning, and mental health benefits, disability income benefits, life and AD&D insurance, a 401(k) plan and match, paid time off, and fitness reimbursements.
 
Fair Chance Statement US & Canada
We are committed to providing equal employment opportunities in compliance with all applicable Federal, Provincial/State and Local laws, including the California Fair Chance Act and any local County Fair Chance Ordinance (or local equivalent). Pursuant to these and other relevant regulations, we consider qualified applicants with criminal histories in a manner consistent with the law.
 
For roles subject to background checks, the following material job duties may be affected by an applicant’s criminal history:
- Access to sensitive or confidential information, such as financial records, proprietary data, or client information.
- Management of cash, company funds, or other valuable assets.
- Work in environments requiring heightened security measures.
- Compliance with contractual or regulatory requirements specific to the position.
 
We evaluate each applicant's criminal history individually, considering its nature, timing, and relevance to the specific job duties, while maintaining our commitment to fair hiring practices and promoting workplace equity.

About Everbridge

Everbridge empowers enterprises and government organizations to anticipate, mitigate, respond to, and recover from critical events. In an unpredictable world, resilient organizations protect their people and operations, adapt under pressure, and return to productivity faster. Our Critical Event Management (CEM) technology combines intelligent automation with comprehensive risk intelligence to help organizations strengthen resilience, keep people safe, and maintain operations.

 

Learn more at everbridge.com, explore the Everbridge blog, and connect with us on social media.

 

Equal Employment Opportunity

Everbridge is an Equal Opportunity Employer. We consider all qualified applicants for employment without regard to race, color, creed, religion, national origin, ancestry, age, sex, pregnancy, sexual orientation, gender identity or expression, disability, protected veteran status, genetic information, marital status, or any other characteristic protected by applicable federal, state, or local law.

 

Employment Practices

Everbridge does not require or administer lie detector (polygraph) tests as a condition of employment or continued employment. In Massachusetts, it is unlawful to require or administer a lie detector test as a condition of employment or continued employment. Employers who violate this law may be subject to criminal penalties and civil liability.

Skills Required

  • Experience designing and operating complex production systems
  • Experience with cloud infrastructure and cloud-native architectures
  • Experience with distributed systems and container platforms
  • Experience with Infrastructure as Code and automation
  • Experience with CI/CD and software delivery practices
  • Experience with observability, monitoring, logging, and telemetry
  • Experience with incident response and operational excellence
  • Experience applying reliability engineering principles, including SLOs, SLIs, capacity planning, and performance optimization
  • Experience writing software or automation using one or more modern programming languages
  • Linux and networking fundamentals
  • Experience in regulated environments such as FedRAMP, DoD, IL4/IL5, SOC 2, or ISO 27001

Everbridge Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Everbridge and has not been reviewed or approved by Everbridge.

  • Healthcare Strength Healthcare coverage appears broad, with medical, dental, vision, telehealth, disability, life/AD&D, and FSA/HSA options described. Wellness elements like an EAP and mental-health support are also included, which can add practical value beyond base pay.
  • Retirement Support Retirement offerings include a 401(k) plan with employer matching and an ESPP, indicating multiple ways to build longer-term financial value. Match details are sometimes specified externally, but the presence of a match and purchase program is consistently referenced.
  • Leave & Time Off Breadth Time-off benefits include flexible PTO plus paid volunteer time, providing both personal flexibility and community-focused leave. Paid holidays and new-hire PTO amounts are also described in some materials, suggesting a reasonably complete time-away menu.

Everbridge Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Vienna, VA
1,437 Employees
Year Founded: 2002

What We Do

Keeping People Safe and Businesses Running. Faster. Everbridge, Inc. (NASDAQ: EVBG) is a global software company that provides enterprise software applications that automate and accelerate organizations’ operational response to critical events in order to Keep People Safe and Businesses Running™. During public safety threats such as active shooter situations, terrorist attacks or severe weather conditions, as well as critical business events including IT outages, cyber-attacks or other incidents such as product recalls or supply-chain interruptions, over 5,300 global customers rely on the company’s Critical Event Management Platform to quickly and reliably aggregate and assess threat data, locate people at risk and responders able to assist, automate the execution of pre-defined communications processes through the secure delivery to over 100 different communication devices, and track progress on executing response plans.

Similar Jobs

Scaled Agile, Inc. Logo Scaled Agile, Inc.

Salesforce Administrator

Artificial Intelligence • Edtech • Productivity • Business Intelligence • Consulting
In-Office or Remote
Boulder, CO, USA
132 Employees
50-58 Hourly

Rain Logo Rain

Vice President Of Engineering

Blockchain • Fintech • Payments • Financial Services • Cryptocurrency • Web3 • Infrastructure as a Service (IaaS)
Remote or Hybrid
New York, NY, USA
100 Employees
320K-450K Annually

HERE Technologies Logo HERE Technologies

Architect

Artificial Intelligence • Automotive • Computer Vision • Information Technology • Internet of Things • Logistics • Software
Remote or Hybrid
2 Locations
6000 Employees
121K-173K Annually

Accuris Logo Accuris

Head of Demand Generation

Information Technology • Machine Learning • Software • Conversational AI • Generative AI • Manufacturing
Remote
CO, USA
1000 Employees
180K-210K Annually

Similar Companies Hiring

Kepler  Thumbnail
Artificial Intelligence • Fintech • Software
New York, New York
9 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Revel.io Thumbnail
Aerospace • Hardware • Robotics • Software
US
50 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account