Staff Site Reliability Engineer

Posted Yesterday
Hiring Remotely in US
Remote
200K-270K Annually
Entry level
Artificial Intelligence • Cybersecurity
The Role
Own and evolve the company-wide SRE strategy, reliability standards, observability practices, incident management, service ownership, SLOs, and on-call operations. Lead cross-functional reliability initiatives, establish dashboards, alerts, runbooks, and escalation paths, improve production readiness and incident response, and operate large-scale distributed systems across AWS and Kubernetes. Participate in a 24/7 on-call rotation and help shape the SRE function.
Summary Generated by Built In

Get to Know Us

Horizon3 is a fast-growing, remote cybersecurity company dedicated to the mission of enabling organizations to proactively find, fix and verify exploitable attack vectors before criminals exploit them. Our flagship product, the NodeZeroTM platform, delivers production-safe autonomous pentests and other key assessment operations that scale across the largest internal, external, cloud, and hybrid cloud environments. NodeZero has been adopted by organizations of all sizes, from small educational institutions to government agencies and Global 100 enterprises. It is used by IT Ops/SecOps teams, consulting pentesters, and MSSPs and MSPs.

We are a fusion of former U.S. Special Operations cyber operators, startup engineers & operators, and formerly frustrated cybersecurity practitioners. We're committed to helping solve our common security problems: ineffective security tools and false positives, resulting in alert fatigue, blind spots, "checkbox” security culture, cybersecurity skills shortage, and the long lead time and expense of hiring outside consultants. Collectively, we are a team of learn it alls, committed to a culture of respect, collaboration, ownership, and results.

We are seeking a hands-on Staff Site Reliability Engineer to own and evolve the reliability strategy, operating model, and engineering-wide standards supporting our platform. This is a foundational role for an experienced engineer who will set technical direction across teams, lead the highest-impact reliability initiatives, and establish the practices and systems that enable engineering to operate production services safely.

What You’ll Do

  • Own and evolve the engineering-wide SRE strategy, operating model, and reliability standards, aligning them to customer impact, business priorities, and risk.

  • Lead cross-functional alignment across Infrastructure, product, service, security, and business stakeholders to improve reliability, observability, incident response, and operational readiness across multiple teams.

  • Establish an organization-wide approach to service ownership, meaningful SLIs and SLOs, and error budgets for critical customer paths and services.

  • Define and drive adoption of observability standards across pipelines and platform components that report on service health, performance and operational risk.

  • Set the standard for dashboards, actionable alerting, runbooks, and escalation paths.

  • Drive end to end complex cross-functional reliability initiatives.

  • Set and raise the engineering wide bar for incident management, incident command, on-call health, post-incident learning, and recovery readiness.

  • Shape the technical direction, operationable model, and growth path of the SRE function.

  • Participate in a 24/7 on-call rotation and help design an on-call model that is sustainable, appropriately staffed, and continuously improved.

What You’ll Bring

  • Experience designing, operating, and troubleshooting large scale distributed systems in production environments.

  • Deep knowledge of reliability engineering, observability, incident management, and production operations, with demonstrated ability to turn that knowledge into standards and practices adopted by others

  • Experience in establishing SLIs, SLOs, actionable alerts, observability, and service ownership.

  • Backend experience building backend systems and automation that reduce optional toil, strengthen safeguards, and operational efficiency.

  • Experience in leading high severity incidents and improving incident response programs.

  • Excellent written and verbal communication skills including technical designs, runbooks, postmortems, and operational documentation.

Required Tech Stack Experience

  • Python and Terraform (Infrastructure as Code), or equivalent automation and infrastructure-as-code tools.

  • Experience with Observability tools such as Datadog, New Relic, Grafana, or equivalent platforms.

  • Experience operating production services in AWS and Kubernetes

  • Experience with CI/CD pipelines such as Gitlab CI, ArgoCD, or GitOps workflows..

Travel Required

We are a fully remote company, and this job may require up to 10% of travel to be successful. Travel primarily consists of team off-sites and in-person project kick-offs.

 

Perks of Horizon3

  • Inclusive Team: We value diversity and promote an inclusive culture where everyone can thrive.

  • Growth Opportunities: Be part of a dynamic and growing team with numerous career development opportunities.

  • Innovative Culture: Work in a collaborative environment that encourages creativity and out-of-the-box thinking.

  • Hybrid & Remote Work: We embrace a mix of remote and hybrid work models depending on role and location, including our Chicago office, where some roles require regular in-office presence.

  • Competitive Compensation: We offer competitive salary, equity and benefits. Our benefits include health, vision & dental insurance for you and your family, a flexible vacation policy, and generous parental leave.

Compensation and Values

At Horizon3, we believe that our people are our greatest asset, and our compensation philosophy reflects this core value. We are committed to fostering an environment where all employees feel valued, respected, and rewarded for their contributions. Our compensation structure is designed to be fair, competitive, and transparent, ensuring that every team member is recognized and compensated equitably across roles, levels, and locations.

In accordance with various State’s transparency regulations, we provide the following salary range information for this position:

  • Base salary range: $199,750 - $270,000 annually. The exact salary will be determined based on the selected candidate’s location, qualifications, experience, and relevant skills.

  • Additional compensation: All full-time roles are eligible for an equity package in the form of stock options.

You Belong Here

Horizon3 is not just an equal opportunity employer - we are a community that values diversity, equity, and inclusion as fundamental principles of our culture and success. We are dedicated to fostering a workplace where everyone feels welcome and respected, regardless of race, color, religion, sex, national origin, age, disability, veteran status, sexual orientation, gender identity or expression, genetic information, marital status, or any other legally protected status by law.

Our commitment to diversity and inclusion means we strive to attract, develop, and retain a workforce that reflects the varied communities we serve. We believe that diverse perspectives drive innovation and strengthen our ability to create cutting-edge cybersecurity solutions. At Horizon3, every team member is valued and supported in an environment that encourages personal and professional growth.

We welcome candidates from all backgrounds and experiences, and we encourage all qualified individuals to apply. Come be a part of Horizon3, where your unique contributions are recognized, and your potential is limitless.

Other Duties

Please note this job description is not designed to cover or contain a comprehensive listing of activities, duties or responsibilities that are required of the employee. Duties, responsibilities, and activities may change at any time with or without notice. 

Application Note

In any materials you submit, you may redact or remove age-identifying information such as age, date of birth, or dates of school attendance or graduation. You will not be penalized for redacting or removing this information.

Skills Required

  • Experience designing, operating, and troubleshooting large-scale distributed systems in production
  • Deep knowledge of reliability engineering, observability, incident management, and production operations
  • Experience establishing SLIs, SLOs, actionable alerts, observability, and service ownership
  • Backend experience building systems and automation to reduce toil and improve operational efficiency
  • Experience leading high-severity incidents and improving incident response programs
  • Excellent written and verbal communication skills, including technical designs, runbooks, postmortems, and operational documentation
  • Experience with Python and Terraform, or equivalent automation and infrastructure-as-code tools
  • Experience with observability tools such as Datadog, New Relic, Grafana, or equivalent platforms
  • Experience operating production services in AWS and Kubernetes
  • Experience with CI/CD pipelines such as GitLab CI, ArgoCD, or GitOps workflows

Horizon3.ai Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Horizon3.ai and has not been reviewed or approved by Horizon3.ai.

  • Fair & Transparent Compensation — Pay is considered competitive across key roles, with employer language and role ranges positioning compensation as market-aligned. Feedback suggests compensation is viewed favorably, particularly in technical positions.
  • Equity Value & Accessibility — Equity is positioned as a core part of total rewards, with stock options broadly available to full-time employees. Feedback suggests equity and related programs enhance overall compensation.
  • Leave & Time Off Breadth — Time off is described as at least four weeks plus generous holidays and vacation packages globally. Feedback suggests this breadth of leave contributes to positive perceptions of total rewards.

Horizon3.ai Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: San Francisco, CA
107 Employees
Year Founded: 2019

What We Do

Horizon3.ai's mission is to help you find and fix attack vectors before attackers can exploit them. NodeZero, our autonomous penetration testing solution, enables organizations to continuously assess the security posture of their enterprise, including external, identity, on-prem, IoT, and cloud attack surfaces. Like APTs, ransomware, and other threat actors, our algorithms discover and fingerprint your attack surface, identifying the ways exploitable vulnerabilities, misconfigurations, harvested credentials, and dangerous product defaults can be chained together to facilitate a compromise. NodeZero is a true self-service SaaS offering that is safe to run in production and requires no persistent or credentialed agents. You will see your enterprise through the eyes of the attacker, identify your ineffective security controls, and ensure your limited resources are spent fixing problems that can actually be exploited. Founded in 2019 by industry, US Special Operations, and US National Security veterans, Horizon3.ai is headquartered in San Francisco, CA, and made in the USA.

Similar Jobs

GitLab Logo GitLab

Site Reliability Engineer

Cloud • Security • Software • Cybersecurity • Automation
Easy Apply
Remote
United States
2500 Employees

Replit Logo Replit

Site Reliability Engineer

Artificial Intelligence • Cloud • Machine Learning • Software • Database • App development • Generative AI
Remote
United States
300 Employees
250K-325K Annually

Fingerprint Logo Fingerprint

Site Reliability Engineer

Information Technology • Security • Software • Cybersecurity
Remote
USA
115 Employees
177K-240K Annually

Finalsite Logo Finalsite

Site Reliability Engineer

Edtech • Information Technology • Software
In-Office or Remote
The Center, IN, USA
563 Employees

Similar Companies Hiring

Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software
New York, New York
30 Employees
Vega Thumbnail
Artificial Intelligence • Automotive • Insurance • Transportation
US
43 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account