Site Reliability Engineer

Posted 6 Days Ago
Santa Clara, CA, USA
In-Office
230K-250K Annually
Senior level
Software • Analytics
The Role
Build and lead the SRE function: define SLOs/SLIs/error budgets, own observability and incident response, embed reliability in SDLC, and scale the SRE team for a distributed SaaS platform.
Summary Generated by Built In

Forward is transforming how the world’s most complex networks are managed and secured. Founded in 2013 by four Stanford Ph.D.s, we built the industry’s first network digital twin — a mathematically precise model of the production network that gives IT teams unmatched visibility, verification, and agility across every major cloud and vendor environment.

Our customers include global leaders such as Goldman Sachs, PayPal, S&P Global, IBM, and Dell, as well as fast-growing enterprises and government agencies. According to IDC, Forward customers realize an average of $14.2 million in annual benefits through improved efficiency and security.

Backed by world-class investors including Andreessen Horowitz, Goldman Sachs, MSD Partners, and Threshold Ventures, Forward offers a people-centric, innovative culture where brilliant minds are shaping the future of network reliability, security, and AI-ready operations.
Forward is looking for a Site Reliability Engineer

About the Role This is not a "keep the lights on" SRE role. As our first or early SRE hire you will be building the reliability engineering function at Forward — defining how we think about availability, observability, incident response, and operational excellence across a complex, distributed SaaS platform. You will work closely with engineering, infrastructure, and product to ensure our platform meets the reliability bar our enterprise customers demand.

If you thrive in environments where you're handed a problem rather than a playbook this role is for you.

What You'll Own

  • Define and drive SRE practices from the ground up — SLOs, SLIs, error budgets, and the frameworks the engineering org will actually use
  • Drive the reliability and operational excellence of the Forward SaaS platform
  • Build and maintain observability infrastructure — logging, metrics, tracing, and alerting — so the team always knows what's happening before customers do
  • Lead incident response: on-call rotations, runbooks, post-mortems, and the follow-through to make sure the same incident doesn't happen twice
  • Partner with engineering teams to embed reliability thinking into the SDLC — capacity planning, load testing, chaos engineering, and production readiness reviews
  • Help define and build the SRE team as the company scales — this is a foundational hire with a path to leadership

What We're Looking For

  • 6+ years of experience in site reliability engineering, DevOps, or infrastructure engineering in a SaaS or cloud environment
  • Proven experience building or significantly maturing an SRE function — not just operating within one someone else built
  • Strong fundamentals in networking — TCP/IP, DNS, routing, switching, firewalls, and load balancing. Experience with network management or observability platforms is a significant plus
  • Hands-on experience with Kubernetes and container orchestration in production environments
  • Deep proficiency with observability tooling — Prometheus, Grafana, Datadog, Splunk, or similar
  • Strong scripting and automation skills in Python, Bash, or similar
  • Experience with cloud platforms — AWS, GCP, or Azure — including infrastructure as code (Terraform, Ansible, or equivalent)
  • Track record of owning and improving incident response processes including blameless post-mortems and SLO-driven reliability improvements
  • Ability to communicate clearly with both engineering teams and non-technical stakeholders — you can explain an outage to a customer-facing team without jargon and explain an SLO to an executive without losing them

Nice to Have

  • Experience supporting enterprise or federal government customers with high availability requirements
  • Experience in a foundational or early SRE hire capacity at a growth stage company

What This Role Is Not

  • A pure ops or NOC role — you are building and engineering, not just monitoring
  • A siloed function — you will be deeply embedded with product and engineering teams
  • A ticket-taker — you will be proactively identifying and solving reliability problems before they become incidents
    The base pay range for this role is between $230,000 and $250,000. Base pay will depend on your skills, qualifications, experience, and location

Skills Required

  • 6+ years of experience in site reliability engineering, DevOps, or infrastructure engineering in a SaaS or cloud environment.
  • Proven experience building or significantly maturing an SRE function.
  • Strong fundamentals in networking (TCP/IP, DNS, routing, switching, firewalls, load balancing).
  • Hands-on experience with Kubernetes and container orchestration in production environments.
  • Deep proficiency with observability tooling such as Prometheus, Grafana, Datadog, or Splunk.
  • Strong scripting and automation skills in Python, Bash, or similar.
  • Experience with cloud platforms (AWS, GCP, or Azure) including infrastructure as code (Terraform, Ansible, or equivalent).
  • Track record of owning and improving incident response processes including blameless post-mortems and SLO-driven reliability improvements.
  • Ability to communicate clearly with both engineering teams and non-technical stakeholders.
  • Experience supporting enterprise or federal government customers with high availability requirements.
  • Experience as an early or foundational SRE hire at a growth-stage company.

Forward Networks Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Forward Networks and has not been reviewed or approved by Forward Networks.

  • Fair & Transparent Compensation Pay is considered competitive for a mid-stage infrastructure/software company across several roles and locations. Signals point to strong totals in technical and select go-to-market positions.
  • Equity Value & Accessibility Equity is broadly offered to all employees, creating ownership potential alongside salary and bonus. As a private company, perceived value can rise with company performance and future liquidity.
  • Healthcare Strength Medical, dental, and vision coverage is described as top-grade for employees and dependents. Company materials consistently highlight strong core health benefits across hiring channels.

Forward Networks Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Palo Alto, CA
70 Employees
Year Founded: 2013

What We Do

The future of network operations is network modeling. Forward Networks' flagship platform Forward Enterprise gives users a mathematically accurate network digital twin. Forward enables perfect network visibility, full path analysis, security policy verification, and change prediction, freeing up time and saving you money.

Similar Jobs

General Motors Logo General Motors

Senior Engineer

Automotive • Big Data • Information Technology • Robotics • Software • Transportation • Manufacturing
Hybrid
2 Locations
165000 Employees
148K-222K Annually

ServiceNow Logo ServiceNow

Site Reliability Engineer

Artificial Intelligence • Cloud • HR Tech • Information Technology • Productivity • Software • Automation
Hybrid
Santa Clara, CA, USA
29000 Employees
166K-290K Annually

DraftKings Logo DraftKings

Site Reliability Engineer

Digital Media • Gaming • Information Technology • Software • Sports • Esports • Big Data Analytics
Remote or Hybrid
United States
6400 Employees
200K-250K Annually

DISQO Logo DISQO

Senior Site Reliability Engineer

AdTech • Big Data • Cloud • Marketing Tech • Software • Analytics
Easy Apply
Hybrid
Los Angeles, CA, USA
272 Employees
170K-190K Annually

Similar Companies Hiring

Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
Kepler  Thumbnail
Fintech • Software
New York, New York
6 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account