Staff Site Reliability Engineer

Posted Yesterday
Hiring Remotely in United States
Remote
235K-436K Annually
Senior level
Cloud • Security • Software • Cybersecurity
The Role
Leads the development of reliable, scalable, and observable SaaS platforms and SRE tooling. Builds reliability libraries, observability pipelines, progressive delivery systems, resilience automation, Kubernetes and infrastructure-as-code components, and incident-learning systems. Designs distributed multi-region services on Azure, leads complex incidents, drives cross-team reliability initiatives, mentors engineers, and establishes SLO, error-budget, architecture, and deployment standards.
Summary Generated by Built In

Veeam is the Data and AI Trust Company, specializing in helping organizations ensure their data and AI are fully understood, secured, and resilient to enable the acceleration of safe AI at scale. As the market leader in both data resilience and data security posture management, Veeam is built for the convergence of identity, data, security, and AI risk. Headquartered in Seattle with offices in more than 30 countries, Veeam protects over 550,000 customers worldwide, who trust Veeam to keep their businesses running. Join us as we go fearlessly forward together, growing, learning, and making a real impact for some of the world’s biggest brands.

About the Role

Veeam is launching a global Site Reliability Engineering (SRE) function to support the rollout and operation of our new SaaS offering: the Veeam Data Cloud. SRE at Veeam is a software engineering discipline that uses code, data, and systems thinking to make product delivery safe and fast. As a Staff Site Reliability Engineer, you’ll lead by building reliable-by-default platforms, tooling, and patterns that product teams adopt at scale. You will serve as a hands-on technical leader within the SRE team, guiding senior engineers, influencing product development teams, and ensuring the systems we operate are built to be reliable, scalable, and observable from the ground up.

You will drive strategic initiatives, mentor others in the practice of SRE, and help define architectural best practices across our platform. This role is pivotal in aligning teams, enforcing high standards, and scaling SRE principles globally within Veeam.

What You Will Build & Own

  • Reliability features as productized code: libraries, services, and controllers (e.g., deployment safety guards, rate‑limiters, circuit breakers, load‑shedding adapters, back‑pressure controls) that product teams import and extend.
  • Observability platform: define the data model and implement telemetry pipelines (metrics + logs + traces), SLI/SLOs, and error‑budget policies; ship SDKs/CLI/plugins that let teams declare SLOs in code and gate releases.
  • Change safety toolchain: progressive delivery primitives (canary, blue/green, feature flags), automated rollback, and release validation - delivered as reusable services/operators and CI/CD integrations.
  • Resilience automation: fault‑injection APIs, chaos experiments, traffic shadowing, and load/perf harnesses baked into pre‑prod and prod pipelines.
  • Golden‑path platform components: Terraform/Pulumi modules, Kubernetes operators, Helm charts, and reference microservice templates (authn/z, config, tenancy, telemetry) with paved‑road docs.
  • Incident learning systems: post‑incident automation (context capture, timeline, action tracking), and code changes that remove classes of failure (not just runbooks).

What you’ll do day‑to‑day

  • Write high‑quality code in one or more of: Go, TypeScript/Node.js, C#, or Java. Design APIs, write tests, and ship iteratively.
  • Lead designs for distributed, multi‑region services (initially on Azure) with a focus on failure modes, graceful degradation, and operability.
  • Partner with Staff/Principal peers across product and platform to align on reliability standards and drive cross‑team adoption.
  • Instrument systems deeply (metrics, logs, traces) and automate detection/response; keep alerting actionable.
  • Lead complex incidents during your daytime; drive blameless learning and land systemic fixes in code.
  • Mentor senior engineers; raise the bar via design reviews, ADRs, and pair programming.

Minimum qualifications

  • 8+ years in software engineering for cloud-based products; significant time designing/operating distributed systems at scale.
  • Strong proficiency in at least one backend language (C#, Java, Go, TypeScript/Node.js) and in writing production‑grade services/libraries.
  • Deep hands‑on with Kubernetes, IaC (Terraform or Pulumi), and CI/CD (e.g., GitHub Actions, GitLab, ArgoCD).
  • Practical observability expertise (metrics, tracing, logging) and experience turning SLOs/error budgets into engineering workflows.
  • Ability to lead cross‑team initiatives, influence architecture, and deliver measurable reliability outcomes.
  • Comfortable with a follow‑the‑sun on‑call model (8 hour daytime rotations) and coverage.

Nice to have

  • Built reliability platforms (SLO/SLO policy engines, progressive delivery, chaos/validation) used by multiple teams.
  • Multi‑cloud or advanced Azure networking/traffic management (cross‑region failover, DNS, gateway, service mesh).
  • Performance engineering at scale (workload modeling, cost/perf tradeoffs, regression detection).
  • Security/compliance aware delivery (SOC 2/ISO/SOx/FedRAMP patterns) as code.

How we work

  • Engineering > firefighting: we prevent repeated incidents by changing code, not adding toil.
  • Data‑driven: SLIs/SLOs and error budgets guide release decisions and reliability investments.
  • Enablement, not gatekeeping: we build paved roads and self‑service so product teams move faster and safer.
  • Blameless learning: incidents are inputs to design; we optimize for fast detection, fast rollback, and fast learning.

On‑call & working hours

  • Standard shifts, aligned to business hours in a 3×8h global rotation.
  • Rotations are designed for fairness (predictable schedules, protected focus time, and compensatory benefits per local policy).

Why Join Veeam?

  • Be a core architect in the rollout of Veeam’s first global SaaS offering—the Veeam Data Cloud.
  • Help shape a modern, engineering-driven SRE practice from the ground up.
  • Influence long-term reliability and architecture across a global product portfolio.
  • Work in a collaborative environment with engineering leaders who value strategic thinking, hands-on problem solving, and customer empathy.
  • Enjoy competitive pay and benefits, flexible work arrangements, and a team culture built on learning, ownership, and impact.

What you'll get

  • Unlimited paid time off, 12 paid holidays including 4 global VeeaMe Days for self-care and 24 paid volunteer hours annually through Veeam Cares
  • Paid parental leave: 8 weeks for all parents, 16 weeks for birthing parents
  • Medical, dental, and vision coverage starting on your first day
  • Mental health support, therapy sessions, and digital wellness tools via our Employee Assistance Program
  • 401(k) retirement plan with company matching contributions
  • Fertility, adoption, and surrogacy support through Maven, plus paid volunteer time
  • AirVet: 24/7 virtual veterinary care at no cost
  • Legal services, identity protection, and supplemental health insurance options
  • Tax-advantaged spending accounts for healthcare, dependent care, and commuting
  • Opportunities to learn and grow through on-demand libraries (LinkedIn Learning, O’Reilly), mentoring, workshops, and learning events like our annual Global Day of Learning

Compensation Transparency

Veeam is committed to pay transparency and equitable compensation. For this role, the compensation range below reflects the expected total target compensation (TTC), inclusive of base pay and a competitive performance-based bonus. For roles with a commission plan, the compensation range represents On Target Earnings (OTE), which includes base salary plus variable commission. When determining compensation, Veeam takes into consideration factors such as experience, education, skills, and geographic zone. Offers are typically made below the midpoint of the range.

In addition to compensation, Veeam provides a comprehensive benefits package, including health coverage, retirement plans, and unlimited time off.

U.S. Geographic Zones & Compensation Ranges (TTC / OTE)
Zone 1: San Francisco Bay Area, New York City Boroughs
$234,840—$436,080 USD
Zone 2: Washington, California (excluding San Francisco Bay Area)
$215,270—$399,740 USD
Zone 3: Texas, Illinois, North Carolina, Colorado, Massachusetts, Pennsylvania, Virginia, Oregon, Nevada, Hawaii, New York (excluding NYC boroughs); Sales roles located in Georgia, Ohio, and Arizona
$195,700—$363,400 USD
Zone 4: All other US locations
$170,259—$316,158 USD

Veeam Software is an equal opportunity employer and does not tolerate discrimination in any form on the basis of race, color, religion, gender, age, national origin, citizenship, disability, veteran status or any other classification protected by federal, state or local law. All your information will be kept confidential.

Personal data collected during the recruitment process will be processed in accordance with our Recruiting Privacy Notice, which explains how your information is collected, used, and handled in connection with hiring activities. By applying for this position, you consent to this processing. 

By submitting your application, you confirm that the information provided, including any supporting documents, is complete and accurate to the best of your knowledge. Any misrepresentation, omission, or falsification may result in disqualification from consideration or, if discovered after employment begins, termination of employment.

Skills Required

  • 8+ years of software engineering experience for cloud-based products
  • Significant experience designing and operating distributed systems at scale
  • Strong proficiency in at least one backend language: C#, Java, Go, or TypeScript/Node.js
  • Experience writing production-grade services and libraries
  • Hands-on experience with Kubernetes
  • Hands-on experience with infrastructure as code using Terraform or Pulumi
  • Experience with CI/CD tools such as GitHub Actions, GitLab, or ArgoCD
  • Practical observability expertise covering metrics, tracing, and logging
  • Experience applying SLOs and error budgets to engineering workflows
  • Ability to lead cross-team initiatives and influence architecture
  • Ability to deliver measurable reliability outcomes
  • Willingness to participate in a follow-the-sun on-call model with daytime rotations
  • Experience building reliability platforms used by multiple teams
  • Multi-cloud experience or advanced Azure networking and traffic management
  • Performance engineering experience at scale
  • Experience with security- and compliance-aware delivery, including SOC 2, ISO, SOx, or FedRAMP patterns

Veeam Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Veeam and has not been reviewed or approved by Veeam.

  • Healthcare Strength — Healthcare coverage is comprehensive with options that include employee-only no-cost tiers, plus mental-health support through an assistance program. Feedback suggests these offerings compare well in tech.
  • Leave & Time Off Breadth — Time off includes unlimited PTO in the U.S., paid company holidays, quarterly company-wide recharge days, and paid volunteer time. Feedback suggests team norms influence how fully this flexibility is utilized.
  • Strong & Reliable Incentives — Sales and pre-sales roles feature meaningful on-target earnings with competitive base and variable structures. Feedback suggests these plans provide strong upside for high performers.

Veeam Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Seattle, WA
4,172 Employees
Year Founded: 2006

What We Do

Veeam provides a single platform for modernizing backup, accelerating hybrid cloud and securing data. Veeam has 400,000+ customers worldwide, including 82% of the Fortune 500 and 69% of the Global 2,000. Veeam’s 100% channel ecosystem includes global partners, as well as HPE, NetApp, Cisco and Lenovo as exclusive resellers, and boasts more than 35K transacting partners worldwide.

Similar Jobs

GitLab Logo GitLab

Site Reliability Engineer

Cloud • Security • Software • Cybersecurity • Automation
Easy Apply
Remote
United States
2500 Employees

Horizon3.ai Logo Horizon3.ai

Site Reliability Engineer

Artificial Intelligence • Cybersecurity
Remote
US
107 Employees
200K-270K Annually

Replit Logo Replit

Site Reliability Engineer

Artificial Intelligence • Cloud • Machine Learning • Software • Database • App development • Generative AI
Remote
United States
300 Employees
250K-325K Annually

NinjaTrader Logo NinjaTrader

Site Reliability Engineer

Fintech • Software • Financial Services
Easy Apply
Remote or Hybrid
Chicago, IL, USA
430 Employees
160K-210K Annually

Similar Companies Hiring

Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Revel Thumbnail
Aerospace • Hardware • Robotics • Software
Marina Del Rey, California
60 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software
New York, New York
30 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account