Staff Site Reliability Engineer

Posted Yesterday
Hiring Remotely in USA
Remote
177K-240K Annually
Expert/Leader
Information Technology • Security • Software • Cybersecurity
Our mission is to empower engineering teams to build safe and seamless online services.
The Role
Lead reliability strategy across eight engineering groups by defining SLIs, SLOs, and error budgets; strengthening incident response, alerting, postmortems, change safety, and failure testing; and coaching teams to own reliability. The role remains hands-on through production investigations, tooling, dashboards, and reference implementations. It also leads AI adoption in incident management and observability while partnering with architecture, platform, and product teams to improve distributed-system resilience.
Summary Generated by Built In

Fingerprint empowers enterprises to detect and stop online fraud with the world’s most accurate device intelligence.  We lead our industry with bleeding-edge identification capabilities and work on turning new ideas and discoveries in the fraud detection space into reality. Our customers range from innovative startups to leading enterprise companies, including Plaid, Dropbox, and Booking.com. 

Fingerprint is a globally dispersed, 100% remote company. We were named on on the 2026 Forbes Best Startup Employers list and ranked #803 on the 2026 Inc. 5000 list of America’s fastest-growing private companies. 

We have raised $77M and are backed by Craft Ventures (Tesla, Facebook, Airbnb ), Nexus Venture Partners ( Postman, Apollo.io, MinIO, Druva) and Uncorrelated Ventures ( Redis, Rollbar,  Gradle).



About the role

You will be Fingerprint's first dedicated Site Reliability Engineer. You will work alongside our Architect on the shape of the platform, with Cloud Platform on the infrastructure that runs it, and with every product team on how they operate what they own — without direct reports. The mandate has three parts. First, make reliability measurable: define SLIs and SLOs for our critical paths, get teams to own them, and make error budgets the shared language for prioritizing reliability against features. Second, raise the operational bar: strengthen incident response, postmortem quality, alerting, and change safety so that we find out first and repeat incidents stop repeating. Third, build the mindset: coach teams to design for failure, test for it deliberately, and treat operability as part of done — so the practices outlive your involvement in any single team.

You report directly to the VP of Engineering. That placement is deliberate: reliability standards need to apply evenly across eight engineering groups, and you need the neutrality to hold every team, including infrastructure, to the same bar.

What you'll do

Make reliability measurable

  • Define SLIs and SLOs for Fingerprint's critical request paths (identification, events, server APIs, client agents) with the teams that own them; make them visible, reviewed, and tied to decisions.
  • Introduce error budgets as the mechanism for balancing reliability investment against feature work, and coach EMs and Staff engineers on using them.
  • Own the reliability metrics that leadership uses to judge progress; be the No Nonsense voice on whether we are actually getting better.

Raise the operational bar

  • Strengthen the incident lifecycle end to end: detection, response, communication, postmortem quality, and follow-up completion. Make the postmortem the most useful document a team writes.
  • Close the "customers find out before we do" gap: drive alert quality, correctness anomaly detection, and escalation design across teams, working with Cloud Platform on shared tooling.
  • Lead reliability reviews for high-risk changes and new services (production readiness, capacity, failure modes, rollback), teaching through review rather than gatekeeping.
  • Introduce deliberate failure testing (game days, chaos exercises) in staging first, then production, to discover gaps and safe limits before customers do.

Build the SRE mindset in teams

  • Embed with teams for time-boxed engagements: pair on their hardest reliability problems, leave behind better practices and a stronger owner, then move on.
  • Develop Staff and Lead engineers as reliability leaders in their own groups — the goal is that every team has someone who thinks like an SRE.
  • Codify practices that stick: production readiness checklists, on-call standards, runbook quality, change safety norms. Make them lightweight enough that teams choose to use them.
  • Partner with the Architect and tech leads so reliability is designed in, not retrofitted.

Stay hands-on

  • Dig into production during incidents and investigations. Write tooling, dashboards, and reference implementations. Be credible with the engineers you are asking to change how they work.
  • Lead AI adoption in reliability practice: set norms for AI-assisted incident investigation, postmortem analysis, runbook authoring, and observability tooling across teams, and shape our runbooks, alerts, and operational data so AI agents can safely help diagnose and operate our systems alongside engineers.
What we're looking for
  • 10+ years of engineering experience, with 3+ years as an SRE, production engineer, or reliability-focused Staff engineer operating across multiple teams — you have owned reliability for a platform, not just for a service you built.
  • Deep experience with SLI/SLO design and error budgets in practice, including the hard part: getting product teams to adopt and act on them.
  • Strong incident leadership: you have run incident response and postmortems for high-severity, customer-facing incidents and materially improved how an organization learns from them.
  • Hands-on depth in distributed systems failure modes — cache/database saturation and cascading failure, retry storms, capacity limits, degradation and load shedding — in a high-throughput, low-latency environment. Fluent in Kubernetes, AWS, and modern observability tooling (Datadog or equivalent).
  • Comfortable reading and writing production code (Go, TypeScript, or similar) and infrastructure as code. You can ship a fix, not just recommend one.
  • Track record of leading through influence: you have changed how teams you did not manage operate, and can explain how adoption actually happened.
  • Teacher's instinct. You have coached engineers into owning reliability and can point to practices that persisted after you stepped back.
  • Exceptional written communication. You make incidents, risks, and trade-offs legible to engineers and executives alike, and you default to async, documented decision-making.
  • AI-native by default. You use AI tools as a normal part of how you investigate incidents, analyze telemetry, write runbooks and postmortems, and build tooling — and you have opinions, from experience, about where they accelerate reliability work and where they don't yet.
  • Operate for an AI-assisted org. You think about how runbooks, alerts, dashboards, and operational data should be structured so that both humans and AI agents can diagnose and act on them safely — legible signals, clear ownership, strong guardrails on automated change.
  • Pragmatism over purity. You know that reliability competes with delivery, and you can make the case for the right investment at the right time — and say when a risk is acceptable.

Nice to have

  • Experience in fraud detection, identity, payments, or other adversarial, real-time domains.
  • Multi-region, cell-based, or failure-isolation architecture experience.
  • Experience with Elasticsearch, Redis, DynamoDB, or Kafka at scale, including their failure modes.
  • Familiarity with FinOps and the reliability/cost trade-off in cloud infrastructure.

Compensation & Transparency

For US-based employees, the cash compensation range for this role is $177,000 – $240,000. We set standard ranges for all US roles based on function, level, and geographic location, benchmarked against similar stage growth companies. To comply with local legislation and provide greater transparency, we share salary ranges on all job postings. However, these ranges are specific to the hiring location and may differ within or outside the US. Offers vary depending on, but not limited to, relevant experience, education, certifications/licenses, skills, training, and market conditions.

Due to regulatory and security reasons, there's a small number of countries where we cannot have Fingerprint teammates based. Additionally, because Fingerprint is an all-remote company and people can join our workforce from almost any country, we do not sponsor visas. Fingerprint teammates need to be authorized to work from their home location.

We are dedicated to creating an inclusive work environment for everyone. We embrace and celebrate the unique experiences, perspectives and cultural backgrounds that each employee brings to our workplace. Fingerprint strives to foster an environment where our employees feel respected, valued and empowered, and our team members are at the forefront in helping us promote and sustain an inclusive workplace. We highly encourage people from underrepresented groups in tech to apply.

If you are applying as a resident of California, please read our CCPA notice here.

If you are applying as a resident of the EU, please read our GDPR notice here.

  • We have noticed a rise in recruiting impersonations across the industry, where scammers attempt to access candidates' personal and financial information through fake interviews and offers. All Fingerprint recruiting email communications will always come from the @fingerprint.com domain. Any outreach claiming to be from Fingerprint via other sources should be ignored.*

Due to regulatory and security reasons, there’s a small number of countries where we cannot have Fingerprint teammates based. Additionally, because Fingerprint is an all-remote company and people can join our workforce from almost any country, we do not sponsor visas. Fingerprint teammates need to be authorized to work from their home location.

We are dedicated to creating an inclusive work environment for everyone. We embrace and celebrate the unique experiences, perspectives and cultural backgrounds that each employee brings to our workplace. Fingerprint strives to foster an environment where our employees feel respected, valued and empowered, and our team members are at the forefront in helping us promote and sustain an inclusive workplace. We highly encourage people from underrepresented groups in tech to apply.

If you are applying as a resident of California, please read our CCPA notice here.

If you are applying as a resident of the EU, please read our GDPR notice here.

**We have noticed a rise in recruiting impersonations across the industry, where scammers attempt to access candidates' personal and financial information through fake interviews and offers. All Fingerprint recruiting email communications will always come from the @fingerprint.com domain. Any outreach claiming to be from Fingerprint via other sources should be ignored.

Skills Required

  • 10+ years of engineering experience
  • 3+ years as an SRE, production engineer, or reliability-focused Staff engineer operating across multiple teams
  • Experience owning reliability for a platform across multiple teams
  • Deep experience designing and implementing SLIs, SLOs, and error budgets
  • Experience driving product-team adoption of reliability metrics and error budgets
  • Experience leading incident response and postmortems for high-severity, customer-facing incidents
  • Experience improving incident learning and follow-up practices
  • Hands-on experience with distributed-systems failure modes, including saturation, cascading failure, retry storms, capacity limits, degradation, and load shedding
  • Experience in high-throughput, low-latency environments
  • Fluency with Kubernetes, AWS, and modern observability tooling such as Datadog
  • Ability to read and write production code in Go, TypeScript, or similar languages
  • Experience with infrastructure as code
  • Track record of leading through influence without direct authority
  • Experience coaching engineers to own reliability practices
  • Exceptional written communication and documented decision-making skills
  • Hands-on use of AI tools for incident investigation, telemetry analysis, runbooks, postmortems, and tooling
  • Ability to design operational data and guardrails for safe AI-assisted diagnosis and operations
  • Pragmatic judgment balancing reliability investment and delivery
  • Experience in fraud detection, identity, payments, or other adversarial real-time domains
  • Multi-region, cell-based, or failure-isolation architecture experience
  • Experience with Elasticsearch, Redis, DynamoDB, or Kafka at scale
  • Familiarity with FinOps and cloud infrastructure reliability-cost tradeoffs

Fingerprint Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Fingerprint and has not been reviewed or approved by Fingerprint.

  • Leave & Time Off Breadth Time off includes flexible or unlimited PTO with a minimum vacation target, plus paid holidays and sick time. Policies emphasize rest and work–life balance in a remote setting.
  • Parental & Family Support Parental leave is fully paid for maternity, paternity, and non-birth parents and is described as generous. This coverage aims to support major life events for families.
  • Flexible Benefits The package supports a remote-first, globally distributed model with flexible hours and asynchronous work. Perks like a home-office stipend, a provided MacBook, and annual meetups complement core benefits.

Fingerprint Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Chicago, Illinois
115 Employees
Year Founded: 2019

What We Do

Fingerprint empowers developers to stop online fraud at the source. We work on turning radical new ideas in the fraud detection space into reality. Our products are developer-focused and our clients range from solo developers to publicly traded companies. Some of our customers include Coinbase, Booking.com, Yahoo, and eBay just to name a few. We are a globally dispersed, 100% remote company with a strong open-source focus. Our flagship open source project is FingerprintJS (16K stars on GitHub). We have raised $77M and are backed by Craft Ventures (previously invested in Tesla, Facebook, Airbnb), Nexus VP (previously invested in Postman & Hasura), and Uncorrelated Ventures (previously invested in Redis, Rollbar & Gradle).

Why Work With Us

Fingerprint offers innovation in a global remote team, a nurturing culture of growth, work-life balance, and impactful projects. Come be part of our exciting journey!

Gallery

Gallery

Similar Jobs

GitLab Logo GitLab

Site Reliability Engineer

Cloud • Security • Software • Cybersecurity • Automation
Easy Apply
Remote
United States
2500 Employees

NBCUniversal Logo NBCUniversal

Site Reliability Engineer

AdTech • Cloud • Digital Media • Information Technology • News + Entertainment • App development
Remote or Hybrid
New York, NY, USA

Anthropic Logo Anthropic

Site Reliability Engineer

Artificial Intelligence • Natural Language Processing • Generative AI
In-Office or Remote
3 Locations
2500 Employees
320K-485K Annually

Finalsite Logo Finalsite

Site Reliability Engineer

Edtech • Information Technology • Software
In-Office or Remote
The Center, IN, USA
563 Employees

Similar Companies Hiring

Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Revel.io Thumbnail
Aerospace • Hardware • Robotics • Software
US
50 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software • Productivity
US
15 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account