Site Reliability Engineer

Posted 2 Days Ago
New York City, NY, USA
Hybrid
Mid level
Software • Analytics • Cybersecurity
The Role
Own production health for Cellebrite’s cloud platform through incident response, monitoring and observability improvements, rolling upgrades, change management, and on-call support. Investigate production issues, coordinate incident response, improve alerting and runbooks, collaborate with support and R&D teams, and apply AI/ML tools to anomaly detection and incident triage.
Summary Generated by Built In
Description

About Cellebrite: 

Cellebrite’s (Nasdaq: CLBT) mission is to enable its global customers to protect and save lives by enhancing digital investigations and intelligence gathering to accelerate justice in communities around the world. Cellebrite’s AI-powered Digital Investigation Platform enables customers to lawfully access, collect, analyze and share digital evidence in legally sanctioned investigations while preserving data privacy. Thousands of public safety organizations, intelligence agencies and businesses rely on Cellebrite’s digital forensic and investigative solutions—available via cloud, on-premises and hybrid deployments—to close cases faster and safeguard communities.

To learn more, visit us at www.cellebrite.com, https://investors.cellebrite.com/investors and find us on social media @Cellebrite. 


What is your mission? 

As a Site Reliability Engineer, you’ll be a key technical owner of production health for Cellebrite’s cloud platform the platform investigators and public safety agencies depend on every day. You’ll work closely with our TCS support team and R&D to triage, investigate, and resolve production issues quickly and thoroughly, while helping modernize how we monitor the platform, roll out changes safely, and respond when things break. The focus of this role is keeping production healthy and continuously raising the bar on how we operate it — not building new infrastructure from scratch. Your work directly supports Cellebrite’s mission to protect lives, accelerate justice, and preserve data privacy. 

Responsibilities: 

Incident Response & Production Health 

  • Own the full incident lifecycle detect, triage by severity/impact, investigate, and drive to resolution engaging TCS, R&D, and DevOps as needed. 
  • Act as a technical responder during major incidents and grow into an incident-commander role coordinating the response across teams in real time. 
  • Lead blameless post-incident reviews and make sure corrective actions actually get implemented, not just documented. 
  • Track recurring issues and support-ticket trends; distinguish patterns that need a permanent fix from one-off noise, and route the former to R&D. 

AI-Driven Monitoring & Observability 

  • Evolve production monitoring toward AI-assisted operations anomaly detection, cross-signal correlation across logs/metrics/traces, and LLM-based triage assistants that cut time-to-diagnosis. 
  • Continuously tune dashboards, alert thresholds, and routing so real issues surface fast and noise doesn’t drown them out. 
  • Evaluate and pilot AI/ML-based observability tooling, and champion adoption of tools like Copilot or log-analysis assistants across the team. 

Rolling Upgrades & Change Management 

  • Plan and execute rolling upgrades, version updates, and patches to production services and infrastructure with zero or minimal downtime. 
  • Partner with R&D and DevOps to define safe rollout and rollback strategies for deployments, and validate system health post-upgrade. 
  • Participate in an on-call rotation to help maintain production uptime SLOs, with upgrade windows and rollback readiness built into the runbooks you own. 

Cross-Team Enablement 

  • Partner with TCS on production tickets, providing deeper technical investigation when issues exceed their level; partner with R&D on code-level fixes, deployments, and architectural input. 
  • Own and continuously improve runbooks and the known-issues knowledge base so TCS resolves more independently over time. 
  • Maintain clear, current documentation of production architecture, known issues, and resolution paths for both TCS and R&D.
Requirements

Requirements: 

  • 3–5 years of experience in a production support, SRE, or operations role. 
  • Solid working experience with AWS (troubleshooting and operating existing infrastructure; deep infra design experience is not required). 
  • Experience with monitoring/observability tools (e.g., Datadog, CloudWatch, Grafana) able to read dashboards, tune alerts, and investigate from logs/metrics/traces. 
  • Working knowledge of Linux system administration and basic networking troubleshooting. 
  • Familiarity with containerized environments (Kubernetes) — enough to check pod health, logs, and restart/rollback safely. 
  • Comfortable with scripting (Python or Bash) to automate repetitive troubleshooting or reporting tasks. 
  • Experience driving rolling upgrades, patching, or zero-downtime deployment processes in a production environment. 
  • Experience collaborating across support and engineering teams — comfortable working with an outsourced support team (e.g., TCS) on one side and R&D engineers on the other, translating between the two. 
  • Strong communication skills in English, written and verbal — clear handoffs and documentation matter as much as fixing things. 
  • Nice to have: exposure to Terraform or other IaC (read/troubleshoot, not necessarily author), basic CI/CD familiarity, and genuine interest in applying AI/ML to monitoring, alerting, and incident triage — not just using AI coding assistants. 

Skills Required

  • 3-5 years of experience in production support, SRE, or operations
  • Experience troubleshooting and operating existing AWS infrastructure
  • Experience with monitoring and observability tools such as Datadog, CloudWatch, or Grafana
  • Working knowledge of Linux system administration and basic networking troubleshooting
  • Familiarity with Kubernetes and containerized environments
  • Scripting experience with Python or Bash
  • Experience with rolling upgrades, patching, or zero-downtime production deployments
  • Experience collaborating across support and engineering teams
  • Strong written and verbal English communication skills
  • Exposure to Terraform or other infrastructure as code
  • Basic CI/CD familiarity
  • Interest or experience applying AI/ML to monitoring, alerting, and incident triage
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Petah Tikva
1,173 Employees
Year Founded: 1999

What We Do

Cellebrite is the leader in digital intelligence and investigative analytics, partnering with public and private organizations to transform how they manage data in investigations to accelerate justice and ensure data security.

Similar Jobs

In-Office or Remote
2 Locations
175633 Employees
85K-193K Annually

Federal Reserve Bank of Boston Logo Federal Reserve Bank of Boston

Site Reliability Engineer

Fintech • Information Technology • Payments • Sharing Economy • Financial Services • Cryptocurrency
In-Office
11 Locations
1200 Employees
147K-234K Annually

MetLife Logo MetLife

Site Reliability Engineer

Fintech • Information Technology • Insurance • Financial Services • Big Data Analytics
Remote or Hybrid
United States
43000 Employees
111K-180K Annually

PwC Logo PwC

Site Reliability Engineer

Artificial Intelligence • Professional Services • Business Intelligence • Consulting • Cybersecurity • Generative AI
Hybrid
58 Locations
370000 Employees
151K-187K Annually

Similar Companies Hiring

Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Revel Thumbnail
Aerospace • Hardware • Robotics • Software
Marina Del Rey, California
60 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software
New York, New York
30 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account