Senior Reliability Engineer - AV Labs

Posted 5 Days Ago
Be an Early Applicant
Sunnyvale, CA, USA
In-Office
180K-200K Annually
Senior level
Logistics • Transportation • 3PL: Third Party Logistics
We reimagine the way the world moves for the better.
The Role
Own sensor and hardware reliability for Uber’s distributed autonomous vehicle fleet. Design observability platforms, telemetry ingestion, alerting, failure detection, SLIs/SLOs, and automated mitigation systems across edge devices and fleet infrastructure. Diagnose hardware, software, networking, and environmental failures while driving technical strategy, cross-team architecture reviews, and proactive analytics to improve sensor uptime, data yield, and operational efficiency.
Summary Generated by Built In
About the Role

We are looking for a hardware focused Senior Reliability Engineer to focus on sensor and hardware system reliability, owning the observability, alerting, and automation that ensures Uber’s in-vehicle sensor data collection systems operate reliably at scale.

This role is centered on maximizing sensor uptime, data yield, and supply hours across a large, geographically distributed fleet. You will design systems that determine when to react to issues impacting data recording capability, whether caused by failing sensors, degraded onboard computers, software regressions, or systemic environmental factors. 

As the technical owner for sensor reliability and observability, you will build the infrastructure that converts low-level signals into actionable intelligence and automated responses. This is a senior role requiring strong software engineering fundamentals, deep systems thinking, and the ability to drive cross-team technical direction without direct authority.

What You Will Do
  • Architect Observability Systems: Design and scale an observability platform capable of ingesting and analyzing real-time health telemetry from thousands of distributed vehicle nodes.
  • Build for Edge Constraints: Develop systems that remain performant despite hardware diversity, intermittent connectivity, and rapid fleet scaling.
  • Define Criticality Models: Establish alerting strategies that distinguish transient anomalies from systemic issues impacting sensor uptime and data yield.
  • Detect Complex Failure Modes: Design detection logic for "silent" failures, such as sensor degradation, compute saturation, or recording pipeline stalls.
  • Scale Through Automation: Design automated detection, triage, and mitigation mechanisms to eliminate manual intervention as the fleet grows.
  • Partner on Mitigation: Collaborate with Operations and Engineering to build safe, automated responses to recurring hardware and software failure scenarios.
  • Drive Operational Efficiency: Build technical interfaces to help Operations surface issues and Engineering diagnose and deploy mitigations rapidly (TTD/TTM).
  • Lead Technical Strategy: Drive reliability-focused design reviews and translate operational pain points into concrete technical requirements and roadmaps.
  • Uncover Proactive Insights: Apply advanced data analytics to identify latent patterns in fleet telemetry, enabling the proactive detection of systemic regressions and hardware degradation before they impact operations.
Basic Qualifications
  • 5+ years of relevant industry experience in software engineering, site reliability, or systems engineering
  • Distributed Systems: Experience with modern observability platforms (e.g., Prometheus, Grafana, ELK) in edge, IoT, or hardware-integrated environments.
  • Language Proficiency: coding skills in one or more of Go, Python, or C++, with experience building and operating production systems.
  • Systems Expertise: Proficiency in Linux internals and shell scripting for triaging and debugging edge devices or hardware-adjacent systems.
  • Engineering Fundamentals: Ability to debug across services, containers (Docker), and networking stacks.
  • Reliability Experience: Proven track record owning reliability, infrastructure, or platform systems for large-scale production workloads.
  • Observability Tooling: Experience designing and operating observability systems (metrics, logging, alerting, and dashboards).
  • Metrics-Driven: Experience defining and implementing SLIs and SLOs for system availability or data yield.
  • Networking Knowledge: Deep understanding of networking protocols (TCP/IP, gRPC, or MQTT) and data handling in bandwidth-constrained environments.
  • Leadership: Experience driving complex technical projects and architectural reviews across multiple teams from design through production.
Preferred Qualifications
  • Experience with modern observability platforms (e.g., Prometheus, Grafana, ELK) in edge, IoT, or hardware-integrated environments.
  • Knowledge of sensor data protocols (e.g., Camera, LiDAR, Radar) or hardware-to-cloud data ingestion pipelines.
  • Experience with "Grey Failure" detection and management in complex, distributed systems.
  • Proven track record in 'Fleet Health' for large-scale hardware deployments (e.g., cloud infrastructure, global server fleets, or industrial IoT) where automation was used to replace manual intervention.
~~ ~~ Responsibilities

For Sunnyvale, CA-based roles: The base salary range for this role is USD $180,000 per year - USD $200,000 per year.

You will be eligible to participate in Uber's bonus program, and may be offered an equity award & other types of comp. All full-time employees are eligible to participate in a 401(k) plan. You will also be eligible for various benefits.

About Us

Ready to Ride?

This isn't the kind of place where you follow a playbook — it's where you help write one. If you're driven by impact, energized by challenge, and ready to shape how the world moves — we'd love to hear from you.

You may be eligible for bonuses, equity, and other compensation, as well as a range of benefits. Explore our benefits.

Offices remain key to collaboration and Uber's culture. Unless approved for full remote work, employees must spend at least 50% of their time in-office. Some roles, like those at greenlight hubs, require full-time in-office presence. Ask your Recruiter for details about this role's requirements.

Uber is proud to be an Equal Opportunity employer. All qualified applicants will receive consideration for employment without regard to sex, gender identity, sexual orientation, race, color, religion, national origin, disability, protected Veteran status, age, or any other characteristic protected by law. We also consider qualified applicants regardless of criminal histories, consistent with legal requirements. If you have a disability or special need that requires accommodation, please let us know by completing this form.

Skills Required

  • 5+ years of relevant industry experience in software engineering, site reliability, or systems engineering
  • Experience with modern observability platforms such as Prometheus, Grafana, or ELK in edge, IoT, or hardware-integrated environments
  • Production coding experience in Go, Python, or C++
  • Proficiency in Linux internals and shell scripting for edge-device or hardware-adjacent debugging
  • Ability to debug across services, Docker containers, and networking stacks
  • Experience owning reliability, infrastructure, or platform systems for large-scale production workloads
  • Experience designing and operating metrics, logging, alerting, and dashboard systems
  • Experience defining and implementing SLIs and SLOs for availability or data yield
  • Deep understanding of TCP/IP, gRPC, or MQTT and data handling in bandwidth-constrained environments
  • Experience driving complex technical projects and architectural reviews across multiple teams from design through production
  • Knowledge of camera, LiDAR, or radar sensor data protocols or hardware-to-cloud ingestion pipelines
  • Experience with grey failure detection and management in complex distributed systems
  • Fleet health experience for large-scale hardware deployments using automation to replace manual intervention

Uber Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Uber and has not been reviewed or approved by Uber.

  • Parental & Family Support Policies provide a minimum of fully paid parental leave for all parents and financial support for fertility, adoption, and surrogacy, with added credits to ease the transition. Programs extend to family medical leave and parenting support resources, indicating depth beyond baseline offerings.
  • Healthcare Strength Healthcare coverage is described as comprehensive across many countries, with medical, dental, vision, life, disability, and mental health benefits, plus allowances where direct plans are not available. Wellness programs and reimbursements further reinforce access to care.
  • Wellbeing & Lifestyle Benefits Monthly ride and meal credits, free office meals/snacks, fitness stipends, onsite gyms, and wellbeing reimbursements create meaningful everyday value. Home‑office stipends, travel medical coverage, and counseling support round out lifestyle-oriented perks.

Uber Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: San Francisco, CA
21,000 Employees
Year Founded: 2009

What We Do

We are Uber. The go-getters. The kind of people who are relentless about our mission to help people go anywhere and get anything. Movement is what we do. It’s our lifeblood. It runs through our veins. It’s what gets us out of bed each morning. It pushes us to constantly reimagine how we can move better. For you. For all the places you want to go. For all the things you want to get. For all the ways you want to earn. Across the entire world. In real-time. At the incredible speed of now.

Why Work With Us

We welcome people from all backgrounds who seek the opportunity to help build a future where everyone and everything can move independently. If you have the curiosity, passion, and collaborative spirit, work with us, and let’s move the world forward, together.

Gallery

Gallery

Similar Jobs

Shield AI Logo Shield AI

Staff Engineer

Aerospace • Artificial Intelligence • Machine Learning • Robotics • Software
In-Office
2 Locations
142K-220K Annually

Shield AI Logo Shield AI

Engineer I, Electromechanical (R6042)

Aerospace • Artificial Intelligence • Machine Learning • Robotics • Software
In-Office
2 Locations
98K-122K Annually

ServiceNow Logo ServiceNow

Full-stack Engineer

Artificial Intelligence • Cloud • HR Tech • Information Technology • Productivity • Software • Automation
Hybrid
San Diego, CA, USA
29000 Employees
150K-262K Annually

ServiceNow Logo ServiceNow

Director, Enterprise Sales - Life Sciences, Bay Area

Artificial Intelligence • Cloud • HR Tech • Information Technology • Productivity • Software • Automation
Remote or Hybrid
Santa Clara, CA, USA
29000 Employees
171K-282K Annually

Similar Companies Hiring

Toro TMS Thumbnail
Cloud • Enterprise Web • Sales • Software • Transportation
Chicago, IL
80 Employees
Axle Health Thumbnail
Artificial Intelligence • Healthtech • Information Technology • Logistics
Santa Monica, CA
25 Employees
Vega Thumbnail
Artificial Intelligence • Automotive • Insurance • Transportation
US
43 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account