Staff Site Reliability Engineer

Posted 10 Days Ago
Be an Early Applicant
Santa Clara, CA, USA
Hybrid
163K-214K Annually
Senior level
Artificial Intelligence • Hardware • Software • Quantum Computing
Building the world’s best quantum computers to solve the world’s most complex problems.
The Role
Lead reliability strategy and standards across regions and services. Design and operate observability and SLO programs, run incident command, execute chaos and disaster-recovery tests, build resilience for stateful/streaming platforms, develop AIOps-driven remediation and self-healing, co-own cloud security posture, and mentor engineers to scale reliability practices.
Summary Generated by Built In

About IonQ: 

IonQ, Inc. [NYSE: IONQ] is the world’s leading quantum platform and merchant supplier - delivering integrated quantum solutions across computing, networking, sensing, and security. IonQ’s newest generation of quantum computers, the IonQ Tempo, is the latest in a line of cutting-edge systems that have been helping customers and partners including Amazon Web Services, and AstraZeneca achieve 20x performance results and accelerate innovation in drug discovery, materials science, financial modeling, logistics, cybersecurity, and defense. In 2025, the company achieved 99.99% two-qubit gate fidelity, setting a world record in quantum computing performance.
Headquartered in College Park, Maryland, IonQ has operations in California, Colorado, Massachusetts, Tennessee, Washington, Italy, South Korea, Sweden, Switzerland, Canada, and the United Kingdom. Our quantum computing services are available through all major cloud providers, while we also meet the needs of networking and sensing customers across land, sea, air, and space. IonQ is making quantum platforms more accessible and impactful than ever before.  

Location: Santa Clara, CA
Travel: Up to 25%
Job ID: 1739

The Role: 

We are seeking a Staff Site Reliability Engineer. As Staff SRE Engineer, you set the technical direction for reliability across regions and services. You own the reliability strategy, define the standards and mechanisms that guide production operations, and raise the bar through design leadership, operational discipline, and mentorship. You remain deeply hands-on by designing and operating observability platforms, defining and governing SLO programs, leading high-severity incident response, building resilience and disaster-recovery automation, improving reliability of stateful and streaming platforms, and creating AI Ops workflows for triage, remediation, and self-healing.

Responsibilities:

  • Production reliability: own service-level objectives, error budgets, and production reliability outcomes end to end, and represent reliability in architecture and scaling decisions.
  • Engineer observability: design and operate the observability stack so production services are fully instrumented and define the standards platform and application teams follow.
  • Govern SLOs and error budgets: define and manage service-level objectives, run regular reviews with service owners, and drive corrective action when services consume error budgets unsafely.
  • Drive resilience: design and execute chaos experiments and validate that failure modes are covered by tested safeguards.
  • Lead incident response: define the incident process and serve as incident commander for the highest-severity incidents, including security incidents within the coverage window.
  • Run on-call and escalation: establish and manage rotations and escalation paths that provide continuous coverage with clean follow-the-sun handoffs.
  • Disaster recovery:  own disaster-recovery testing and failover validation against defined recovery objectives and turn exercise findings into architectural and operational improvements.
  • Cloud security posture: co-own cloud security posture management, runtime vulnerability detection, and configuration-compliance monitoring with DevSecOps.
  • Data, streaming, and AI Ops: own reliability of stateful and streaming services, capacity planning and rightsizing, and autonomous agents for triage, predictive alerting, remediation, and self-healing.
  • Scale the team and broaden impact: mentor engineers at different seniority levels, set standards adopted across teams, and align Architecture, DevSecOps, Cloud Operations, and Product Development behind a shared reliability roadmap.

Requirements:

  • 7+ years of production engineering experience with recent hands-on reliability work.
  • Hands-on, recent experience operating large-scale, fault-tolerant production systems on AWS or GCP.
  • Observability ownership: have instrumented production systems and governed service-level objectives and error budgets, not only installed dashboards.
  • Resilience practice: have designed and executed failure experiments or disaster-recovery exercises with real failover validation.
  • Incident command: have personally commanded serious SEV1/SEV2 incidents and driven root cause through to a systemic fix.
  • Demonstrated ownership of reliability outcomes with measurable results, such as availability, mean time to recovery, and error-budget adherence.
  • Evidence of multi-team technical leadership through standards, review, coaching, and mechanisms adopted beyond one service or team.

Preferred Qualifications:

  • Proven production experience with cloud security posture management, runtime vulnerability detection, and workload protection across cloud and distributed environments.
  • Strong experience prioritizing risk using identity, workload, and exposure-path context to focus remediation on issues that materially increase attack likelihood and operational impact.
  • Experience with autonomous remediation and self-healing workflows powered by AIOps, including Amazon Bedrock Agent Core or equivalent agentic automation frameworks.
  • Hands-on experience in capacity management, resource rightsizing, efficiency engineering, and practical cost optimization based on FinOps principles.
  • Experience with load-balancing design and operations, including health-based failover, global traffic management, and performance optimization for highly available services.
  • Experience with AI traffic management via an LLM gateway, including request routing, policy enforcement, rate limiting, model fallback, latency optimization, cost controls, and observability for multi-model or multi-provider environments.
  • Ability to connect networking, security, and reliability considerations into cohesive platform design decisions that improve resilience, performance, and operability.

The approximate base salary range for this position is $163,430 - $213,972. The total compensation package includes base, bonus, equity, and a range of benefit options found on our career site.

Compensation will vary based on individual factors such as education, qualifications, and experience of the final candidate(s), specific office location, and calibration against relevant market data and internal team equity. Posted base salary figures are subject to change as new market data becomes available. Our benefits include comprehensive medical, dental, and vision plans, matching 401(k), unlimited PTO and paid holidays, parental/adoption leave, legal insurance, and a home technology stipend. Details of participation in these benefit plans will be provided when a candidate receives an offer of employment. 

At IonQ, we believe in fair treatment, access, opportunity, and advancement for all while striving to identify and eliminate barriers. We empower employees to thrive by fostering a culture of autonomy, productivity, and respect. We are dedicated to creating an environment where individuals can feel welcomed, respected, supported, and valued.
 
We are committed to equity and justice. We welcome different voices and viewpoints and do not discriminate on the basis of race, religion, ancestry, physical and/or mental disability, medical condition, genetic information, marital status, sex, gender, gender identity, gender expression, transgender status, age, sexual orientation, military or veteran status, or any other basis protected by law. We are proud to be an Equal Employment Opportunity employer.


US Technical Jobs. The position you are applying for will require access to technology that is subject to U.S. export control and government contract restrictions.  Employment with IonQ is contingent on either verifying “U.S. Person” (e.g., U.S. citizen, U.S. national, U.S. permanent resident, or lawfully admitted into the U.S. as a refugee or granted asylum) status for export controls and government contracts work, obtaining any necessary license, and/or confirming the availability of a license exception under U.S. export controls.  Please note that in the absence of confirming you are a U.S. Person for export control and government contracts work purposes, IonQ may choose not to apply for a license or decline to use a license exception (if available) for you to access export-controlled technology that may require authorization, and similarly, you may not qualify for government contracts work that requires U.S. Persons, and IonQ may decline to proceed with your application on those bases alone.  Accordingly, we will have some additional questions regarding your immigration status that will be used for export control and compliance purposes, and the answers will be reviewed by compliance personnel to ensure compliance with federal law.  

US Non-Technical Jobs. Due to applicable export control laws and regulations, candidates must be a U.S. citizen or national, U.S. permanent resident (i.e., current Green Card holder), or lawfully admitted into the U.S. as a refugee or granted asylum. Accordingly, we will have some additional questions regarding your immigration status that will be used for export control and compliance purposes, and the answers will be reviewed by compliance personnel to ensure compliance with federal law.


If you are interested in being a part of our team and mission, we encourage you to apply! 

Skills Required

  • 7+ years of production engineering experience with recent hands-on reliability work
  • Hands-on experience operating large-scale, fault-tolerant production systems on AWS or GCP
  • Instrumented production systems and governed service-level objectives and error budgets
  • Designed and executed failure experiments or disaster-recovery exercises with real failover validation
  • Personally commanded serious SEV1/SEV2 incidents and driven root cause to systemic fix
  • Demonstrated ownership of reliability outcomes with measurable results (availability, MTTR, error-budget adherence)
  • Evidence of multi-team technical leadership through standards, reviews, coaching, and cross-team adoption
  • U.S. Person status or ability to meet export-control and government-contract compliance requirements

IonQ Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about IonQ and has not been reviewed or approved by IonQ.

  • Healthcare Strength Comprehensive medical, dental, and vision coverage is described alongside HSA/FSA options, disability and life insurance, and mental health support. Inclusive elements such as transgender health benefits are also part of the package.
  • Parental & Family Support Paid maternity, paternity, and bonding leave is described as fully paid for eligible employees. Additional leave types such as bereavement leave are also included in time-off provisions.
  • Retirement Support A 401(k) plan with company matching up to 5% is included as part of the core package. Vesting is noted as applying over time, indicating the match is structured as a longer-term retention benefit.

IonQ Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: College Park, MD
415 Employees
Year Founded: 2015

What We Do

Quantum computers are a revolutionizing technology — they have the potential to transform business, society, and the planet for the better, and IonQ is at the forefront of this revolution. After over 25 years of academic research, IonQ was founded in 2015 by Chris Monroe and Jungsang Kim with $2 million in seed funding from New Enterprise Associates, a license to core technology from the University of Maryland and Duke University, and the goal of taking trapped ion quantum computing out of the lab and into the market. In the following three years, we raised an additional $20 million from GV, Amazon Web Services, and NEA, and built two of the world’s most accurate quantum computers. In 2019, we raised another $55 million in a round led by Samsung and Mubadala, and announced partnerships with Microsoft and Amazon Web Services to make our quantum computers available via the cloud. In 2020 and 2021, we built additional generations of high performance quantum hardware, added Google Cloud Marketplace to our cloud partner roster and announced a series of collaborations and business partnerships with leading academic and commercial institutions. On October 1st, 2021, IonQ began trading as IONQ on the New York Stock Exchange, making it the world's first public pure-play quantum computing company. We remain hard at work realizing the world-changing potential of quantum computing.

Why Work With Us

We’re growing a passionate, diverse team of collaborative, creative people. We believe in pursuing innovative, challenging work with integrity, alongside team members we can learn from and grow with.

Gallery

Gallery

Similar Jobs

ServiceNow Logo ServiceNow

Site Reliability Engineer

Artificial Intelligence • Cloud • HR Tech • Information Technology • Productivity • Software • Automation
Hybrid
Santa Clara, CA, USA
29000 Employees
166K-290K Annually
In-Office
5 Locations
6000 Employees
194K-267K Annually
In-Office
San Francisco, CA, USA
6000 Employees
174K-239K Annually

Domino Data Lab Logo Domino Data Lab

Site Reliability Engineer

Artificial Intelligence • Machine Learning
Easy Apply
Remote or Hybrid
US
200 Employees
200K-230K Annually

Similar Companies Hiring

Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
Kepler  Thumbnail
Fintech • Software
New York, New York
6 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account