Lead Site Reliability Engineer

Posted Yesterday
Be an Early Applicant
Hiring Remotely in United States
Remote
112K-179K Annually
Expert/Leader
Aerospace • Information Technology • Security • Cybersecurity • Defense
Do the can't be done.
The Role
Leads site reliability engineering for mission-critical cloud platforms, focusing on disaster recovery drills, platform rebuild validation, infrastructure automation, Kubernetes operations, CI/CD, observability, incident response, and system resilience. Collaborates across engineering, security, applications, and data teams while driving remediation and modernization efforts. The role requires deep AWS, Infrastructure as Code, containerization, automation, monitoring, and failure-analysis expertise, plus eligibility for Public Trust clearance.
Summary Generated by Built In
Responsibilities

Peraton is seeking a Lead Site Reliability Engineer to join our team of qualified, diverse individuals. The ideal candidate will play a critical role in ensuring the reliability, resilience, and recoverability of mission essential platforms by leading infrastructure level disaster recovery drills, platform rebuild validation, and automated deployment processes. This engineer will partner closely with cross functional teams to maintain and enhance complex cloud-based environments, integrate modern automation solutions, and support largescale modernization and continuity efforts across high visibility programs.


Responsibilities

The Lead Site Reliability Engineer’s responsibilities shall include, but are not limited to:

  • Supporting full lifecycle platform portability and disaster recovery (DR) drill execution, including validation of platform rebuild procedures and DR playbooks.
  • Executing infrastructure level drill activities to ensure the platform can be fully rebuilt within the 48hour recovery target.
  • Verifying end to end data completeness, integrity, and accuracy during drill exercises, documenting results and remediation recommendations.
  • Identifying exit readiness gaps across infrastructure, deployment automation, monitoring, and data recovery processes, and driving corrective actions with engineering teams.
  • Designing, implementing, and supporting automated IaaC workflows utilizing Terraform, AWS CloudFormation, and standardized CI/CD pipelines.
  • Managing and optimizing Kubernetes clusters and containerized workloads (Docker), including cluster provisioning, scaling, and workload reliability improvements.
  • Building and maintaining observability solutions using CloudWatch, Datadog, and other monitoring/alerting tools to ensure service reliability and proactive incident response.
  • Developing automation, tooling, and scripts using Python or Java to reduce manual processes and enhance operational repeatability.
  • Collaborating with platform engineering, security, applications, and data teams to ensure consistent, secure, and compliant platform operations.
  • Participating in on‑call rotations, root cause analyses, and incident response activities to improve system resilience and operational excellence.
Qualifications

Required Qualifications

  • Bachelor’s degree and 8–10 years of relevant SRE, DevOps, cloud engineering, or infrastructure engineering experience; or 12 years of experience with a high school diploma. 
  • Expert level hands on knowledge of AWS services across compute, networking, storage, IAM, and serverless components.
  • Strong experience with Infrastructure as Code (Terraform, CloudFormation) and infrastructure automation principles.
  • Experience building CI/CD deployment pipelines and progressive delivery mechanisms using GitHub actions or similar tools
  •  Deep understanding of Kubernetes administration, container orchestration, and Docker based deployments.
  • Proven experience validating DR processes, performing system rebuilds, and conducting data integrity checks.
  • Experience building monitoring tools like dashboards, metrics, logs, and alerting systems using CloudWatch, Datadog, or similar observability tools.
  • Proficiency with programming/scripting languages such as Python, Java, or C# or Go.
  • Experience debugging complex failure modes, including cascading failures, network partitions, backpressure, and eventual consistency issues.
  • Strong analytical and documentation skills with the ability to clearly communicate technical findings to cross functional teams.
  • Ability to work in a fast-paced environment supporting high visibility, mission critical systems.
  • Ability to obtain a Public Trust clearance. 
  • US Citizen or Green Card Holder. 

Preferred Qualifications

  • AWS DevSecOps Engineer certification (preferred).
  • Additional AWS certifications (Solutions Architect, SysOps, Developer) and/or Kubernetes certifications (CKA, CKAD).
  • Familiarity with Zero Trust security models and cloud security best practices.
  • Experience with GitLab, Jenkins, or similar CI/CD platforms.
  • Experience with highly regulated environments (healthcare, finance, DHS, DoD, CMS, etc.).
  • Experience supporting federal, defense, or largescale enterprise programs involving legacy-to-cloud modernization.
  • Prior involvement in large‑scale DR drills, continuity of operations (COOP), or portability/executable readiness assessments.
Peraton Overview

Peraton is a next-generation national security company that drives missions of consequence spanning the globe and extending to the farthest reaches of the galaxy. As the world’s leading mission capability integrator and transformative enterprise IT provider, we deliver trusted, highly differentiated solutions and technologies to protect our nation and allies. Peraton operates at the critical nexus between traditional and nontraditional threats across all domains: land, sea, space, air, and cyberspace. The company serves as a valued partner to essential government agencies and supports every branch of the U.S. armed forces. Each day, our employees do the can’t be done by solving the most daunting challenges facing our customers. Visit peraton.com to learn how we’re keeping people around the world safe and secure.

Target Salary Range$112,000 - $179,000. This represents the typical salary range for this position. Salary is determined by various factors, including but not limited to, the scope and responsibilities of the position, the individual’s experience, education, knowledge, skills, and competencies, as well as geographic location and business and contract considerations. Depending on the position, employees may be eligible for overtime, shift differential, and a discretionary bonus in addition to base pay. EEOEEO: Equal opportunity employer, including disability and protected veterans, or other characteristics protected by law.

Skills Required

  • Bachelor’s degree and 8–10 years of relevant SRE, DevOps, cloud engineering, or infrastructure engineering experience; or 12 years of experience with a high school diploma.
  • Expert hands-on knowledge of AWS services across compute, networking, storage, IAM, and serverless components.
  • Strong experience with Infrastructure as Code using Terraform and CloudFormation.
  • Experience building CI/CD deployment pipelines and progressive delivery mechanisms using GitHub Actions or similar tools.
  • Deep understanding of Kubernetes administration, container orchestration, and Docker-based deployments.
  • Experience validating disaster recovery processes, performing system rebuilds, and conducting data integrity checks.
  • Experience building monitoring dashboards, metrics, logs, and alerting systems using CloudWatch, Datadog, or similar tools.
  • Proficiency with Python, Java, C#, or Go.
  • Experience debugging complex failure modes, including cascading failures, network partitions, backpressure, and eventual consistency issues.
  • Strong analytical and documentation skills with the ability to communicate technical findings to cross-functional teams.
  • Ability to work in a fast-paced environment supporting high-visibility, mission-critical systems.
  • Ability to obtain a Public Trust clearance.
  • US Citizen or Green Card Holder.
  • AWS DevSecOps Engineer certification.
  • Additional AWS certifications such as Solutions Architect, SysOps, or Developer, and/or Kubernetes certifications such as CKA or CKAD.
  • Familiarity with Zero Trust security models and cloud security best practices.
  • Experience with GitLab, Jenkins, or similar CI/CD platforms.
  • Experience with highly regulated environments such as healthcare, finance, DHS, DoD, or CMS.
  • Experience supporting federal, defense, or large-scale enterprise programs involving legacy-to-cloud modernization.
  • Prior involvement in large-scale disaster recovery drills, continuity of operations, or portability and executable readiness assessments.

Peraton Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Peraton and has not been reviewed or approved by Peraton.

  • Healthcare Strength — Healthcare offerings are described as broad, including medical, dental, vision, mental health support, and specialty programs like virtual physical therapy and surgical/cancer care navigation. Medical coverage is also characterized as a standout part of the overall package in terms of perceived quality.
  • Retirement Support — Retirement support includes a 401(k) with employer matching and is repeatedly positioned as a strong component of total rewards. The 401(k) benefit is frequently singled out as one of the most valued parts of the package.
  • Leave & Time Off Breadth — Time-off benefits are positioned as robust, including PTO, holidays, paid sick days, and floating holidays, with some roles offering notably generous accrual. Leave breadth is reinforced by related programs such as short/long-term disability and bereavement leave.

Peraton Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Reston, VA
18,000 Employees
Year Founded: 2017

What We Do

Peraton is a next-generation national security company that drives missions of consequence spanning the globe and extending to the farthest reaches of the galaxy. As the world’s leading mission capability integrator and transformative enterprise IT provider, we deliver trusted, highly differentiated solutions and technologies to protect our nation and allies. Peraton operates at the critical nexus between traditional and nontraditional threats across all domains: land, sea, space, air, and cyberspace. The company serves as a valued partner to essential government agencies and supports every branch of the U.S. armed forces. Each day, our employees do the can’t be done by solving the most daunting challenges facing our customers. Visit peraton.com to learn how we’re keeping people around the world safe and secure.

Why Work With Us

We are fearlessly solving the world’s most complex challenges to keep people safe and secure. You can be a part of protecting and promoting that freedom around the world.

Gallery

Gallery

Similar Jobs

MetLife Logo MetLife

Site Reliability Engineer

Fintech • Information Technology • Insurance • Financial Services • Big Data Analytics
Remote or Hybrid
United States
43000 Employees
111K-180K Annually

MetLife Logo MetLife

Site Reliability Engineer

Fintech • Information Technology • Insurance • Financial Services • Big Data Analytics
Remote or Hybrid
United States
43000 Employees
111K-180K Annually

MetLife Logo MetLife

Site Reliability Engineer

Fintech • Information Technology • Insurance • Financial Services • Big Data Analytics
Remote or Hybrid
United States
43000 Employees
111K-180K Annually
Remote or Hybrid
2 Locations
289097 Employees

Similar Companies Hiring

Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees
Outpost Space Thumbnail
Aerospace • Defense
US
24 Employees
Revel Thumbnail
Aerospace • Hardware • Robotics • Software
Marina Del Rey, California
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account