Senior Cloud Site Reliability Engineer (SRE)

Posted 2 Hours Ago
Hiring Remotely in United States
Remote
104K-166K Annually
Senior level
Aerospace • Information Technology • Security • Cybersecurity • Defense
Do the can't be done.
The Role
Designs and develops Python-based AWS cloud solutions, reliability tools, automation, observability, and Terraform infrastructure. Builds CI/CD pipelines, defines SRE standards and SLOs, supports incident management, and improves platform reliability through monitoring, postmortems, testing, and proactive toil reduction. Collaborates with cross-functional Agile teams to deliver scalable cloud automation across the Federal Reserve System.
Summary Generated by Built In
Responsibilities

Peraton is looking for a Senior Cloud Site Reliability Engineer (SRE) who will be responsible for designing and developing advanced Python-based AWS cloud solutions and engineering reliability tools for the Cloud Foundation Services (CFS) platform in the Infrastructure, Platforms & Operations organization. This person will apply software engineering practices including Infrastructure-as-Code (IaC) with Terraform to build scalable, reusable solutions and utilities that enhance platform reliability across the Federal Reserve System.


Work Location: This is a remote position


What You Will Do

  • Design, develop, and maintain reliability solutions and SRE utilities using Python in AWS environments to reduce toil, improve cloud platform reliability, and industrialize SRE practices across the system
    • Build automation scripts, APIs, and utilities in Python to reduce toil and improve platform reliability.
    • Implement observability and monitoring solutions (Grafana, AWS CloudWatch) leveraging Python for custom metrics and dashboards
  • Build and optimize Infrastructure as Code (IaC) using Terraform to manage AWS resources related to SRE solutions, incorporating cost-efficient design principles
    • Optimize Infrastructure as Code (IaC) with Terraform for AWS resources, integrating Python-based workflows.
  • Develop CI/CD pipelines and automated testing to ensure code quality, reliability, and rapid delivery of the solutions
  • Define SRE standards, best practices, and guidelines for adoption across teams; establish SRE metrics like SLI, SLOs, etc. 
  • Apply software engineering best practices including version control, code reviews, test-driven development, and documentation to all development
  • Participate in incident management and on-call rotation, providing technical support for SRE tools, troubleshooting production issues, and collaborating with teams to reduce incident recurrence through proactive detection and pattern analysis
  • Stay current with emerging AWS services, SRE methodologies, and cloud-native development technologies, and drive adoption of innovative solutions
  • Collaborate within Agile and Scaled Agile frameworks with cross-functional teams to deliver integrated cloud automation solutions
  • Produce clear, blameless postmortems with actionable items and documented failure scenarios
Qualifications

Basic Qualifications

  • Must be a U.S. Citizen with the ability to obtain and maintain the required Public Trust level Clearance
  • Bachelors Degree and 8 years of experience, or a High School diploma or equivalent and 12 years of experience
  • Must have 5+ years of advanced Python development experience, building enterprise-grade, highly available tools, APIs, and utilities for AWS
  • 7+ years of extensive experience in software development with focus on reliability and platform engineering
  • 3+ years of hands-on experience developing solutions in AWS environments with deep understanding of core services (EC2, VPC, S3, Lambda, IAM, CloudFormation, EventBridge, Step Functions etc.) and resource cost optimization
  • 3+ years of experience applying SRE principles including observability, toil automation, SLIs/SLOs and reliability engineering
  • Expert-level proficiency with Infrastructure as Code (IaC) using Terraform, including module development and state management
  • Strong experience with CI/CD pipelines, automated testing frameworks, and DevOps practices
  • Experience with observability tools and practices including Grafana, AWS CloudWatch, AWS Canary
  • Experience defining, implementing, and managing SLOs/SLIs and error budgets; familiarity with conducting RCAs and producing postmortem documentation
  • Working experience in Agile and Scaled Agile environments and familiarity with ITSM processes (incident, change, and problem management), resilience testing and chaos engineering practices

Preferred Qualifications

  • Experience with GoLang or additional programming languages is a plus
  • Bachelors Degree in Computer Science, Information Systems, or similar
Peraton Overview

Peraton is a next-generation national security company that drives missions of consequence spanning the globe and extending to the farthest reaches of the galaxy. As the world’s leading mission capability integrator and transformative enterprise IT provider, we deliver trusted, highly differentiated solutions and technologies to protect our nation and allies. Peraton operates at the critical nexus between traditional and nontraditional threats across all domains: land, sea, space, air, and cyberspace. The company serves as a valued partner to essential government agencies and supports every branch of the U.S. armed forces. Each day, our employees do the can’t be done by solving the most daunting challenges facing our customers. Visit peraton.com to learn how we’re keeping people around the world safe and secure.

Target Salary Range$104,000 - $166,000. This represents the typical salary range for this position. Salary is determined by various factors, including but not limited to, the scope and responsibilities of the position, the individual’s experience, education, knowledge, skills, and competencies, as well as geographic location and business and contract considerations. Depending on the position, employees may be eligible for overtime, shift differential, and a discretionary bonus in addition to base pay. EEOEEO: Equal opportunity employer, including disability and protected veterans, or other characteristics protected by law.

Skills Required

  • U.S. citizenship and ability to obtain and maintain the required Public Trust clearance
  • Bachelor’s degree and 8 years of experience, or high school diploma/equivalent and 12 years of experience
  • 5+ years of advanced Python development experience building enterprise-grade, highly available tools, APIs, and AWS utilities
  • 7+ years of software development experience focused on reliability and platform engineering
  • 3+ years of hands-on AWS development experience with core AWS services and resource cost optimization
  • 3+ years applying SRE principles, including observability, toil automation, SLIs, SLOs, and reliability engineering
  • Expert-level Terraform Infrastructure as Code experience, including module development and state management
  • Strong experience with CI/CD pipelines, automated testing frameworks, and DevOps practices
  • Experience with Grafana, AWS CloudWatch, AWS Canary, and observability practices
  • Experience defining, implementing, and managing SLOs, SLIs, and error budgets
  • Familiarity with root-cause analysis, postmortem documentation, ITSM, resilience testing, and chaos engineering
  • Experience working in Agile and Scaled Agile environments
  • Experience with GoLang or additional programming languages
  • Bachelor’s degree in Computer Science, Information Systems, or a similar field

Peraton Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Peraton and has not been reviewed or approved by Peraton.

  • Healthcare Strength — Healthcare offerings are described as broad, including medical, dental, vision, mental health support, and specialty programs like virtual physical therapy and surgical/cancer care navigation. Medical coverage is also characterized as a standout part of the overall package in terms of perceived quality.
  • Retirement Support — Retirement support includes a 401(k) with employer matching and is repeatedly positioned as a strong component of total rewards. The 401(k) benefit is frequently singled out as one of the most valued parts of the package.
  • Leave & Time Off Breadth — Time-off benefits are positioned as robust, including PTO, holidays, paid sick days, and floating holidays, with some roles offering notably generous accrual. Leave breadth is reinforced by related programs such as short/long-term disability and bereavement leave.

Peraton Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Reston, VA
18,000 Employees
Year Founded: 2017

What We Do

Peraton is a next-generation national security company that drives missions of consequence spanning the globe and extending to the farthest reaches of the galaxy. As the world’s leading mission capability integrator and transformative enterprise IT provider, we deliver trusted, highly differentiated solutions and technologies to protect our nation and allies. Peraton operates at the critical nexus between traditional and nontraditional threats across all domains: land, sea, space, air, and cyberspace. The company serves as a valued partner to essential government agencies and supports every branch of the U.S. armed forces. Each day, our employees do the can’t be done by solving the most daunting challenges facing our customers. Visit peraton.com to learn how we’re keeping people around the world safe and secure.

Why Work With Us

We are fearlessly solving the world’s most complex challenges to keep people safe and secure. You can be a part of protecting and promoting that freedom around the world.

Gallery

Gallery

Similar Jobs

NVIDIA Logo NVIDIA

Senior Software Engineer

Artificial Intelligence • Computer Vision • Hardware • Robotics • Metaverse
In-Office or Remote
2 Locations
21960 Employees
152K-288K Annually

Cato Networks Logo Cato Networks

Senior Site Reliability Engineer

Information Technology • Security • Cybersecurity
Remote
United States
931 Employees

Oscar Health Logo Oscar Health

Senior Software Engineer

Healthtech • Insurance
In-Office or Remote
San Francisco, CA, USA
2200 Employees
181K-237K Annually

Oscar Health Logo Oscar Health

Senior Software Engineer

Healthtech • Insurance
In-Office or Remote
Boston, MA, USA
2200 Employees
181K-237K Annually

Similar Companies Hiring

Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees
Outpost Space Thumbnail
Aerospace • Defense
US
24 Employees
Revel Thumbnail
Aerospace • Hardware • Robotics • Software
Marina Del Rey, California
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account