Cloud Platforms Engineer

Posted Yesterday
Be an Early Applicant
Fairfax, VA, USA
In-Office
180K-210K Annually
Senior level
Artificial Intelligence • Cloud • Information Technology • Security • Software
The Role
Design, deploy, and maintain highly available, fault-tolerant Azure Government cloud infrastructure using Terraform and IaC. Implement HA/DR, automation, monitoring, incident response, and reliability standards (SLIs/SLOs/RTOs/RPOs). Support lifecycle management, runbooks, disaster recovery testing, and cross-team collaboration to ensure secure, scalable, recoverable mission-critical environments.
Summary Generated by Built In
Job Summary & Responsibilities

Everforth ECS is seeking an experienced Cloud Platforms Engineer specializing in reliability and resiliency to work in our Fairfax, VA office in a hybrid capacity.


Everforth ECS is seeking an experienced Cloud Platforms Engineer specializing in reliability and resiliency to design, operate, and continuously improve an Azure Government cloud infrastructure supporting mission-critical workloads for multiple coalition Mission Partner Network enclaves in support of the DoW community. This position will have a strong focus on infrastructure reliability, automation, High Availability (HA), Disaster Recovery (DR), and Infrastructure as Code (IaC), with Terraform serving as a primary platform for provisioning and managing cloud resources.


The Cloud Platforms Engineer will work across cloud infrastructure, platform engineering, security, and application teams to ensure production environments remain reliable, scalable, recoverable, secure, and operationally sustainable. The ideal candidate combines a strong cloud architecture skill set with hands-on operational experience and an automation-first approach to infrastructure management.


Key Responsibilities:

  • Design, deploy, and maintain highly available, fault-tolerant cloud infrastructure using Terraform and Infrastructure as Code principles.
  • Architect and implement High Availability (HA) and Disaster Recovery (DR) solutions, including infrastructure redundancy, automated failover, backup and restoration, geographic resiliency, and recovery procedures.
  • Manage the complete lifecycle of cloud infrastructure, including provisioning, configuration, operating system and platform maintenance, patching, upgrades, vulnerability remediation, and decommissioning.
  • Develop automation for infrastructure deployment and routine operational activities to reduce manual administration, configuration drift, and the potential for human error.
  • Implement and maintain monitoring, logging, alerting, and observability capabilities to identify infrastructure degradation, capacity constraints, performance issues, and potential service disruptions before they impact users.
  • Participate in incident response, troubleshooting, and root-cause analysis for production infrastructure events and develop corrective actions to prevent recurrence.
  • Define and maintain infrastructure reliability standards, including Service Level Indicators (SLIs), Service Level Objectives (SLOs), availability targets, recovery time objectives (RTOs), and recovery point objectives (RPOs).
  • Develop, maintain, and regularly validate disaster recovery procedures through recovery exercises, failover testing, and infrastructure restoration testing.
  • Evaluate cloud infrastructure capacity, performance, availability, and scalability and recommend architectural or operational improvements.
  • Partner with application development, cybersecurity, DevOps, and platform engineering teams to establish standardized deployment patterns and resilient cloud architectures.
  • Maintain infrastructure documentation, operational procedures, architecture diagrams, runbooks, and recovery procedures required to support production environments.
  • Provide technical leadership and guidance regarding cloud infrastructure reliability, resiliency, automation, and operational best practices.
  • Other duties, as assigned.

Salary Range: $180,000 - $210,000

General Description of Benefits 


Preferred Qualifications
  • U.S. Citizen.
  • Active DoD Secret security clearance.
  • Bachelor’s degree with 5+ years of related work experience.
  • Ability to obtain a DoD 8140 IAT Level II Security+ (or higher) within 60 days of hire.
  • Ability to work in a hybrid capacity in Fairfax, VA (up to 3 days in office).
  • Ability to travel <20% throughout the lifespan of the Program to CONUS / OCONUS customer sites and government installations.
  • Strong experience designing, deploying, and supporting highly available, fault-tolerant cloud infrastructure.
  • Hands-on experience with Terraform and Infrastructure as Code (IaC) principles for provisioning and managing cloud resources.
  • Knowledge of High Availability (HA) and Disaster Recovery (DR) architecture, including redundancy, automated failover, backup and restoration, geographic resiliency, and recovery planning.
  • Experience managing the full cloud infrastructure lifecycle, including provisioning, configuration, patching, upgrades, vulnerability remediation, maintenance, and decommissioning.
  • Strong infrastructure automation skills with an emphasis on reducing manual administration, configuration drift, and operational error.
  • Experience implementing and operating monitoring, logging, alerting, and observability solutions for production infrastructure.
  • Strong troubleshooting and diagnostic skills, including incident response, root-cause analysis, and corrective action development.
  • Understanding of Site Reliability Engineering concepts, including SLIs, SLOs, availability targets, RTOs, and RPOs.
  • Experience developing and validating disaster recovery procedures, including failover exercises, recovery testing, and infrastructure restoration.
  • Ability to evaluate infrastructure capacity, performance, scalability, availability, and resiliency and recommend architectural or operational improvements.
  • Experience working collaboratively with application development, cybersecurity, DevOps, and platform engineering teams.
  • Ability to develop and maintain technical documentation, including architecture diagrams, operational procedures, runbooks, and recovery documentation.
  • Strong understanding of cloud infrastructure security, vulnerability management, and operational best practices.
  • Demonstrated ability to provide technical leadership and guidance in infrastructure reliability, resiliency, automation, and cloud operations.
  • Experience supporting production, enterprise, regulated, or mission-critical environments is highly desirable.
  • Strong problem-solving and decision-making capabilities, with a proven ability to weigh the relative costs and benefits of potential actions and identify the most appropriate solution.
  • Highly developed interpersonal and oral/written communication skills, with the ability to effectively and professionally interact with a diverse set of stakeholders (from peers to end-users to executive management).

Skills Required

  • U.S. Citizen
  • Active DoD Secret security clearance
  • Bachelor's degree with 5+ years of related work experience
  • Ability to obtain DoD 8140 IAT Level II Security+ (or higher) within 60 days of hire
  • Ability to work hybrid in Fairfax, VA (up to 3 days in office)
  • Ability to travel <20% to CONUS/OCONUS customer sites and government installations
  • Strong experience designing, deploying, and supporting highly available, fault-tolerant cloud infrastructure
  • Hands-on experience with Terraform and Infrastructure as Code (IaC)
  • Knowledge of High Availability (HA) and Disaster Recovery (DR) architecture and planning
  • Experience managing full cloud infrastructure lifecycle (provisioning, patching, upgrades, vulnerability remediation, decommissioning)
  • Strong infrastructure automation skills to reduce manual administration and configuration drift
  • Experience implementing and operating monitoring, logging, alerting, and observability solutions
  • Strong troubleshooting, incident response, root-cause analysis, and corrective action development skills
  • Understanding of Site Reliability Engineering concepts (SLIs, SLOs, RTOs, RPOs)
  • Experience developing and validating disaster recovery procedures, failover exercises, and recovery testing
  • Ability to evaluate infrastructure capacity, performance, scalability, availability, and resiliency and recommend improvements
  • Experience collaborating with application development, cybersecurity, DevOps, and platform engineering teams
  • Ability to develop and maintain technical documentation, architecture diagrams, runbooks, and recovery documentation
  • Strong understanding of cloud infrastructure security and vulnerability management
  • Demonstrated ability to provide technical leadership and guidance in infrastructure reliability and operations
  • Experience supporting production, enterprise, regulated, or mission-critical environments
  • Strong problem-solving and communication skills

ECS Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about ECS and has not been reviewed or approved by ECS.

  • Healthcare Strength ECS advertises multiple national-network medical plan options with HSA eligibility alongside dental and vision coverage. Coverage generally begins quickly and is paired with company-paid short- and long-term disability, adding stability to the health package.
  • Retirement Support A 401(k) with Safe Harbor and immediate vesting on employer contributions is emphasized, with an employer match available. Access to an employee stock purchase plan via the parent company provides an additional savings avenue.
  • Parental & Family Support Paid parental leave up to 30 days, adoption assistance, and other family-oriented leaves are highlighted. Feedback suggests these offerings add meaningful value beyond base pay for many roles.

ECS Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Elkhorn, NE
2,129 Employees
Year Founded: 1993

What We Do

ECS, a segment of ASGN (NYSE: ASGN), delivers advanced solutions and services in cloud, cybersecurity, artificial intelligence (AI), machine learning (ML), application and IT modernization, and science and engineering. The company solves critical, complex challenges for customers across the U.S. public sector, defense, intelligence and commercial industries. ECS maintains partnerships with leading cloud, cybersecurity, and AI/ML providers and holds specialized certifications in their technologies. Headquartered in Fairfax, Virginia, ECS has more than 3,400 employees throughout the U.S. and has been recognized as a Top Workplace by The Washington Post for the last five years.

Similar Jobs

Capital One Logo Capital One

Lead Software Engineer

Fintech • Machine Learning • Payments • Software • Financial Services
Hybrid
4 Locations
55000 Employees
179K-246K Annually

Capital One Logo Capital One

Lead Software Engineer

Fintech • Machine Learning • Payments • Software • Financial Services
Hybrid
4 Locations
55000 Employees
179K-246K Annually

ECS Logo ECS

Cloud Platforms Engineer

Artificial Intelligence • Cloud • Information Technology • Security • Software
In-Office
Fairfax, VA, USA
2129 Employees
180K-210K Annually

Synergy ECP Logo Synergy ECP

Senior Systems Engineer

Information Technology • Software • Consulting • Cybersecurity
In-Office
Sterling, VA, USA
72 Employees
170K-210K Annually

Similar Companies Hiring

Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees
Kepler  Thumbnail
Fintech • Software
New York, New York
6 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account