SRE Platform Engineer

Posted 2 Days Ago
Hiring Remotely in Washington, DC, USA
In-Office or Remote
115K-252K Annually
Expert/Leader
Information Technology • Consulting • Defense
The Role
Operate and improve reliable Azure Government platforms supporting AI applications, OIGChat, and enterprise data systems. Responsibilities include monitoring, observability, incident response, root-cause analysis, performance and cost optimization, capacity planning, disaster recovery, high availability, deployment readiness, and operational documentation. The role also supports Azure Databricks, data pipelines, AI model endpoints, secure government cloud environments, and mission-critical federal operations.
Summary Generated by Built In
Job Title: SRE Platform Engineer

Job Category: Information Technology

Time Type: Full time

Minimum Clearance Required to Start: None

Employee Type: Regular

Percentage of Travel Required: None

Type of Travel: None

* * *

The Opportunity:
CACI is seeking a seasoned Site Reliability (SRE) Platform Engineer to support the Department of Homeland Security (DHS) Office of the Inspector General (OIG). This role offers a unique opportunity to ensure the reliability, performance, and availability of cutting-edge AI and data analytics platforms that strengthen national security oversight through investigative, audit, and inspection operations.
As an SRE Platform Engineer, you will be the operational guardian of mission-critical systems including OIG Chat—a large language model assistant—enterprise data platforms, and AI-powered applications that federal oversight professionals depend on daily. You will monitor production environments, implement comprehensive alerting and observability, troubleshoot complex incidents, optimize performance and costs, and ensure high availability through proactive capacity planning and reliability engineering. From analyzing performance metrics to identify bottlenecks to supporting disaster recovery and operational readiness, you will apply site reliability engineering principles to maintain service excellence. Working within secure Azure Government environments, you will build operational resilience for a transformational program of national importance. Join us to make a meaningful impact by ensuring mission-critical AI and data capabilities are always available, performant, and reliable.

Responsibilities:

Monitor, maintain, and support production and non-production environments for OIGChat, AI applications, and enterprise data platforms to ensure availability, performance, service health, and adherence to service level objectives (SLOs)
Implement comprehensive observability including alerting, dashboards, health checks, synthetic monitoring, and log analysis using Azure Monitor, Application Insights, Log Analytics, or equivalent tools to enable proactive incident detection
Lead incident response and troubleshooting efforts including root cause analysis, defect resolution, dependency updates, integration validation, and coordination of emergency changes to restore service rapidly
Analyze performance and usage metrics across applications, APIs, AI model endpoints, data pipelines, and infrastructure to identify and remediate bottlenecks, latency issues, resource constraints, and efficiency opportunities
Support capacity planning, resource sizing, autoscaling configuration, and cost optimization for compute, storage, and AI model consumption to balance performance requirements with fiscal responsibility
Implement and maintain backup/restore processes, disaster recovery procedures, high availability architectures, and business continuity capabilities to ensure data protection and operational resilience
Monitor data platform availability including Azure Databricks clusters, data pipelines, storage services, and analytical workloads with alerting for pipeline failures, job errors, and performance degradation
Support deployment and operational readiness for pilot applications and new capabilities including pre-production validation, performance testing, runbook development, and go-live coordination
Provide surge support for complex technical issues, large-scale data collection analysis, analytical environment optimization, and specialized troubleshooting requiring deep platform knowledge
Develop and maintain operational documentation including runbooks, troubleshooting guides, architecture diagrams, incident post-mortems, and knowledge transfer materials to support sustainable operations
Qualifications:
Required – 
Bachelor's degree + 15 years of experience in site reliability engineering, DevOps, platform engineering, systems administration, or related field; equivalencies considered (Master's + 12 years; 21 years with no degree; AA + 17 years)
Must be able to obtain a Active DHS/ EOD Clearance as required.
Extensive experience with Site Reliability Engineering (SRE) principles including monitoring, observability, incident response, capacity planning, performance optimization, and reliability engineering practices
Proven expertise with Azure cloud services including compute, storage, networking, monitoring, and platform-as-a-service (PaaS) offerings with deep understanding of operational best practices
Strong experience with monitoring and observability tools (Azure Monitor, Application Insights, Grafana, Prometheus, ELK stack) and implementing alerting, dashboards, and log aggregation
Demonstrated ability to troubleshoot complex technical issues across application, platform, and infrastructure layers with strong analytical and problem-solving skills
Desired –
Experience operating AI/ML platforms, large language model services (Azure OpenAI), data analytics platforms (Databricks, Synapse), or high-scale cloud applications in production environments
Hands-on experience with Azure Government or other secure government cloud environments (AWS GovCloud) with understanding of compliance monitoring, security operations, and federal operational requirements
Background in federal government, mission-critical systems, or 24/7 operational environments with experience supporting incident response, change management, and operational excellence programs


 
 

-

What You Can Expect:

 A culture of integrity.

At CACI, we place character and innovation at the center of everything we do. As a valued team member, you’ll be part of a high-performing group dedicated to our customer’s missions and driven by a higher purpose – to ensure the safety of our nation.

An environment of trust.

CACI values the unique contributions that every employee brings to our company and our customers - every day. You’ll have the autonomy to take the time you need through a unique flexible time off benefit and have access to robust learning resources to make your ambitions a reality.

A focus on continuous growth.

Together, we will advance our nation's most critical missions, build on our lengthy track record of business success, and find opportunities to break new ground — in your career and in our legacy.

Pay Range:

There are a host of factors that can influence final salary including, but not limited to, geographic location, Federal Government contract labor categories and contract wage rates, relevant prior work experience, specific skills and competencies, education, and certifications. Our employees value the flexibility at CACI that allows them to balance quality work and their personal lives. We offer competitive compensation, benefits and learning and development opportunities. Our broad and competitive mix of benefits options is designed to support and protect employees and their families. At CACI, you will receive comprehensive benefits such as; healthcare, wellness, financial, retirement, family support, continuing education, and time off benefits.

Since this position can be worked in more than one location, the range shown is the national average for the position.

The proposed salary range for this position is:

$114,600-$252,100

CACI is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, pregnancy, sexual orientation, age, national origin, disability, status as a protected veteran, or any other protected characteristic.

Skills Required

  • Bachelor's degree and 15 years of experience in site reliability engineering, DevOps, platform engineering, systems administration, or a related field; equivalent combinations include a master's degree plus 12 years, 21 years without a degree, or an associate degree plus 17 years.
  • Must be able to obtain an active DHS/EOD clearance.
  • Extensive experience applying Site Reliability Engineering principles, including monitoring, observability, incident response, capacity planning, performance optimization, and reliability engineering.
  • Proven expertise with Azure cloud services, including compute, storage, networking, monitoring, and platform-as-a-service offerings.
  • Strong experience with Azure Monitor, Application Insights, Grafana, Prometheus, ELK Stack, alerting, dashboards, and log aggregation.
  • Demonstrated ability to troubleshoot complex issues across application, platform, and infrastructure layers.
  • Experience operating AI/ML platforms, Azure OpenAI, Databricks, Synapse, or high-scale cloud applications in production.
  • Experience with Azure Government, AWS GovCloud, secure government cloud environments, compliance monitoring, security operations, and federal operational requirements.
  • Background supporting federal government, mission-critical systems, or 24/7 operational environments, including incident response and change management.

CACI International Inc Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about CACI International Inc and has not been reviewed or approved by CACI International Inc.

  • Healthcare Strength — Pay is supported by a broad set of health-plan options across multiple national carriers, with telemedicine and tax-advantaged accounts included. Dental and vision choices are also described as multi-option, which strengthens overall coverage breadth.
  • Leave & Time Off Breadth — Flexible Time Off is available for many salaried-exempt roles, while hourly roles accrue PTO, creating multiple time-off pathways by employment class. Paid disability coverage, fixed holidays, and paid leave programs (including parental leave and other leave types) further round out time-off and leave coverage.
  • Retirement Support — Retirement support includes a 401(k) match structure described as 50% up to 8% of pay (effective 4%) and an Employee Stock Purchase Plan with a discount. Tuition reimbursement and certification support also add to the overall rewards package that complements core retirement benefits.

CACI International Inc Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Reston, VA
17,673 Employees
Year Founded: 1962

What We Do

CACI’s approximately 23,000 talented employees are vigilant in providing the unique expertise and distinctive technology that address our customers’ greatest enterprise and mission challenges. Our culture of good character, innovation, and excellence drives our success and earns us recognition as a Fortune World's Most Admired Company. As a member of the Fortune 1000 Largest Companies, the Russell 1000 Index, and the S&P MidCap 400 Index, we consistently deliver strong shareholder value. Visit us at www.caci.com.

Similar Jobs

GitLab Logo GitLab

Site Reliability Engineer

Cloud • Security • Software • Cybersecurity • Automation
Easy Apply
Remote
3 Locations
2500 Employees
223K-380K Annually

DroneUp Logo DroneUp

Site Reliability Engineer

Information Technology
Remote
United States
38 Employees
125K-150K Annually

Rapid7 Logo Rapid7

Senior Director, Customer Innovation

Artificial Intelligence • Cloud • Information Technology • Sales • Security • Software • Cybersecurity
Remote or Hybrid
United States
2400 Employees
211K-285K Annually

Rapid7 Logo Rapid7

Vector Command Specialist

Artificial Intelligence • Cloud • Information Technology • Sales • Security • Software • Cybersecurity
Remote or Hybrid
United States
2400 Employees
89K-121K Annually

Similar Companies Hiring

NODA AI Thumbnail
Artificial Intelligence • Information Technology • Software • Cybersecurity
Sydney, AU
54 Employees
Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees
Outpost Space Thumbnail
Aerospace • Defense
US
24 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account