FedRAMP now defines compliance as a set of Key Security Indicators (KSIs); measurable security outcomes that automation validates continuously. Holding an estate at that standard day after day is an operations job. Coalfire organizes its delivery engineering into capability-focused teams, and the Run teams keep authorized client environments available and provably compliant long after the build team leaves. As a Senior Site Reliability Engineer you own one operational capability, such as remidation, monitoring and alerting or backup and recovery, and you design how Coalfire observes a client estate and holds it in a compliant state. You own the automation and the procedure that each managed environment inherits, you carry escalation for the hardest operational problems, and you set and continually raise the engineering bar for the engineers around you. If you are driven by a desire to innovate, excel at operational excellence, and thrive in a collaborative environment, come be part of a team committed to making the world a more secure and complaint place.
What You'll Do
- Own one operational capability for the managed estate, such as monitoring and alerting or backup and recovery, including its automation, its runbooks, and the service standard each managed environment inherits.
- Design how Coalfire observes a regulated cloud environment: telemetry and log pipelines, service-level objectives, alert quality, and the escalation paths behind them. Make alerts useful and actionable.
- Build and maintain the continuous-monitoring evidence pipeline so you can prove the authorized state each day, in machine-readable form where the framework allows it.
- Own backup and recovery engineering for your capability: recovery procedure you have tested, recovery objectives you can measure, and the automation that runs them during an outage.
- Serve as the senior escalation point in client environments. Diagnose and troubleshoot beyond the runbook, resolve the incident, then fix the runbook so the next engineer on call does not repeat the diagnosis.
- Lead incident and problem management for your domain, including incident command on major events, blameless post-incident review, and corrective action that closes the problem out.
- Automate operational toil out of the estate with infrastructure-as-code, pipelines, and scripting. Decide what Coalfire automates once for the whole estate and what stays specific to one client.
- Partner with Engagement Architects and the Build teams on transition into managed operations: operational readiness review, monitoring and runbook coverage, and the service commitments Coalfire can meet at go-live.
- Hold on-call for your capability and improve the rotation: coverage, alert actionability, and the load it places on the team.
- Represent operational posture in front of clients, and support renewals and expansions with a credible account of what Coalfire operates for them and at what cost.
- Mentor and lead Site Reliability Engineers and junior staff: review designs as well as changes, set operational standards, and grow named individuals in your domain.
- Author and peer review code, runbooks, operational design documentation, and the compliance artifacts that evidence continuous monitoring, inclusive of vendor best practices.
What You'll Bring
- Automation-first mindset with deep Infrastructure-as-Code (Terraform or equivalent), CI/CD, and scripting (Python, Go, or similar); working use of policy-as-code
- Deep operational command of at least one major cloud platform (AWS, Azure, or GCP) and its native monitoring, logging, and recovery services, with working knowledge of a second
- Observability engineering depth: metrics, logging and log pipelines, distributed tracing, SLI and SLO definition, and alert design that produces action
- Demonstrated incident response and incident command capability, and the discipline to convert a resolved incident into automation or procedure
- Backup, recovery, and resilience engineering, including tested recovery procedure against stated recovery objectives
- Working command of NIST 800-53, FedRAMP, or comparable security control frameworks, and the judgment to map continuous-monitoring obligations to real operational design
- Ability to lead technical conversations with clients about operational posture, risk, and trade-offs with both engineers and executives
- The instinct to solve a problem once and package the solution so it works across the managed estate
- Demonstrated ability to mentor engineers and improve the output of a team
- Excellent communication, organizational, and problem-solving skills
- Effective documentation skills, including technical diagrams, runbooks, and written descriptions
- Ability to work independently and as part of a team with a professional attitude and demeanor
- Critical thinking, and the ability to balance security and availability requirements against mission needs
- BS or above in a related Information Technology field or equivalent combination of education and experience
- 5+ years in site reliability engineering, cloud operations, platform engineering, or managed services
- 5+ years operating production cloud environments in AWS, Azure, or GCP, including monitoring, incident response, and automation
- Demonstrated experience owning an operational capability, monitoring platform, or reusable automation used by multiple teams or clients
- Experience as the senior operational escalation point on client-facing managed services, including incident command on major events
- Advanced experience with Infrastructure-as-Code and orchestration/automation tools such as Terraform and Ansible
- Experience transitioning environments from build into steady-state operations
- Professional- or specialty-level certification in AWS, Azure, or GCP (associate-level considered with equivalent demonstrated depth)
REQUIRED CERTIFICATIONS:
EDUCATION:
Bachelor’s degree (four-year college or university) or equivalent combination of education and work experience.
Bonus Points
- Direct familiarity with FedRAMP Rev 5 and 20x, Key Security Indicators (KSIs), and continuous monitoring obligations
- Experience with OSCAL or JSON machine-readable compliance formats
- Recognized depth in a specialty area: SIEM and log pipelines, observability platforms, vulnerability and patch management at scale, or disaster-recovery engineering
- Experience operating a monitored compliance environment against contractual service levels
- Relevant certifications such as cloud security or DevOps specialty certifications, CISSP, or GIAC
- Previous experience supporting clients from within a professional services or managed services organization
- Experience contributing to renewals, expansions, and pre-sales technical solutioning for managed services
- Familiarity with configuration baseline standards such as CIS Benchmarks and DISA STIG
- Familiarity with frameworks such as FedRAMP, FISMA, HIPAA, HITRUST, or PCI
Skills Required
- Automation-first experience with Infrastructure as Code using Terraform or equivalent, CI/CD, scripting, and policy as code
- Deep operational experience with at least one major cloud platform: AWS, Azure, or GCP; working knowledge of a second
- Experience with metrics, logging, log pipelines, distributed tracing, SLI/SLO definition, and actionable alert design
- Demonstrated incident response and incident command capability
- Backup, recovery, and resilience engineering experience with tested recovery procedures
- Working knowledge of NIST 800-53, FedRAMP, or comparable security control frameworks
- Ability to lead technical client discussions concerning operational posture, risk, and trade-offs
- Demonstrated ability to mentor engineers and improve team output
- Excellent communication, organizational, documentation, and problem-solving skills
- Ability to work independently and collaboratively while balancing security, availability, and mission requirements
- Bachelor's degree in a related information technology field or equivalent education and experience
- At least 5 years of experience in site reliability engineering, cloud operations, platform engineering, or managed services
- At least 5 years operating production cloud environments, including monitoring, incident response, and automation
- Experience owning an operational capability, monitoring platform, or reusable automation used by multiple teams or clients
- Experience serving as the senior operational escalation point for client-facing managed services and leading major incidents
- Advanced experience with Terraform, Ansible, and orchestration or automation tools
- Experience transitioning environments from build into steady-state operations
- Professional- or specialty-level certification in AWS, Azure, or GCP; associate-level certification may be considered with equivalent depth
- Familiarity with FedRAMP Rev 5 and 20x, KSIs, and continuous-monitoring obligations
- Experience with OSCAL or JSON machine-readable compliance formats
- Specialty expertise in SIEM and log pipelines, observability, vulnerability and patch management, or disaster recovery
- Experience operating monitored compliance environments against contractual service levels
- Cloud security or DevOps specialty certifications, CISSP, or GIAC
- Professional services or managed services client-support experience
- Experience contributing to renewals, expansions, or pre-sales technical solutioning
- Familiarity with CIS Benchmarks, DISA STIG, FISMA, HIPAA, HITRUST, or PCI
Coalfire Compensation & Benefits Highlights
The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Coalfire and has not been reviewed or approved by Coalfire.
-
Leave & Time Off Breadth — Flexible paid time off and paid parental leave are prominently offered, with remote/WFH support enabling time away when workload allows.
-
Healthcare Strength — Comprehensive medical, dental, vision, wellness resources, and an EAP are part of the core package. Carrier coverage and plan options are regularly highlighted across employer materials.
-
Retirement Support — A company‑matched 401(k) is included alongside other financial and development perks. This retirement benefit is consistently featured across benefits overviews.
Coalfire Insights
What We Do
Coalfire is the cybersecurity advisor that helps private and public sector organizations avert threats, close gaps, and effectively manage risk. By providing independent and tailored advice, assessments, technical testing, and cyber engineering services, we help clients develop scalable programs that improve their security posture, achieve their business objectives, and fuel their continued success. Coalfire has been a cybersecurity thought leader for more than 20 years and has offices throughout the United States and Europe.








