Senior Site Reliability Engineer - FedRAMP

Posted Yesterday
Hiring Remotely in U.S.
Remote
130K-160K Annually
Senior level
Cloud • Security
The Role
Own reliability, availability, performance, and capacity for production SaaS services across Azure, AWS, and a FedRAMP High environment. Build observability, SLOs, monitoring, infrastructure automation, and remediation workflows; lead incident response, postmortems, support escalations, and disaster recovery efforts. Manage Terraform, CI/CD, Kubernetes, WAF, networking, and observability costs while improving on-call operations and collaborating with Security, Product, Support, and Development.
Summary Generated by Built In

About Delinea:
Delinea is a pioneer in securing human and machine identities through intelligent, centralized authorization, empowering organizations to seamlessly govern their interactions across the modern enterprise. Leveraging AI-powered intelligence, Delinea’s leading cloud-native Identity Security Platform applies context throughout the entire identity lifecycle – across cloud and traditional infrastructure, data, SaaS applications, and AI. It is the only platform that enables you to discover all identities – including workforce, IT administrator, developers, and machines – assign appropriate access levels, detect irregularities, and respond to threats in real-time. With deployment in weeks, not months, 90% fewer resources to manage than the nearest competitor, and a 99.995% uptime, Delinea delivers robust security and operational efficiency without compromise. Learn more about Delinea on Delinea.com, LinkedIn, X, and YouTube.

Join our passionate, global team at Delinea and help us make the world a safer and more secure place. Our success is driven by world-class product leadership, outstanding engineers, and strategic investment from TPG. We value diversity, innovation, and a culture of respect and fairness. If you're ready to push boundaries and challenge the status quo in security, we want to hear from you.
 

Apply today to help us achieve our mission.

Summary:

Delinea is looking for a Senior Site Reliability Engineer to join our Cloud Engineering team. You will own the availability and performance of several production SaaS services running in Azure and AWS, including our FedRAMP High environment. This is a hands-on role focused on measuring reliability, tuning detection, responding to incidents, and removing the manual work that keeps engineers awake at night. Prior experience in a government cloud environment is welcome but not required; we will teach you the compliance side.

What You'll Do:

  • Own reliability for a set of production SaaS services end to end: availability, performance, and capacity.

  • Define service-level indicators, set service-level objectives and error budgets, and use them to prioritize reliability work with engineering and product teams.

  • Build and tune monitoring in Datadog and Azure Monitor, including threshold, composite, and anomaly detection monitors, synthetic checks, dashboards, and alert routing that shorten time to detect and cut alert noise.

  • Automate response. Wire monitors into remediation workflows for known failure conditions and replace manual runbook steps with code.

  • Participate in an on-call rotation, lead incident response for high-severity events, and coordinate resolution across support, engineering, and product.

  • Write post-incident reviews and customer-facing root cause analyses, then drive the preventive actions to completion rather than filing and forgetting them.

  • Work support escalations: reproduce the issue, diagnose it from logs, traces, and network captures, resolve it or route it with evidence, and feed recurring patterns back into the product backlog.

  • Build and maintain infrastructure as code with Terraform and Azure DevOps pipelines.

  • Administer the web application firewall: rule tuning, rate limiting, false positive triage, and coordination with the security team.

  • Own the cost of our observability platform: Datadog index and retention spend, custom metric and APM volume, log ingestion rates, and S3 and blob storage for archived logs. Decide what gets indexed, sampled, or archived without losing the signal needed to troubleshoot.

  • Operate within our FedRAMP High environment following established change control and continuous monitoring processes, and contribute to the runbooks and standard operating procedures that support it.

  • Improve how the team runs on-call: rotation design, escalation paths, alert quality, runbook coverage, and handoff between regions.

  • Partner across Support, Security, Product, and Development to make sure new services ship with monitoring, runbooks, and SLOs in place.

What You'll Bring:

  • 8+ years in Site Reliability Engineering, DevOps, cloud operations, or production engineering for a SaaS product.

  • Hands-on Azure experience across AKS, App Service, Azure SQL, Redis, Service Bus, Front Door, and Storage, including cloud networking and cloud security fundamentals.

  • Production experience with an observability platform such as Datadog: metrics, logs, APM, dashboards, and monitor design. You should be comfortable writing queries, not just reading dashboards.

  • Demonstrated ownership of SLIs, SLOs, and error budgets for services you supported.

  • Incident response experience: you have run a bridge, made calls under pressure, and written the postmortem afterward.

  • Kubernetes in production, including ingress, deployments, resource limits, and troubleshooting failing workloads.

  • Infrastructure as code with Terraform, plus CI/CD pipeline creation and troubleshooting (Azure DevOps preferred).

  • Scripting in PowerShell and Python, and fluency with YAML and JSON.

  • Strong networking and web fundamentals: DNS, TLS and certificate chains, load balancing, reverse proxies, firewalls, and packet-level troubleshooting.

  • Knowledge of redundancy, backup, and disaster recovery strategies in cloud environments.

  • Clear written communication. Your RCAs will be read by customers and executives.

  • Willingness to participate in an on-call rotation covering weekends and emergencies.

  • Up to 10% travel.

 

We'd Love to See:

  • Exposure to a regulated or audited environment: FedRAMP, NIST 800-53, SOC 2, ISO 27001, PCI, or similar. Direct FedRAMP experience is a plus, not a requirement.

  • AWS experience alongside Azure, including CloudFormation and SES.

  • Hands-on experience with Jenkins pipeline creation and maintenance, SaltStack configuration management, and Consul for service discovery and distributed system coordination.

  • Advanced log analysis in the ELK stack, and/or CloudWatch Logs Insights QL.

  • Web application firewall administration (Imperva, Azure WAF, Cloudflare) and bot or rate-limiting rule tuning.

  • A track record of controlling observability tooling spend at scale: index tuning, sampling strategy, retention tiering, and archive and rehydration workflows.

  • Microsoft Entra ID and troubleshooting SAML and OIDC authentication flows.

  • Large-scale, multi-region, geo-redundant architectures and hands-on disaster recovery testing.

  • Experience with Jira Service Management, PagerDuty, or similar on-call and alerting tooling, and with public status page communication.

  • Running game days or chaos exercises, and mentoring engineers newer to SRE practice.

For this Job, Delinea is not considering candidates that need any type of US work authorization now or in the future. This includes, but is not limited to: F1-OPT, F1-CPT, H-1B, TN, L-1, J1, etc.

Why work at Delinea?

  • We're passionate problem-solvers helping the world's largest organizations protect what matters most: their human and machine identities.

  • We invest in people who are smart, self-motivated, and collaborative.

  • What we offer in return is meaningful work, a culture of innovation and great career progression.

At Delinea, our core values are STRONG and guide our behaviors and success:

  • Spirited - We bring energy and passion to everything we do

  • Trust - We act with integrity and deliver on our commitments

  • Respect - We listen, value different perspectives, and work as one team

  • Ownership - We take initiative and follow through

  • Nimble - We adapt quickly in a fast-changing environment

  • Global - We embrace diverse people and ideas to drive better outcomes

We believe weaving these core values into our day-to-day actions, and our process for hiring, evaluating, and promoting employees, helps us cultivate a work environment that embraces collaboration and camaraderie.

We take care of our employees. We offer competitive salaries, a meaningful bonus program, and excellent benefits, including healthcare insurance, as well as pension/retirement matching, comprehensive life insurance, an employee assistance program, time off plans, and paid company holidays.

Delinea is an Equal Opportunity and Affirmative Action employer and prohibits discrimination and harassment of any type with regard to race, color, religion, age, sex, national origin, disability status, genetics, protected veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by federal, state or local laws.

Upon conditional offer of employment, candidates are required to complete comprehensive criminal background check, verification of education, and verification of employment, per employment policy. In addition, all publicly posted social media sites may be reviewed.

 

 

 

 


Skills Required

  • 8+ years of experience in Site Reliability Engineering, DevOps, cloud operations, or production engineering for a SaaS product
  • Hands-on Azure experience with AKS, App Service, Azure SQL, Redis, Service Bus, Front Door, Storage, cloud networking, and cloud security
  • Production experience with an observability platform such as Datadog, including metrics, logs, APM, dashboards, monitor design, and query writing
  • Experience owning SLIs, SLOs, and error budgets
  • Incident response experience, including leading incident bridges and writing postmortems
  • Production Kubernetes experience, including ingress, deployments, resource limits, and workload troubleshooting
  • Infrastructure as code experience with Terraform
  • CI/CD pipeline creation and troubleshooting; Azure DevOps preferred
  • Scripting experience with PowerShell and Python, plus fluency with YAML and JSON
  • Strong networking and web fundamentals, including DNS, TLS, certificate chains, load balancing, reverse proxies, firewalls, and packet-level troubleshooting
  • Knowledge of cloud redundancy, backup, and disaster recovery strategies
  • Strong written communication for customer- and executive-facing root cause analyses
  • Willingness to participate in on-call rotations covering weekends and emergencies
  • Up to 10% travel
  • Experience in regulated or audited environments such as FedRAMP, NIST 800-53, SOC 2, ISO 27001, or PCI
  • AWS experience with CloudFormation and SES
  • Jenkins pipeline creation and maintenance, SaltStack, and Consul experience
  • Advanced ELK Stack log analysis or CloudWatch Logs Insights QL
  • Web application firewall administration and bot or rate-limiting rule tuning
  • Experience controlling observability tooling costs through indexing, sampling, retention, and archiving strategies
  • Microsoft Entra ID and SAML/OIDC authentication troubleshooting
  • Large-scale, multi-region, geo-redundant architecture and disaster recovery testing experience
  • Experience with Jira Service Management, PagerDuty, or similar on-call tooling and public status page communication
  • Experience running game days or chaos exercises and mentoring engineers

Delinea Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Delinea and has not been reviewed or approved by Delinea.

  • Healthcare Strength Company materials highlight comprehensive medical, dental, vision, disability, HSA/FSA options, and an EAP as part of the core package. Feedback suggests health and protection coverage is a well-rounded pillar of total rewards.
  • Leave & Time Off Breadth The package includes discretionary/unlimited PTO alongside paid company holidays. Feedback suggests the breadth of time off is attractive, though real usage can depend on team norms.
  • Parental & Family Support Paid parental leave is defined for both primary and secondary caregivers. Feedback suggests these provisions are meaningful support for caregivers.

Delinea Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: San Francisco, CA
794 Employees
Year Founded: 2022

What We Do

Delinea is a leading provider of privileged access management (PAM) solutions that make security seamless for the modern, hybrid enterprise. Our solutions empower organizations to secure critical data, devices, code, and cloud infrastructure to help reduce risk, ensure compliance, and simplify security. Delinea removes complexity and defines the boundaries of access for thousands of customers worldwide, including over half of the Fortune 100. Our customers range from small businesses to the world's largest financial institutions, intelligence agencies, and critical infrastructure companies. As organizations continue their digital transformations and move to the cloud, they are faced with increasingly complex privileged access requirements for the expanded threatscape. But the opposite of complex isn’t simple – it’s seamless. At Delinea, we believe every user should be treated like a privileged user and wants seamless, secure access, even as administrators want privileged access controls without excess complexity. Our solutions put privileged access at the center of cybersecurity by defining the boundaries of access. With Delinea, privileged access is more accessible. Get to know our industry-leading privileged access management solutions: - Delinea Secret Server: Secure privileges for service, application, root, and administrator accounts across your enterprise with our enterprise-grade PAM solution. Available both on-premise or in the cloud. https://delinea.com/products/secret-server/ - Delinea Cloud Suite: A unified PAM platform for managing privileged access in multi-cloud infrastructure to seamlessly secure access and protect against identity-based cyberattacks. https://delinea.com/products/cloud-suite/ - Delinea Server Suite: Secure and comprehensive access control to on-premises infrastructure, centrally managed from Active Directory, minimizing risk across all Linux, UNIX, and Windows systems. https://delinea.com/products/server-suite

Similar Jobs

AuthZed Logo AuthZed

Senior Site Reliability Engineer

Artificial Intelligence • Information Technology • Software • Database
Remote
2 Locations
30 Employees
150K-195K Annually

MongoDB Logo MongoDB

Site Reliability Engineer

Big Data • Cloud • Software • Database
Easy Apply
Remote or Hybrid
10 Locations
5550 Employees
127K-249K Annually

Snapsheet (snapsheetapp) Logo Snapsheet (snapsheetapp)

Senior Software Engineer

Information Technology • Insurance • Professional Services • Software
In-Office or Remote
Chicago, IL, USA
593 Employees
130K-165K Annually

Similar Companies Hiring

Toro TMS Thumbnail
Cloud • Enterprise Web • Sales • Software • Transportation
Chicago, IL
80 Employees
Credal.ai Thumbnail
Software • Security • Productivity • Machine Learning • Artificial Intelligence
Brooklyn, NY
Milestone Systems Thumbnail
Artificial Intelligence • Security • Software • Analytics • Big Data Analytics
Lake Oswego, OR
1500 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account