Site Reliability Engineer (SRE) – II

Posted Yesterday
Be an Early Applicant
2 Locations
In-Office
125K-187K Annually
Mid level
Cloud • Information Technology • Security • Software
The Role
Provides 24/7 monitoring, incident response, troubleshooting, and maintenance across RHEL bare-metal servers, AWS/GovCloud, and Kubernetes EKS environments. Responsibilities include observability with Prometheus, Grafana, Kibana, and Elasticsearch; vulnerability remediation and FedRAMP compliance; AWS networking and load-balancing operations; GitLab and GitOps deployments using ArgoCD and Argo Workflows; infrastructure changes with Terraform; and maintaining operational documentation and automation scripts.
Summary Generated by Built In

At F5, we strive to bring a better digital world to life. Our teams empower organizations across the globe to create, secure, and run applications that enhance how we experience our evolving digital world. We are passionate about cybersecurity, from protecting consumers from fraud to enabling companies to focus on innovation. 
 

Everything we do centers around people. That means we obsess over how to make the lives of our customers, and their customers, better. And it means we prioritize a diverse F5 community where each individual can thrive.

Role Summary
We are seeking a proactive and detail-oriented Site Reliability Engineer II (SRE II) to join our 24/7 Operations team in a hybrid capacity. In this role, you will provide round-the-clock, eyes-on-glass monitoring, proactive incident response, operational maintenance, and continuous compliance support across our hybrid environment—spanning bare-metal on-premises Red Hat Enterprise Linux (RHEL) servers, AWS (Commercial and GovCloud), and Kubernetes (EKS) infrastructure operating under strict FedRAMP standards.


As an SRE II, your primary responsibility is maintaining platform availability, hardware reliability, and security posture through real-time telemetry monitoring via Prometheus and Grafana, log troubleshooting and root cause analysis in Kibana and Elasticsearch, rapid incident triage in Slack, execution of automated deployments and GitOps continuous delivery via GitLab CI/CD, ArgoCD, and Argo Workflows, and consistent enforcement of FedRAMP security controls (NIST SP 800-53). You will operate in a structured shift model covering weekdays and weekends to ensure 24/7/365 hybrid platform uptime.


24/7 Shift & On-Call Expectations

This position requires active participation in a 24/7/365 operational shift model:

• Round-the-Clock Coverage: Active "eyes-on-glass" monitoring during scheduled shifts (day, evening, night, and weekend rotations).

• Weekend & Holiday Rotation: Scheduled weekend shifts and holiday coverage to ensure continuous operational readiness across both on-premises data centers and cloud regions.

• On-Call Escalations: Primary and secondary on-call responsibilities during and outside regular shift windows to meet stringent FedRAMP Incident Response (IR) SLAs.

• Real-time Collaboration: Continuous presence in operational Slack channels and ChatOps bridge rooms for instant incident mobilization.

Key Responsibilities

1. 24/7 Eyes-on Monitoring, Log Troubleshooting & Incident Triage

• Maintain continuous monitoring of production telemetry across bare-metal RHEL servers, EKS clusters, and AWS cloud environments.

• Act as first responder to high-priority alerts generated by Prometheus and Alertmanager, performing immediate triage and root cause investigation.

• Perform deep-dive log troubleshooting using Elasticsearch and Kibana (building queries, analyzing container/application stdout logs, systemd/journald logs, and ingress traffic logs) to isolate error patterns and service failures.

• Coordinate incident resolution bridges in Slack, engaging secondary on-call engineers, security teams, or infrastructure engineers as required.

• Maintain rigorous adherence to Incident Response SLAs and document timeline logs for post-incident reviews (RCAs).

2. Bare-Metal On-Premises RHEL Server Support & Hardware Operations

• Provide operational support for bare-metal on-premises Linux servers (RHEL), including OS configuration, system maintenance, and day-to-day lifecycle administration.

• Diagnose physical network interface bonding, link aggregation, VLAN tagging, and local server connectivity issues on bare-metal systems.

• Coordinate vendor hardware dispatch requests for failing server components and oversee physical parts replacements.

3. FedRAMP Security & Vulnerability Remediation

• Execute day-to-day security operational duties aligned with FedRAMP High/Moderate (NIST SP 800-53) standards.

• Perform routine vulnerability patching (CVE remediation) across bare-metal RHEL servers, cloud AMIs, Kubernetes worker nodes, and container images within strict regulatory timelines.

• Apply operational system hardening based on DISA STIG guidelines and maintain FIPS 140 compliance configurations across both cloud and on-premise operating systems.

• Ensure strict Role-Based Access Control (RBAC), SSH key management, and security boundary enforcement across operational environments.

4. Kubernetes (EKS), AWS Infrastructure & Networking Operations

• Support day-to-day operations and node maintenance for Amazon EKS clusters across AWS Commercial and AWS GovCloud environments.

• Perform Layer 4 (L4) and Layer 7 (L7) secure load balancing troubleshooting, including AWS Application Load Balancers (ALB), Network Load Balancers (NLB), ingress controllers, SSL/TLS certificate termination, and traffic routing issues.

• Perform Linux administration tasks including kernel parameter tuning (sysctl), storage expansion, log rotation, and system troubleshooting.

• Execute infrastructure changes and updates using Infrastructure as Code (Terraform) in alignment with change control procedures.

• Monitor key AWS cloud infrastructure components including VPCs, Security Groups, EC2, IAM, S3, and KMS.

5. CI/CD & GitOps Deployment Operations (GitLab, ArgoCD, Argo Workflows)

• Execute, monitor, and troubleshoot automated deployment pipelines using GitLab CI/CD.

• Manage application state, synchronization, and rollouts across Kubernetes clusters using ArgoCD (GitOps paradigm).

• Monitor, execute, and troubleshoot operational batch processes, system maintenance tasks, and automated pipelines using Argo Workflows.

• Facilitate application releases and configuration rollouts using Helm charts, Kustomize, and GitOps workflows.

• Validate build pipeline compliance, container security scanning results, and image signature verifications prior to production deployment.

6. Documentation & Operational Runbooks

• Maintain accurate, step-by-step incident runbooks, standard operating procedures (SOPs), and triage workflows for both cloud and bare-metal environments.

• Write detailed post-incident reports and post-mortems for production impact events.

• Identify manual operational overhead (toil) and implement shell scripts (Bash/Python) to streamline routine monitoring, hardware checks, and maintenance tasks.

Technical Skills & Qualifications


Required Qualifications

• Citizenship & Regulatory Standard: US Citizenship required due to FedRAMP / AWS GovCloud security requirements.

• Experience: 3–5 years of hands-on experience in SRE, DevOps, Systems Administration, or Hybrid Infrastructure Support roles.

• 24/7 Shift Availability: Willingness and capability to work in a 24/7 rotational shift schedule (including nights, weekends, and holidays).

• Bare-Metal & On-Premises Linux Administration: Strong hands-on experience administering bare-metal Linux servers (Red Hat Enterprise Linux / RHEL), including hardware diagnostics, out-of-band management (IPMI, iDRAC, iLO), RAID configuration, LVM, and physical NIC bonding.

• Log Troubleshooting (Kibana/Elasticsearch): Direct experience querying and analyzing log streams in Kibana and Elasticsearch (KQL/Lucene queries, index management, log pattern matching for microservices and cluster components).

• L4/L7 Secure Load Balancing: Practical experience troubleshooting Layer 4 (NLB) and Layer 7 (ALB / Ingress Controllers) secure load balancing, mTLS, TLS termination, health checks, and secure traffic routing.

• AWS & GovCloud: Solid operational experience with AWS core services (EC2, VPC, IAM, S3, KMS) and familiarity with AWS GovCloud operating models.

• Kubernetes (EKS): Hands-on operational experience with Amazon EKS / Kubernetes (kubectl, Helm, pod lifecycle management, node pool maintenance).

• Observability Tools: Experience working with Prometheus, PromQL metrics queries, Alertmanager, and Grafana dashboard visualization.

• CI/CD & Continuous Delivery: Hands-on operational experience executing and troubleshooting deployments with GitLab CI/CD, managing GitOps syncing with ArgoCD, and orchestrating Kubernetes workflows using Argo Workflows.

• ChatOps & Communication: Experience utilizing Slack for operational messaging, ChatOps commands, and incident war rooms.

• FedRAMP / Security Hardening: Understanding of FedRAMP/NIST SP 800-53 controls, vulnerability patching cycles, and DISA STIG hardening on both bare-metal and cloud Linux systems.


Preferred Qualifications

• Experience with automation scripting using Python or Bash.

• Basic working knowledge of Terraform for infrastructure maintenance.

• Red Hat Certified System Administrator (RHCSA) or Red Hat Certified Engineer (RHCE).

• AWS Certified SysOps Administrator or Certified Kubernetes Administrator (CKA).

Key Performance Indicators (KPIs)

• Monitoring MTTA (Mean Time to Acknowledge): Rapid response times for eyes-on alert notifications.

• MTTR (Mean Time to Resolve): Efficient triage, log analysis in Kibana, and mitigation of production service disruptions across cloud and bare-metal environments.

• Hardware Availability & Uptime: Rapid isolation and remediation of physical server component failures and RAID disk faults.

• Deployment & GitOps Stability: High release success rates and swift remediation of failed ArgoCD syncs or Argo Workflows runs.

• Shift Readiness & Escalation Precision: Clear shift handovers and timely escalation management during 24/7 windows.

• CVE Remediation Timeliness: Strict compliance with patching SLAs for bare-metal OS, container images, and cloud AMIs.

• Runbook Currency: Continuous refinement of operational runbooks based on shift learnings.


#LI-KA1

The Job Description is intended to be a general representation of the responsibilities and requirements of the job. However, the description may not be all-inclusive, and responsibilities and requirements are subject to change.

The annual base pay for this position is: $124,800.00 - $187,200.00

F5 maintains broad salary ranges for its roles in order to account for variations in knowledge, skills, experience, geographic locations, and market conditions, as well as to reflect F5’s differing products, industries, and lines of business. The pay range referenced is as of the time of the job posting and is subject to change.

You may also be offered incentive compensation, bonus, restricted stock units, and benefits. More details about F5’s benefits can be found at the following link: https://www.f5.com/company/careers/benefits. F5 reserves the right to change or terminate any benefit plan without notice. 

Please note that F5 only contacts candidates through F5 email address (ending with @f5.com) or auto email notification from Workday (ending with f5.com or @myworkday.com).

Equal Employment Opportunity

It is the policy of F5 to provide equal employment opportunities to all employees and employment applicants without regard to unlawful considerations of race, religion, color, national origin, sex, sexual orientation, gender identity or expression, age, sensory, physical, or mental disability, marital status, veteran or military status, genetic information, or any other classification protected by applicable local, state, or federal laws. This policy applies to all aspects of employment, including, but not limited to, hiring, job assignment, compensation, promotion, benefits, training, discipline, and termination.  F5 offers a variety of reasonable accommodations for candidates. Requesting an accommodation is completely voluntary. F5 will assess the need for accommodations in the application process separately from those that may be needed to perform the job. Request by contacting [email protected].

Skills Required

  • US citizenship due to FedRAMP and AWS GovCloud security requirements
  • 3-5 years of hands-on experience in SRE, DevOps, systems administration, or hybrid infrastructure support
  • Availability for a 24/7 rotational shift schedule including nights, weekends, and holidays
  • Hands-on bare-metal Linux administration with RHEL, hardware diagnostics, IPMI/iDRAC/iLO, RAID, LVM, and physical NIC bonding
  • Experience querying and analyzing logs in Kibana and Elasticsearch using KQL/Lucene
  • Experience troubleshooting L4/L7 secure load balancing, NLB, ALB, ingress controllers, mTLS, TLS termination, health checks, and secure traffic routing
  • Operational experience with AWS core services including EC2, VPC, IAM, S3, and KMS, plus familiarity with AWS GovCloud
  • Hands-on operational experience with Amazon EKS and Kubernetes, including kubectl, Helm, pod lifecycle management, and node pool maintenance
  • Experience with Prometheus, PromQL, Alertmanager, and Grafana
  • Operational experience with GitLab CI/CD, ArgoCD, and Argo Workflows
  • Experience using Slack for operational messaging, ChatOps commands, and incident war rooms
  • Understanding of FedRAMP, NIST SP 800-53, vulnerability patching, and DISA STIG hardening
  • Automation scripting experience with Python or Bash
  • Basic working knowledge of Terraform
  • RHCSA or RHCE certification
  • AWS Certified SysOps Administrator or Certified Kubernetes Administrator certification

F5 Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about F5 and has not been reviewed or approved by F5.

  • Equity Value & Accessibility — Equity grants and an employee stock purchase plan are positioned as meaningful parts of total compensation, with RSUs and a discount ESPP commonly included. Pay packages for many technical roles are considered competitive when equity is taken into account.
  • Leave & Time Off Breadth — Paid vacation that increases with tenure, sick time, paid holidays, and paid family leave are prominently featured. Additional programs like volunteer time and periodic wellness long weekends are highlighted as part of the time-off ecosystem.
  • Inclusive Benefits Coverage — Health plans include travel support for specific care (such as reproductive and gender‑affirming services) and mental health resources, alongside comprehensive medical, dental, and vision coverage. These elements are presented as part of a broad, inclusive approach to healthcare.

F5 Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Seattle, WA
5,847 Employees

What We Do

F5 application services ensure that applications are always secure and perform the way they should—in any environment and on any device. F5 (NASDAQ: FFIV) powers applications from development through their entire life cycle, across any multi-cloud environment, so our customers – enterprise businesses, service providers, governments, and consumer brands—can deliver differentiated, high-performing, and secure digital experiences.

Similar Jobs

Microsoft Logo Microsoft

Site Reliability Engineer

Software • Quantum Computing • Metaverse • Infrastructure as a Service (IaaS)
In-Office
Reston, VA, USA
206870 Employees
102K-219K Annually

Akamai Technologies Logo Akamai Technologies

Site Reliability Engineer

Cloud • Security • Software • Cybersecurity
In-Office or Remote
2 Locations
10285 Employees
138K-171K Annually

Akamai Technologies Logo Akamai Technologies

Site Reliability Engineer

Cloud • Security • Software • Cybersecurity
In-Office or Remote
2 Locations
10285 Employees
146K-264K Annually

Akamai Technologies Logo Akamai Technologies

Site Reliability Engineer

Cloud • Security • Software • Cybersecurity
In-Office or Remote
2 Locations
10285 Employees
95K-171K Annually

Similar Companies Hiring

Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Revel Thumbnail
Aerospace • Hardware • Robotics • Software
Marina Del Rey, California
60 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software
New York, New York
30 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account