We are seeking an experienced Lead Site Reliability Engineer to join our engineering team and drive the reliability, scalability, and performance of our critical systems. This role combines deep technical expertise in software engineering, cloud infrastructure, and DevOps practices to ensure our services meet the highest standards of availability and operational excellence.
Responsibilities
System Reliability & Performance
• Design, implement, and maintain highly available, scalable, and resilient systems across cloud infrastructure
• Establish and monitor SLIs, SLOs, and SLAs to ensure optimal system performance
• Lead incident response, conduct root cause analysis, and implement preventive measures
• Develop and maintain disaster recovery and business continuity plans
Infrastructure & Automation
• Architect and manage cloud infrastructure on AWS using Infrastructure as Code (Terraform)
• Automate deployment pipelines, monitoring, and operational workflows
• Optimize cloud resource utilization and cost management
Engineering & Development
• Build and maintain internal tools and services to improve operational efficiency
• Collaborate with development teams to implement reliability best practices
• Conduct code reviews and provide technical guidance on system design
• Develop monitoring solutions, alerting systems, and observability frameworks
Security & Compliance
• Integrate security practices into CI/CD pipelines (SAST/DAST)
• Implement and maintain security controls across infrastructure and applications
• Ensure compliance with industry standards and regulatory requirements
• Conduct security assessments and vulnerability management
Leadership & Collaboration
• Mentor junior SRE team members and promote SRE culture across the organization
• Partner with software engineering teams to improve system reliability
• Drive technical initiatives and contribute to architectural decisions
• Document processes, runbooks, and technical specifications
Software Engineering:
- Strong proficiency in Java, Python, and Node.js
- Experience with microservices architecture and distributed systems
- Solid understanding of data structures, algorithms, and design patterns
- Proficiency in writing clean, maintainable, and testable code
Cloud Infrastructure (AWS):
- Extensive experience with AWS services including:
- Compute: Lambda, ECS, EC2, Fargate
- Storage: S3, EBS, EFS
- Database: RDS, DynamoDB, Aurora
- Networking: VPC, Route53, CloudFront, API Gateway
- Monitoring: CloudWatch, X-Ray
- AWS certifications (Solutions Architect, DevOps Engineer) preferred
DevOps & CI/CD:
- Expert-level knowledge of GitLab (CI/CD pipelines, runners, GitOps)
- Advanced Terraform skills for infrastructure provisioning and management
- Experience with containerization (Docker) and orchestration (Kubernetes/ECS)
- Proficiency with configuration management tools
Security:
- Hands-on experience with SAST (Static Application Security Testing) tools
- Knowledge of DAST (Dynamic Application Security Testing) methodologies
- Understanding of security best practices, OWASP Top 10, and compliance frameworks
- Experience with secrets management and identity access management (IAM)
Monitoring & Observability:
- Experience with monitoring tools (Grafana, Datadog, New Relic, or similar)
- Log aggregation and analysis (CloudWatch Logs, Splunk)
- Distributed tracing with aws X-Ray
Qualifications
- Bachelor's degree in Computer Science, Engineering, or related field, or equivalent practical experience
- 7+ years of experience in Site Reliability Engineering, DevOps, or related roles
- 3+ years in a lead or senior technical position
- Proven track record of managing large-scale production systems
- Experience with on-call rotations and incident management
- GenAI based Applications: Working knowledge of LLMs and agentic applications a plus
- Experience with serverless architectures and event-driven systems
- Familiarity with chaos engineering principles and practices
- Background in Agile/Scrum methodologies
- Experience with multi-cloud or hybrid cloud environments
The selected candidate will reside within a reasonable commuting distance, as defined by the employing Reserve Bank, and will work full-time onsite.
- Eligible Locations for Hire: Richmond, VA, San Francisco, CA
- The following Reserve Bank locations are preferred due to the concentration of System IT team members in these locations: San Francisco, and Richmond, VA
Base Salary Range: Min: $146,700 Mid: $190,500 Max: $234,300 (Location: San Francisco)
The listed salary is applicable to 12th District/San Francisco. Final offers are determined by factors including the candidate’s qualifications, internal alignment considerations, district assignment, and geographic location.
The Bank is committed to providing reasonable accommodations to individuals with disabilities to participate in the job application or interview process, perform essential job functions and receive other benefits and privileges of employment. The SF Fed is an Equal Opportunity Employer. If you need any assistance or accommodations due to a disability, please let us know at [email protected].
Full Time / Part TimeFull timeRegular / TemporaryRegularJob Exempt (Yes / No)YesJob CategoryInformation Technology Family GroupWork ShiftFirst (United States of America)The Federal Reserve Banks are committed to equal employment opportunity for employees and job applicants in compliance with applicable law and to an environment where employees are valued for their differences.
Always verify and apply to jobs on Federal Reserve System Careers (https://rb.wd5.myworkdayjobs.com/FRS) or through verified Federal Reserve Bank social media channels.
Privacy Notice
Skills Required
- Bachelor's degree in Computer Science, Engineering, or a related field, or equivalent practical experience
- 7+ years of experience in Site Reliability Engineering, DevOps, or related roles
- 3+ years in a lead or senior technical position
- Proven track record managing large-scale production systems
- Experience with on-call rotations and incident management
- Strong proficiency in Java, Python, and Node.js
- Experience with microservices architecture and distributed systems
- Solid understanding of data structures, algorithms, and design patterns
- Proficiency writing clean, maintainable, and testable code
- Extensive experience with AWS cloud infrastructure and services
- Expert-level knowledge of GitLab CI/CD pipelines, runners, and GitOps
- Advanced Terraform skills for infrastructure provisioning and management
- Experience with Docker and Kubernetes or ECS
- Proficiency with configuration management tools
- Hands-on experience with SAST tools
- Knowledge of DAST methodologies
- Understanding of OWASP Top 10 and compliance frameworks
- Experience with secrets management and IAM
- Experience with monitoring and observability tools such as Grafana, Datadog, New Relic, or similar
- Experience with log aggregation and analysis
- Experience with distributed tracing using AWS X-Ray
- Experience with serverless architectures and event-driven systems
- Familiarity with chaos engineering principles and practices
- Background in Agile/Scrum methodologies
- Experience with multi-cloud or hybrid cloud environments
- Working knowledge of LLMs and agentic applications
- AWS certification such as Solutions Architect or DevOps Engineer
What We Do
This page is dedicated to Federal Reserve System career and employment related information only. Comments not pertaining to Fed recruiting will be removed. The Fed - Make a world of difference in the global economy OUR BANK has one of the most recognizable brands around the world. The Federal Reserve is the central bank of the United States—one of the world's most influential, trusted and prestigious financial organizations. The Federal Reserve is charged with the important mission of promoting a strong economy and a stable financial system and fulfills this responsibility by formulating national monetary policy, supervising and regulating banks and bank holding companies, and providing financial services for banks and the U.S. government. OUR PEOPLE are diverse in background and ideas, which allows for ongoing creativity and innovation. Ultimately, they are the ones who push our high-performance, exchange-driven culture forward. Why Our People Choose Us: Our reputation precedes us There will always be room for personal growth Our people are first You’ll find the right balance Your responsibilities will be meaningful We hope that you will be our future colleague. Find your preferred locations around the United States and explore the breadth of opportunity available at the Federal Reserve. Atlanta https://www.frbatlanta.org/ Boston http://www.bostonfed.org/ Chicago https://www.chicagofed.org/ Cleveland https://www.clevelandfed.org/ Dallas http://dallasfed.org/ Kansas City https://www.kansascityfed.org/ Minneapolis https://www.minneapolisfed.org/ New York http://www.newyorkfed.org/ Philadelphia https://www.philadelphiafed.org/ Richmond https://www.richmondfed.org/ San Francisco http://www.frbsf.org/ St. Louis https://www.stlouisfed.org/ Board http://www.federalreserve.gov/









