Maximum of 25 job preferences reached.
Top Site Reliability Engineer Jobs
Financial Services
Designs, maintains, monitors, and optimizes applications and cloud infrastructure using SRE practices. Builds CI/CD pipelines, infrastructure as code, observability, SLO-based alerting, and automated remediation. Participates in incident response, on-call operations, post-incident reviews, and reliability improvements. Uses authorized AI tools for troubleshooting while validating recommendations and protecting sensitive operational data. Supports containerized platforms, cloud resiliency, scalability, and adoption of SRE best practices.
Top Skills:
AlbAnsibleAWSAzureBashCi/CdCloudwatchDatadogDockerDynatraceGitlab CiGrafanaJenkinsKubernetesNlbPrometheusPythonRoute 53Service MeshSplunkTerraform
Financial Services
Independently execute small-to-medium reliability projects, write maintainable code, triage and resolve incidents, remove operational toil, maintain cloud infrastructure, implement observability and SLOs, use enterprise-authorized AI for troubleshooting and post-incident analysis, and collaborate across teams to improve reliability and CI/CD practices.
Top Skills:
Ci/Cd ToolingCloud InfrastructureContainersEnterprise-Authorized Ai CapabilitiesLinuxObservability (Slo/SliTelemetry)Windows
Artificial Intelligence • Big Data • Healthtech • Software • Biotech
Owns resilience, performance, scalability, and operational cost improvements for a cloud-native platform. Designs and delivers full-stack solutions, builds monitoring and analysis capabilities, performs architectural and code reviews, and resolves complex production issues. The role requires pragmatic modernization judgment, AI-assisted development experience, customer or user accountability, and participation in a weekend on-call rotation.
Top Skills:
Ai-Assisted DevelopmentCloud-Native PlatformsDashboardsFull-Stack Web ApplicationsGrafanaLoggingMonitoring
Fintech • Financial Services
Leads Site Reliability Engineering for enterprise workplace technology platforms, driving reliability, availability, observability, incident management, automation, and operational resilience. Provides technical leadership and mentorship while designing, coding, deploying, and maintaining infrastructure solutions. Uses Python, monitoring data, AI-assisted workflows, automation, Kubernetes, and Infrastructure as Code to improve platform stability, proactively remediate issues, and reduce operational risk. Collaborates with engineering teams, leaders, customers, and vendors while ensuring compliance with technology policies and controls.
Top Skills:
Agentic AiAgileAPIsContainersElasticGenerative AiGrafanaInfrastructure As CodeKubernetesLarge Language Models (Llms)Low-Code/No-Code PlatformsPrometheusPrompt EngineeringPythonRetrieval-Augmented Generation (Rag)Splunk
Fintech • Financial Services
Leads Site Reliability Engineering for enterprise workplace technology platforms, improving stability, availability, performance, observability, and operational resilience. Provides technical leadership, develops Python automation, analyzes telemetry and outages, and designs infrastructure solutions. Drives AI-assisted incident triage, root cause analysis, proactive remediation, workflow automation, and self-healing capabilities. Collaborates with engineers, business teams, and vendors while managing risks, controls, modernization initiatives, and continuous reliability improvements.
Top Skills:
Agentic AiAgileAPIsContainersElasticGenerative AiGrafanaInfrastructure As CodeKubernetesLarge Language ModelsLow-Code/No-Code PlatformsPrometheusPrompt EngineeringPythonRetrieval-Augmented GenerationSplunk
Artificial Intelligence • Professional Services • Business Intelligence • Consulting • Cybersecurity • Generative AI
Leads teams designing and deploying AI-driven enterprise and cloud security solutions. Responsibilities include developing machine learning systems, integrating data infrastructure and pipelines, performing advanced modeling and analysis, and building AI applications with Python, Java, and C++. The role involves client engagement, strategic problem-solving, stakeholder validation, coaching, process innovation, and operational excellence while addressing complex cybersecurity and privacy challenges.
Top Skills:
Ai SystemsAWSC++Cloud SecurityData EngineeringData PipelinesDatabricksGCPJavaMachine LearningAzurePythonScikit-LearnSnowflakeTensorFlow
Artificial Intelligence • Professional Services • Business Intelligence • Consulting • Cybersecurity • Generative AI
Leads the design and development of AI-driven enterprise and cloud security solutions. Manages teams delivering data, analytics, and machine learning engineering projects; oversees AI model implementation, data infrastructure, pipelines, and platform deployments. Uses Python and C++ for algorithm development, analyzes data, improves data quality, ensures technical compliance, mentors staff, manages client expectations, and drives innovation across cybersecurity and privacy initiatives.
Top Skills:
AWSC++DatabricksGCPAzurePythonSnowflake
Fintech • Financial Services
Leads Forge’s Site Reliability Engineering team, improving system availability, observability, incident response, disaster recovery, automation, and production operations. Partners with Platform, Engineering, Security, Compliance, Risk, and Product teams to strengthen reliability and operational maturity in a regulated environment. Responsibilities include technical design, troubleshooting, team hiring, coaching, performance management, and career development while helping deliver secure, scalable, highly reliable products.
Top Skills:
Amazon CloudwatchAnsibleAWSAzureCi/CdCloud InfrastructureDatadogDistributed SystemsInfrastructure-As-CodeKubernetesObservabilityTerraform
Fintech • Financial Services
Leads Forge’s Site Reliability Engineering team and drives system availability, incident management, observability, disaster recovery, automation, and production support. Partners with Platform, Engineering, Security, Compliance, Risk, and Product teams to improve reliability and operational maturity in a regulated environment. The role combines hands-on technical leadership with hiring, coaching, performance management, architecture contributions, troubleshooting, and development of secure, scalable infrastructure.
Top Skills:
Amazon CloudwatchAnsibleAWSAzureCi/CdDatadogInfrastructure As CodeKubernetesTerraform
Machine Learning • Payments • Security • Software • Financial Services
Administer and stabilize enterprise Windows Server and VMware environments, responding to incidents, patching, vulnerabilities, monitoring alerts, and production outages. Perform root cause analysis, capacity planning, disaster recovery testing, and resiliency improvements. Develop automation and monitoring dashboards, implement SLAs and SLOs, support change management, and mentor junior engineers while improving infrastructure reliability and reducing incident resolution times.
Top Skills:
Microsoft Windows ServerPowershellVMware
Aerospace • Information Technology • Software • Cybersecurity • Design • Defense • Manufacturing
Operate and support mission-critical developer platforms (GitLab, CI/CD runners, Jira, Confluence, PostgreSQL, Artifactory, SonarQube). Monitor health, triage incidents, automate operations with IaC/Ansible/scripts, assist cloud/on-prem administration (AWS/Azure/Linux/containers), maintain runbooks, support backups/patching, and collaborate with developers and cybersecurity to apply SRE practices and improve delivery metrics.
Top Skills:
AnsibleArtifactoryAWSAzure DevopsBashC#C++ConfluenceContainer PlatformsGitlabGitlab Ci/CdGitlab RunnersInfrastructure As Code (Iac)JavaJenkinsJIRALinuxAzurePostgresPowershellPythonSonarqubeSQLVirtualization
5 Days AgoSaved
Artificial Intelligence • Cloud • Fintech • Information Technology • Insurance • Financial Services • Big Data Analytics
Leads a cloud SRE shared-services organization supporting AWS application platforms. Responsibilities include managing and mentoring engineers, modernizing infrastructure, governing Terraform modules and AWS patterns, building self-service provisioning and Harness CI/CD automation, establishing SLOs and error budgets, resolving critical incidents, implementing observability, improving resiliency, and planning disaster recovery. The leader also drives GenAI adoption, stakeholder alignment, security, governance, cost optimization, and engineering best practices across technology teams.
Top Skills:
Agentic AiAmazon CloudwatchAWSCi/CdGenerative AiHarnessInfrastructure As CodeJfrogKubernetesNew RelicTerraform Hcl
New
Track Smarter, Apply Better.
Ditch the spreadsheets. Organize your job search with our freeApplication Tracker.
Use For Free
Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
Leads AI-assisted site reliability engineering across Azure and AWS. Designs observability, incident response, automation, resiliency testing, disaster recovery, chaos engineering, and recovery-validation capabilities. Establishes OpenTelemetry, SLI, SLO, error-budget, and reliability-scorecard standards; improves alert quality and operational insights; creates human-in-the-loop mitigation workflows; and mentors engineers while driving cross-functional reliability improvements.
Top Skills:
AnsibleAWSAzureDatadogGrafanaHelmKubernetesLlmsOpentelemetryPrometheusPulumiRagTerraform
Fintech • Financial Services
Lead Site Reliability Engineer responsible for driving stability, resiliency, performance, and security of enterprise platforms. Define SLIs/SLOs, implement observability and automation, lead incident management and toil reduction, advise leadership, enforce reliability standards, and collaborate across development, product, and operations to ensure high availability and continuous improvement.
Top Skills:
AppdynamicsGrafanaLinuxOpenshift Container PlatformOpentelemetrySplunkSplunk ObservabilityWindows
Fintech • Financial Services
Lead SRE responsible for platform stability, resiliency, performance, and security. Define SLIs/SLOs, drive observability and automation, lead incident response and toil reduction, collaborate across teams, and enforce reliability standards for cloud and on‑prem environments.
Top Skills:
AngularAppdynamicsAWSAzureGlassboxGrafanaJavaJavaScriptJSONKubernetesLinuxNode.jsOpenshift Container PlatformOpentelemetryPcfPksPrometheusPythonRubyShell ScriptingSplunkSplunk ObservabilityVMwareWindows
Digital Media • Information Technology • News + Entertainment
Own and support production infrastructure for FreeWheel's Streaming Hub. Design and implement cloud and Kubernetes infrastructure, automate with Terraform/Ansible and Python/Go, troubleshoot incidents, improve observability and CI/CD, and ensure platform reliability and scalability during high-traffic streaming events.
Top Skills:
Amazon EksAnsibleAWSDockerGo (Golang)IamJenkinsKubernetesLoad BalancerOracle Cloud Infrastructure (Oci)PythonRoute 53Security GroupsTerraformVpc
Digital Media • Information Technology • News + Entertainment
Lead Data SRE responsible for ensuring reliability, scalability, and performance of data platforms. Design monitoring, automate operations, optimize performance, respond to incidents, plan capacity, enforce security/compliance, document systems, and collaborate with engineering, data science, and product teams.
Top Skills:
AerospikeAnsibleApache KafkaAWSAws S3AzureCassandraContainerizationDockerElk StackGCPGoGrafanaHadoopHdfsJavaKubernetesMicroservicesMySQLNoSQLPostgresPrometheusPythonScalaSnowflakeSparkTerraform
Artificial Intelligence • Legal Tech • Software
The Staff Site Reliability Engineer will architect reliability strategies, manage observability, and drive operational excellence across teams, particularly in large-scale production systems.
Top Skills:
Cloud InfrastructureDistributed Systems
Financial Services
Administer and improve global block, file, and object storage across on-premises and cloud environments. Manage performance, capacity, resiliency, replication, backup, disaster recovery, incident response, root-cause analysis, and operational automation. Build observability and AI-assisted AIOps workflows with validation, guardrails, auditability, and human review. Maintain runbooks, SLOs, SLIs, error budgets, and on-call readiness while partnering with infrastructure and application teams.
Top Skills:
AnsibleBashCephChefCloud-Native StorageCloudFormationContainer Storage Interface (Csi)DatadogDell Emc IsilonDell Emc PowerstoreElasticGoGrafanaHardware Security ModulesHitachiIbm StorageKafkaKey Management ServicesKubernetesLinuxNetappOpensearchOpentelemetryPrometheusPuppetPure StoragePythonSecrets ManagementServicenowSplunkTerraform
Financial Services
Lead and scale Site Reliability Engineering practices across application and platform teams. Own non-functional requirements, drive resiliency, security, monitoring, automation, and incident post-mortems. Coach engineers, influence stakeholders, adopt AI-assisted reliability workflows, and measure reliability through stability metrics while ensuring traceability, auditability, and guardrails.
Top Skills:
AICi/CdContainersDockerEcsGitlabGoGraphQLJavaScriptJenkinsKafkaKubernetesOpentelemetryPythonTerraform
Financial Services
Lead SRE role owning reliability, resiliency design, incident leadership and mentorship. Drive SRE practices, observability, CI/CD and container orchestration improvements, and responsibly adopt enterprise AI to accelerate incident triage and operational workflows.
Top Skills:
.NetCi/CdContainer OrchestrationContainersEnterprise AiJavaNetworkingObservabilityPythonService Level Objective AlertingSpring BootTelemetry Collection
Financial Services
Lead SRE role responsible for driving reliability, resiliency design, incident leadership, and mentoring. Improve stability via data-driven analytics, define SLOs/error budgets, adopt AI-assisted SRE workflows with proper guardrails, and lead CI/CD, observability, containerization, and operational readiness across the SDLC.
Top Skills:
.NetCi/CdContainer OrchestrationDockerEnterprise Ai ToolsJavaKubernetesMonitoringNetworkingObservabilityPythonSlosSpring BootTelemetry
Financial Services
Lead SRE responsible for driving reliability culture, conducting resiliency design reviews, leading incident response, mentoring engineers, defining SLOs/error budgets, adopting AI-assisted reliability workflows, and improving observability, CI/CD, and containerized platforms to ensure scalable, secure, and resilient production services.
Top Skills:
.NetCi/CdContainer OrchestrationContainersEnterprise AiJavaNetworkingObservabilityPythonSpring BootTelemetry
Financial Services
Lead SRE for Sales Execution platforms responsible for stability, availability, resiliency, incident leadership, RCA, and operational maturity. Partner with Front Office, Product, Development, and Infrastructure to drive SRE adoption, observability, automation, and AI-assisted reliability workflows while mentoring engineers and owning outcomes for business-critical services.
Top Skills:
AnsibleAWSAzureCi/CdContainersDynatraceEnterprise-Authorized AiGCPGeneosGrafanaItilKubernetesMicroservicesOpenshiftPowershellPythonShellSplunkTerraform
Big Data • Fintech • Information Technology • Business Intelligence • Financial Services • Cybersecurity • Big Data Analytics
The Staff Site Reliability Engineer will lead reliability strategies, manage high-risk initiatives, and enhance engineering standards while ensuring system reliability and operational excellence within a hybrid work environment.
Top Skills:
BashCi/CdDatabase ArchitectureGoGoogle Cloud PlatformInfrastructure-As-CodeKubernetesMonitoring PlatformsPulumiPythonTerraform
Let Your Resume Do The Work
Upload your resume to be matched with jobs you're a great fit for.
Success! We'll use this to further personalize your experience.
Top Companies Hiring Site Reliability Engineers
See AllPopular Job Searches
All Software Engineer Jobs
.NET Developer Jobs
Aerospace Thermal Engineering Jobs
AI Engineer Jobs
Android Developer Jobs
Automation Engineer Jobs
Backend Developer Jobs
Blockchain Developer Jobs
C# Jobs
C++ Jobs
Cloud Architect Jobs
Cloud Engineer Jobs
Design Engineer Jobs
DevOps Engineer Jobs
Director Of Engineering Jobs
Electrical Engineering Jobs
Embedded Software Engineer Jobs
Engineering Jobs
Engineering Manager Jobs
Environmental Engineering Jobs
Field Engineer Jobs
Front End Developer Jobs
Full Stack Developer Jobs
Game Developer Jobs
Golang Jobs
Hardware Engineer Jobs
Industrial Engineering Jobs
iOS Developer Jobs
Java Developer Jobs
Javascript Developer Jobs
Linux Jobs
Manufacturing Engineer Jobs
Mechanical Engineering Jobs
Network Engineer Jobs
PHP Developer Jobs
Process Engineer Jobs
Project Engineer Jobs
Prompt Engineering Jobs
Python Jobs
QA Jobs
Robotics Engineer Jobs
Ruby on Rails Jobs
Salesforce Administrator Jobs
Salesforce Developer Jobs
Scala Jobs
Sharepoint Developer Jobs
Site Reliability Engineer Jobs
Software Engineering Manager Jobs
Solutions Architect Jobs
SQL Developer Jobs
Structural Engineer Jobs
System Engineer Jobs
Test Engineer Jobs
Web Developer Jobs
All Filters
Total selected ()
No Results
No Results























