Top Site Reliability Engineer Jobs

YesterdaySaved
Hybrid
Jersey City, NJ, USA
Mid level
Mid level
Financial Services
Designs, maintains, monitors, and optimizes applications and cloud infrastructure using SRE practices. Builds CI/CD pipelines, infrastructure as code, observability, SLO-based alerting, and automated remediation. Participates in incident response, on-call operations, post-incident reviews, and reliability improvements. Uses authorized AI tools for troubleshooting while validating recommendations and protecting sensitive operational data. Supports containerized platforms, cloud resiliency, scalability, and adoption of SRE best practices.
Top Skills: AlbAnsibleAWSAzureBashCi/CdCloudwatchDatadogDockerDynatraceGitlab CiGrafanaJenkinsKubernetesNlbPrometheusPythonRoute 53Service MeshSplunkTerraform
Reposted YesterdaySaved
Hybrid
Chicago, IL, USA
Junior
Junior
Financial Services
Independently execute small-to-medium reliability projects, write maintainable code, triage and resolve incidents, remove operational toil, maintain cloud infrastructure, implement observability and SLOs, use enterprise-authorized AI for troubleshooting and post-incident analysis, and collaborate across teams to improve reliability and CI/CD practices.
Top Skills: Ci/Cd ToolingCloud InfrastructureContainersEnterprise-Authorized Ai CapabilitiesLinuxObservability (Slo/SliTelemetry)Windows
3 Days AgoSaved
Hybrid
Boston, MA, USA
88K-168K Annually
Entry level
88K-168K Annually
Entry level
Artificial Intelligence • Big Data • Healthtech • Software • Biotech
Owns resilience, performance, scalability, and operational cost improvements for a cloud-native platform. Designs and delivers full-stack solutions, builds monitoring and analysis capabilities, performs architectural and code reviews, and resolves complex production issues. The role requires pragmatic modernization judgment, AI-assisted development experience, customer or user accountability, and participation in a weekend on-call rotation.
Top Skills: Ai-Assisted DevelopmentCloud-Native PlatformsDashboardsFull-Stack Web ApplicationsGrafanaLoggingMonitoring
3 Days AgoSaved
Hybrid
Iselin, NJ, USA
119K-224K Annually
Senior level
119K-224K Annually
Senior level
Fintech • Financial Services
Leads Site Reliability Engineering for enterprise workplace technology platforms, driving reliability, availability, observability, incident management, automation, and operational resilience. Provides technical leadership and mentorship while designing, coding, deploying, and maintaining infrastructure solutions. Uses Python, monitoring data, AI-assisted workflows, automation, Kubernetes, and Infrastructure as Code to improve platform stability, proactively remediate issues, and reduce operational risk. Collaborates with engineering teams, leaders, customers, and vendors while ensuring compliance with technology policies and controls.
Top Skills: Agentic AiAgileAPIsContainersElasticGenerative AiGrafanaInfrastructure As CodeKubernetesLarge Language Models (Llms)Low-Code/No-Code PlatformsPrometheusPrompt EngineeringPythonRetrieval-Augmented Generation (Rag)Splunk
3 Days AgoSaved
Hybrid
Irving, TX, USA
119K-224K Annually
Senior level
119K-224K Annually
Senior level
Fintech • Financial Services
Leads Site Reliability Engineering for enterprise workplace technology platforms, improving stability, availability, performance, observability, and operational resilience. Provides technical leadership, develops Python automation, analyzes telemetry and outages, and designs infrastructure solutions. Drives AI-assisted incident triage, root cause analysis, proactive remediation, workflow automation, and self-healing capabilities. Collaborates with engineers, business teams, and vendors while managing risks, controls, modernization initiatives, and continuous reliability improvements.
Top Skills: Agentic AiAgileAPIsContainersElasticGenerative AiGrafanaInfrastructure As CodeKubernetesLarge Language ModelsLow-Code/No-Code PlatformsPrometheusPrompt EngineeringPythonRetrieval-Augmented GenerationSplunk
3 Days AgoSaved
Hybrid
9 Locations
124K-280K Annually
Senior level
124K-280K Annually
Senior level
Artificial Intelligence • Professional Services • Business Intelligence • Consulting • Cybersecurity • Generative AI
Leads teams designing and deploying AI-driven enterprise and cloud security solutions. Responsibilities include developing machine learning systems, integrating data infrastructure and pipelines, performing advanced modeling and analysis, and building AI applications with Python, Java, and C++. The role involves client engagement, strategic problem-solving, stakeholder validation, coaching, process innovation, and operational excellence while addressing complex cybersecurity and privacy challenges.
Top Skills: Ai SystemsAWSC++Cloud SecurityData EngineeringData PipelinesDatabricksGCPJavaMachine LearningAzurePythonScikit-LearnSnowflakeTensorFlow
3 Days AgoSaved
Hybrid
9 Locations
99K-232K Annually
Senior level
99K-232K Annually
Senior level
Artificial Intelligence • Professional Services • Business Intelligence • Consulting • Cybersecurity • Generative AI
Leads the design and development of AI-driven enterprise and cloud security solutions. Manages teams delivering data, analytics, and machine learning engineering projects; oversees AI model implementation, data infrastructure, pipelines, and platform deployments. Uses Python and C++ for algorithm development, analyzes data, improves data quality, ensures technical compliance, mentors staff, manages client expectations, and drives innovation across cybersecurity and privacy initiatives.
Top Skills: AWSC++DatabricksGCPAzurePythonSnowflake
4 Days AgoSaved
Easy Apply
Hybrid
San Francisco, CA, USA
Easy Apply
150K-220K Annually
Expert/Leader
150K-220K Annually
Expert/Leader
Fintech • Financial Services
Leads Forge’s Site Reliability Engineering team, improving system availability, observability, incident response, disaster recovery, automation, and production operations. Partners with Platform, Engineering, Security, Compliance, Risk, and Product teams to strengthen reliability and operational maturity in a regulated environment. Responsibilities include technical design, troubleshooting, team hiring, coaching, performance management, and career development while helping deliver secure, scalable, highly reliable products.
Top Skills: Amazon CloudwatchAnsibleAWSAzureCi/CdCloud InfrastructureDatadogDistributed SystemsInfrastructure-As-CodeKubernetesObservabilityTerraform
4 Days AgoSaved
Easy Apply
Hybrid
New York, NY, USA
Easy Apply
150K-220K Annually
Expert/Leader
150K-220K Annually
Expert/Leader
Fintech • Financial Services
Leads Forge’s Site Reliability Engineering team and drives system availability, incident management, observability, disaster recovery, automation, and production support. Partners with Platform, Engineering, Security, Compliance, Risk, and Product teams to improve reliability and operational maturity in a regulated environment. The role combines hands-on technical leadership with hiring, coaching, performance management, architecture contributions, troubleshooting, and development of secure, scalable infrastructure.
Top Skills: Amazon CloudwatchAnsibleAWSAzureCi/CdDatadogInfrastructure As CodeKubernetesTerraform
4 Days AgoSaved
Hybrid
Pittsburgh, PA, USA
86K-158K Annually
Senior level
86K-158K Annually
Senior level
Machine Learning • Payments • Security • Software • Financial Services
Administer and stabilize enterprise Windows Server and VMware environments, responding to incidents, patching, vulnerabilities, monitoring alerts, and production outages. Perform root cause analysis, capacity planning, disaster recovery testing, and resiliency improvements. Develop automation and monitoring dashboards, implement SLAs and SLOs, support change management, and mentor junior engineers while improving infrastructure reliability and reducing incident resolution times.
Top Skills: Microsoft Windows ServerPowershellVMware
Reposted 4 Days AgoSaved
In-Office
Berkeley, MO, USA
99K-171K Annually
Junior
99K-171K Annually
Junior
Aerospace • Information Technology • Software • Cybersecurity • Design • Defense • Manufacturing
Operate and support mission-critical developer platforms (GitLab, CI/CD runners, Jira, Confluence, PostgreSQL, Artifactory, SonarQube). Monitor health, triage incidents, automate operations with IaC/Ansible/scripts, assist cloud/on-prem administration (AWS/Azure/Linux/containers), maintain runbooks, support backups/patching, and collaborate with developers and cybersecurity to apply SRE practices and improve delivery metrics.
Top Skills: AnsibleArtifactoryAWSAzure DevopsBashC#C++ConfluenceContainer PlatformsGitlabGitlab Ci/CdGitlab RunnersInfrastructure As Code (Iac)JavaJenkinsJIRALinuxAzurePostgresPowershellPythonSonarqubeSQLVirtualization
5 Days AgoSaved
Hybrid
New York, NY, USA
185K-265K Annually
Senior level
185K-265K Annually
Senior level
Artificial Intelligence • Cloud • Fintech • Information Technology • Insurance • Financial Services • Big Data Analytics
Leads a cloud SRE shared-services organization supporting AWS application platforms. Responsibilities include managing and mentoring engineers, modernizing infrastructure, governing Terraform modules and AWS patterns, building self-service provisioning and Harness CI/CD automation, establishing SLOs and error budgets, resolving critical incidents, implementing observability, improving resiliency, and planning disaster recovery. The leader also drives GenAI adoption, stakeholder alignment, security, governance, cost optimization, and engineering best practices across technology teams.
Top Skills: Agentic AiAmazon CloudwatchAWSCi/CdGenerative AiHarnessInfrastructure As CodeJfrogKubernetesNew RelicTerraform Hcl
New

Track Smarter, Apply Better.

Ditch the spreadsheets. Organize your job search with our freeApplication Tracker.

Use For Free
Application Tracker Preview
5 Days AgoSaved
In-Office or Remote
Eden Prairie, MN, USA
135K-231K Annually
Expert/Leader
135K-231K Annually
Expert/Leader
Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
Leads AI-assisted site reliability engineering across Azure and AWS. Designs observability, incident response, automation, resiliency testing, disaster recovery, chaos engineering, and recovery-validation capabilities. Establishes OpenTelemetry, SLI, SLO, error-budget, and reliability-scorecard standards; improves alert quality and operational insights; creates human-in-the-loop mitigation workflows; and mentors engineers while driving cross-functional reliability improvements.
Top Skills: AnsibleAWSAzureDatadogGrafanaHelmKubernetesLlmsOpentelemetryPrometheusPulumiRagTerraform
Reposted 5 Days AgoSaved
Hybrid
Saint Louis, MO, USA
119K-187K Annually
Senior level
119K-187K Annually
Senior level
Fintech • Financial Services
Lead Site Reliability Engineer responsible for driving stability, resiliency, performance, and security of enterprise platforms. Define SLIs/SLOs, implement observability and automation, lead incident management and toil reduction, advise leadership, enforce reliability standards, and collaborate across development, product, and operations to ensure high availability and continuous improvement.
Top Skills: AppdynamicsGrafanaLinuxOpenshift Container PlatformOpentelemetrySplunkSplunk ObservabilityWindows
Reposted 5 Days AgoSaved
Hybrid
Charlotte, NC, USA
119K-187K Annually
Senior level
119K-187K Annually
Senior level
Fintech • Financial Services
Lead SRE responsible for platform stability, resiliency, performance, and security. Define SLIs/SLOs, drive observability and automation, lead incident response and toil reduction, collaborate across teams, and enforce reliability standards for cloud and on‑prem environments.
Top Skills: AngularAppdynamicsAWSAzureGlassboxGrafanaJavaJavaScriptJSONKubernetesLinuxNode.jsOpenshift Container PlatformOpentelemetryPcfPksPrometheusPythonRubyShell ScriptingSplunkSplunk ObservabilityVMwareWindows
Reposted 6 Days AgoSaved
Hybrid
Reston, VA, USA
109K-164K Annually
Mid level
109K-164K Annually
Mid level
Digital Media • Information Technology • News + Entertainment
Own and support production infrastructure for FreeWheel's Streaming Hub. Design and implement cloud and Kubernetes infrastructure, automate with Terraform/Ansible and Python/Go, troubleshoot incidents, improve observability and CI/CD, and ensure platform reliability and scalability during high-traffic streaming events.
Top Skills: Amazon EksAnsibleAWSDockerGo (Golang)IamJenkinsKubernetesLoad BalancerOracle Cloud Infrastructure (Oci)PythonRoute 53Security GroupsTerraformVpc
Reposted 6 Days AgoSaved
Hybrid
Reston, VA, USA
152K-228K Annually
Expert/Leader
152K-228K Annually
Expert/Leader
Digital Media • Information Technology • News + Entertainment
Lead Data SRE responsible for ensuring reliability, scalability, and performance of data platforms. Design monitoring, automate operations, optimize performance, respond to incidents, plan capacity, enforce security/compliance, document systems, and collaborate with engineering, data science, and product teams.
Top Skills: AerospikeAnsibleApache KafkaAWSAws S3AzureCassandraContainerizationDockerElk StackGCPGoGrafanaHadoopHdfsJavaKubernetesMicroservicesMySQLNoSQLPostgresPrometheusPythonScalaSnowflakeSparkTerraform
Reposted 6 Days AgoSaved
In-Office
New York City, NY, USA
273K-369K Annually
Senior level
273K-369K Annually
Senior level
Artificial Intelligence • Legal Tech • Software
The Staff Site Reliability Engineer will architect reliability strategies, manage observability, and drive operational excellence across teams, particularly in large-scale production systems.
Top Skills: Cloud InfrastructureDistributed Systems
Mid level
Financial Services
Administer and improve global block, file, and object storage across on-premises and cloud environments. Manage performance, capacity, resiliency, replication, backup, disaster recovery, incident response, root-cause analysis, and operational automation. Build observability and AI-assisted AIOps workflows with validation, guardrails, auditability, and human review. Maintain runbooks, SLOs, SLIs, error budgets, and on-call readiness while partnering with infrastructure and application teams.
Top Skills: AnsibleBashCephChefCloud-Native StorageCloudFormationContainer Storage Interface (Csi)DatadogDell Emc IsilonDell Emc PowerstoreElasticGoGrafanaHardware Security ModulesHitachiIbm StorageKafkaKey Management ServicesKubernetesLinuxNetappOpensearchOpentelemetryPrometheusPuppetPure StoragePythonSecrets ManagementServicenowSplunkTerraform
Reposted 7 Days AgoSaved
Hybrid
Plano, TX, USA
Senior level
Senior level
Financial Services
Lead and scale Site Reliability Engineering practices across application and platform teams. Own non-functional requirements, drive resiliency, security, monitoring, automation, and incident post-mortems. Coach engineers, influence stakeholders, adopt AI-assisted reliability workflows, and measure reliability through stability metrics while ensuring traceability, auditability, and guardrails.
Top Skills: AICi/CdContainersDockerEcsGitlabGoGraphQLJavaScriptJenkinsKafkaKubernetesOpentelemetryPythonTerraform
Reposted 7 Days AgoSaved
Hybrid
Plano, TX, USA
Senior level
Senior level
Financial Services
Lead SRE role owning reliability, resiliency design, incident leadership and mentorship. Drive SRE practices, observability, CI/CD and container orchestration improvements, and responsibly adopt enterprise AI to accelerate incident triage and operational workflows.
Top Skills: .NetCi/CdContainer OrchestrationContainersEnterprise AiJavaNetworkingObservabilityPythonService Level Objective AlertingSpring BootTelemetry Collection
Reposted 7 Days AgoSaved
Hybrid
Jersey City, NJ, USA
Senior level
Senior level
Financial Services
Lead SRE role responsible for driving reliability, resiliency design, incident leadership, and mentoring. Improve stability via data-driven analytics, define SLOs/error budgets, adopt AI-assisted SRE workflows with proper guardrails, and lead CI/CD, observability, containerization, and operational readiness across the SDLC.
Top Skills: .NetCi/CdContainer OrchestrationDockerEnterprise Ai ToolsJavaKubernetesMonitoringNetworkingObservabilityPythonSlosSpring BootTelemetry
Reposted 7 Days AgoSaved
Hybrid
2 Locations
Senior level
Senior level
Financial Services
Lead SRE responsible for driving reliability culture, conducting resiliency design reviews, leading incident response, mentoring engineers, defining SLOs/error budgets, adopting AI-assisted reliability workflows, and improving observability, CI/CD, and containerized platforms to ensure scalable, secure, and resilient production services.
Top Skills: .NetCi/CdContainer OrchestrationContainersEnterprise AiJavaNetworkingObservabilityPythonSpring BootTelemetry
Reposted 7 Days AgoSaved
Hybrid
New York, NY, USA
Senior level
Senior level
Financial Services
Lead SRE for Sales Execution platforms responsible for stability, availability, resiliency, incident leadership, RCA, and operational maturity. Partner with Front Office, Product, Development, and Infrastructure to drive SRE adoption, observability, automation, and AI-assisted reliability workflows while mentoring engineers and owning outcomes for business-critical services.
Top Skills: AnsibleAWSAzureCi/CdContainersDynatraceEnterprise-Authorized AiGCPGeneosGrafanaItilKubernetesMicroservicesOpenshiftPowershellPythonShellSplunkTerraform
Reposted 8 Days AgoSaved
Hybrid
6 Locations
113K-188K Annually
Senior level
113K-188K Annually
Senior level
Big Data • Fintech • Information Technology • Business Intelligence • Financial Services • Cybersecurity • Big Data Analytics
The Staff Site Reliability Engineer will lead reliability strategies, manage high-risk initiatives, and enhance engineering standards while ensuring system reliability and operational excellence within a hybrid work environment.
Top Skills: BashCi/CdDatabase ArchitectureGoGoogle Cloud PlatformInfrastructure-As-CodeKubernetesMonitoring PlatformsPulumiPythonTerraform
All Filters
JobType
New Jobs
Job Category
Experience
Industry
Company Name
Company Size

Sign up now Access later

Create Free Account