Maximum of 25 job preferences reached.
Top Site Reliability Engineer Jobs
Information Technology • Software • Financial Services • Quantitative Trading
The Site Reliability Engineer will provide support and diagnose issues within a real-time, distributed environment, focusing on large-scale application and infrastructure management, with basic required skills in UNIX/Linux, networking, SQL, and scripting languages.
Top Skills:
BashPythonSQLTcp/IpUdpUnix/Linux
Fintech • Information Technology • Financial Services
Own reliability, stability, and performance of low-latency algorithmic and sequencer-based trading platforms. Troubleshoot production issues, manage releases and configurations, improve observability and automation, support client onboarding (FIX/API), conduct performance testing, mentor peers, and drive SRE best practices across teams.
Top Skills:
AWSBashContainerizationFix ProtocolGithub CopilotJavaJenkinsJvmKshLinuxMonitoring/Observability ToolsPythonSQL
Healthtech • Payments • Software
The Senior SRE I will design and maintain automation for infrastructure provisioning, monitor system health, resolve production incidents, and mentor junior SREs, ensuring reliability and operational efficiency across cloud platforms.
Top Skills:
AnsibleAWSAzureBashCloudFormationDatadogDockerGCPGithub ActionsGitlab CiGoGrafanaJavaJenkinsKubernetesPrometheusPythonRubySplunkTerraform
Information Technology • Insurance • Software
Define and own enterprise reliability, scalability, and performance for production services. Drive architectural standards, observability strategy, SLO/SLI and error-budget governance, lead incident command for high-severity events, and foster a blameless, engineering-first operations culture across cloud, hybrid data centers, and customer-hosted environments.
Top Skills:
.NetAWSC#Ci/CdInfrastructure-As-CodeJavaKubernetesLinuxObservabilityPythonReactRelational DatabasesWindows
Artificial Intelligence • Marketing Tech • Software
Lead technical reliability initiatives across a multi-cloud, multi-region active-active content platform. Architect and evolve core services, observability and logging, automation and capacity planning. Mentor engineers, drive cross-team reliability projects, define standards (IaC, SLOs, on-call) and proactively improve platform scalability and incident outcomes.
Top Skills:
Apache KafkaApache PulsarAWSCassandraChefEksGCPGkeGoGrafana AlloyGrafana LokiKubernetesLinuxNode.jsPrometheusPythonRubyScylladbShell ScriptingTempoTerraformThanos
Fintech • Financial Services
Lead and consult on large-scale systems and network infrastructure, resolve complex production issues, drive technical changes, and collaborate with engineering teams to improve reliability and observability across cloud and distributed platforms.
Top Skills:
AiopsAirflowArtifactoryAWSAzureBigpandaElastic ApmElasticsearchGitGradleGrafanaGroovyHarness IoJaegerJenkinsKafkaKibanaKubernetesLinuxLogstashMavenNetcoolOcpPksRemedyServicenowSpinnakerTerraformUdeployUnixVMwareWindowsZipkin
Fintech • Financial Services
Lead SRE responsible for leading large-scale systems and network infrastructure initiatives, resolving complex production issues, consulting on change/design, and improving observability and hosting platform reliability across cloud and distributed environments while collaborating with technical peers and managers.
Top Skills:
AiopsAirflowArtifactoryAWSAzureBig PandaElastic ApmElasticsearchGitGradleGrafanaGroovyHarness IoJaegerJenkinsKafkaKibanaKubernetesLinuxLogstashMavenNetcoolOcpPksRemedyServicenowSpinnakerTerraformUdeployUnixVMwareWindowsZipkin
Fintech • Financial Services
Lead and advise on large-scale systems and network infrastructure initiatives, resolve escalated SRE issues, design technical changes, and collaborate with engineering and management to improve observability, hosting platforms, and production reliability.
Top Skills:
AiopsAirflowArtifactoryAWSAzureBigpandaElastic ApmElasticsearchGitGradleGrafanaGroovyHarness IoJaegerJenkinsKafkaKibanaKubernetesLinuxLogstashMavenNetcoolOcpPksRemedyServicenowSpinnakerTerraformUdeployUnixVMwareWindowsZipkin
Fintech • Financial Services
Lead and consult on large-scale systems and network infrastructure planning, drive complex SRE initiatives, analyze escalated technical issues, make technical change decisions, and collaborate with engineering and management to resolve systems support problems and improve reliability.
Top Skills:
AiopsAirflowArtifactoryAWSAzureBig PandaElastic ApmElasticsearchGitGradleGrafanaGroovyHarness IoJaegerJenkinsKafkaKibanaKubernetesLinuxLogstashMavenNetcoolOcpPksRemedyServicenowSpinnakerTerraformUdeployUnixVMwareWindowsZipkin
Artificial Intelligence • Marketing Tech • Software
Lead SRE to provide technical leadership for a multi-cloud, multi-region active-active content platform. Define automation and IaC strategy, design core platform and logging/observability systems, establish capacity and performance frameworks, lead cross-team reliability initiatives, improve on-call and incident response, mentor engineers, and proactively address systemic reliability and scaling challenges.
Top Skills:
Apache KafkaApache PulsarAWSCassandraChefEksGCPGkeGoGrafana AlloyGrafana LokiKubernetesLinuxNode.jsPrometheusPythonRubyScylladbShell ScriptingTempoTerraformThanos
Aerospace • Information Technology • Software • Cybersecurity • Design • Defense • Manufacturing
Lead technical strategy and architecture for developer tooling and platforms (GitLab, CI/CD, Jira, Confluence, PostgreSQL, Artifactory, SonarQube). Define SLIs/SLOs, reliability standards, automation, IaC, backups, DR, and security controls. Lead incidents, root-cause analysis, upgrades, migrations, and mentor SREs while partnering with stakeholders to ensure secure, scalable, and supportable platforms. Participate in after-hours escalation as needed.
Top Skills:
AnsibleArtifactoryAWSAzure DevopsCi/CdConfluenceContainersDockerGitlabGitlab Ci/CdGCPInfrastructure As CodeJenkinsJIRAKubernetesAzureMonitoring/ObservabilityPostgresRunnersSecrets ManagementSonarqubeVirtualization
Aerospace • Information Technology • Software • Cybersecurity • Design • Defense • Manufacturing
Operate and support mission-critical developer platforms (GitLab, CI/CD runners, Jira, Confluence, PostgreSQL, Artifactory, SonarQube). Monitor health, triage incidents, automate operations with IaC/Ansible/scripts, assist cloud/on-prem administration (AWS/Azure/Linux/containers), maintain runbooks, support backups/patching, and collaborate with developers and cybersecurity to apply SRE practices and improve delivery metrics.
Top Skills:
AnsibleArtifactoryAWSAzure DevopsBashC#C++ConfluenceContainer PlatformsGitlabGitlab Ci/CdGitlab RunnersInfrastructure As Code (Iac)JavaJenkinsJIRALinuxAzurePostgresPowershellPythonSonarqubeSQLVirtualization
New
Track Smarter, Apply Better.
Ditch the spreadsheets. Organize your job search with our freeApplication Tracker.
Use For Free
Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
Define and scale SRE standards across teams, implement SLOs/SLIs/error budgets, build observability and resiliency patterns, drive automation and AIOps, improve reliability for large-scale Azure cloud systems, and influence engineering and platform teams.
Top Skills:
Ai/MlAiopsAutomationAzureError BudgetsIncident ManagementLogsObservability (MetricsOpentelemetrySlisSlosTracing)
Artificial Intelligence • Information Technology • Machine Learning • Natural Language Processing • Productivity • Software • Generative AI
The SRE will ensure the reliability of backend systems, scale Kubernetes-based control planes, and improve automation mechanisms while managing incident processes.
Top Skills:
AWSAzureDockerGCPJavaKubernetesLinuxTerraform
Reposted 3 Days AgoSaved
Easy Apply
Easy Apply
Big Data • Cloud • Software • Database
Maintain and improve multi-cloud Kubernetes infrastructure, CI/CD (Argo Workflows/ArgoCD), observability, and networking. Build reliable continuous deployment tooling and onboarding flows, provide internal support, collaborate across Platform Engineering, contribute upstream (open-source/operators), and participate in a 24/7 on-call rotation to resolve deployment infrastructure issues.
Top Skills:
AlertingArgo WorkflowsArgocdAWSAzureCi/CdContainersDnsGCPGoKubernetesLinuxLoad BalancerObservabilityPythonService MeshTcp/IpTls
Blockchain • Fintech • Payments • Consulting • Cryptocurrency • Cybersecurity • Quantum Computing
Ensure reliability, scalability, and performance of Mastercard applications by implementing observability, automation, CI/CD, and cloud infrastructure best practices. Support production readiness, triage incidents, perform root-cause analysis and blameless post-mortems, mentor developers, and drive operational standards, capacity planning, and risk/compliance activities to maximize service availability and customer experience.
Top Skills:
AWSAzureBashBitbucketCi/CdContainerizationDynatraceGCPGoJenkinsLinux/UnixOrchestrationPcfPythonSplunkXlr
Fintech • Machine Learning • Payments • Software • Financial Services
Lead technical, second-line oversight of SRE and cloud engineering practices. Perform deep-dive risk analyses of cloud architectures, resiliency, CI/CD, observability, and Gen AI integrations. Produce data-driven risk findings, mitigation recommendations, and executive-facing reports while partnering with first-line engineers and leadership to ensure robust controls and operational reliability.
Top Skills:
AWSAzureCi/CdCloud-NativeContainerizationDatadogElkGCPGenerative AiKubernetesPagerdutyPrometheusSplunk
Financial Services
Designs, implements, monitors, and optimizes cloud-based application infrastructure and reliability. Uses IaC/NaC, observability, SLOs, CI/CD pipelines, and enterprise-authorized AI to prevent and resolve incidents, improve availability and scalability, and mentor peers on SRE best practices.
Top Skills:
Ci/CdCloudContainer OrchestrationContainersContinuous DeliveryContinuous IntegrationEnterprise-Authorized AiInfrastructure As CodeJavaMonitoringNetwork As CodeNetworkingObservabilityPysparkPythonService Level Objectives (Slos)Spring BootTelemetry Collection
Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
Lead and scale multidisciplinary engineering teams to deliver cloud-native, AI-enabled enterprise applications and data platforms. Own architecture, modernization, DevOps/MLOps, observability, and AI adoption (RAG, vector DBs, conversational AI). Partner with executives and cross-functional teams to define roadmaps, ensure reliability, and drive engineering excellence and organizational growth.
Top Skills:
.NetAWSAzureCi/CdCloud-Native Application ArchitecturesData LakehouseDatabricksDevOpsDistributed Data Processing PlatformsEvent-Driven ArchitecturesFlinkGoogle Cloud PlatformJavaKafkaMicroservicesMlopsOraclePower BIPythonRest ApisSparkSQLSQL Server
Digital Media • Gaming • Information Technology • Software • Sports • Esports • Big Data Analytics
Lead long-term strategy and architecture for cloud and on‑prem platform infrastructure, driving Kubernetes and multi‑cloud reliability, IaC/GitOps automation, observability, SLO/SLI/error‑budget practices, incident leadership, AI‑augmented tooling adoption, and mentorship of senior engineers to improve platform resilience and developer experience.
Top Skills:
Amazon Elastic Kubernetes Service (Eks)AutoscalingAWSCapacity PlanningCi/CdGitopsGoGoogle Cloud PlatformGoogle Kubernetes Engine (Gke)Identity And Access ManagementInfrastructure As CodeKubernetesLinuxNetworkingObservabilityOperatorsPulumiPythonRke2StorageTerraform
Aerospace • Hardware • Information Technology • Security • Software • Cybersecurity • Defense
Design, deploy, and operate large-scale, distributed IaaS/PaaS/SaaS systems with automation and monitoring. Provide primary operational support, instrument production KPIs, resolve performance issues, and collaborate across DevOps, Security, and IT teams to scale systems through automation.
Top Skills:
AnsibleAzure StackCephHelm ChartsIaasJdfsJuniperKubernetesNfsOpenstackPaasS3SaaSSecurity+TerraformVMware
Aerospace • Hardware • Information Technology • Security • Software • Cybersecurity • Defense
Lead and operate reliability for large-scale distributed systems: deploy and monitor IaaS/PaaS/SaaS, automate scaling and deployments, instrument production for KPIs, troubleshoot performance, provide primary operational support, and collaborate with DevOps, Security, and IT Operations to ensure continuous service delivery.
Top Skills:
AnsibleAzure StackCephHelm ChartsIaasJdfsJuniperKubernetesNfsOpenstackPaasS3SaaSTerraformVMware
Aerospace • Hardware • Information Technology • Security • Software • Cybersecurity • Defense
Design, deploy, automate, and operate large-scale distributed IaaS/PaaS/SaaS systems. Implement monitoring, performance tuning, and automation (Ansible/Helm). Provide operational support for networking and storage (Juniper, NFS/Ceph/S3), collaborate with cross-functional teams, and ensure service reliability and continuous improvement.
Top Skills:
AnsibleAzure StackCephHelm ChartsIaasJdfsJuniperNfsOpenstackPaasS3SaaSSecurity+VMware
Aerospace • Hardware • Information Technology • Security • Software • Cybersecurity • Defense
Build, deploy, automate, monitor, and operate large-scale distributed IaaS/PaaS/SaaS systems. Provide primary operational support, instrument performance metrics, scale via automation, troubleshoot complex issues, and collaborate across DevOps, Security, and IT teams.
Top Skills:
AnsibleAzure StackCephHelmIaasJdfsJuniperKubernetesNfsOpenstackPaasS3SaaSSecurity+TerraformVMware
Aerospace • Hardware • Information Technology • Security • Software • Cybersecurity • Defense
Design, deploy, automate, monitor, and operate large-scale distributed IaaS/PaaS/SaaS systems. Provide primary operational support, instrument production KPIs, resolve complex service issues, scale systems via automation, and collaborate with DevOps, Security, and IT teams to improve reliability and performance.
Top Skills:
AnsibleAzure StackCephHelm ChartsIaasJdfsJuniperKubernetesNfsOpenstackPaasS3SaaSTerraformVMware
Let Your Resume Do The Work
Upload your resume to be matched with jobs you're a great fit for.
Success! We'll use this to further personalize your experience.
Top Companies Hiring Site Reliability Engineers
See AllPopular Job Searches
All Software Engineer Jobs
.NET Developer Jobs
Aerospace Thermal Engineering Jobs
AI Engineer Jobs
Android Developer Jobs
Automation Engineer Jobs
Backend Developer Jobs
Blockchain Developer Jobs
C# Jobs
C++ Jobs
Cloud Architect Jobs
Cloud Engineer Jobs
Design Engineer Jobs
DevOps Engineer Jobs
Director Of Engineering Jobs
Electrical Engineering Jobs
Embedded Software Engineer Jobs
Engineering Jobs
Engineering Manager Jobs
Environmental Engineering Jobs
Field Engineer Jobs
Front End Developer Jobs
Full Stack Developer Jobs
Game Developer Jobs
Golang Jobs
Hardware Engineer Jobs
Industrial Engineering Jobs
iOS Developer Jobs
Java Developer Jobs
Javascript Developer Jobs
Linux Jobs
Manufacturing Engineer Jobs
Mechanical Engineering Jobs
Network Engineer Jobs
PHP Developer Jobs
Process Engineer Jobs
Project Engineer Jobs
Prompt Engineering Jobs
Python Jobs
QA Jobs
Robotics Engineer Jobs
Ruby on Rails Jobs
Salesforce Administrator Jobs
Salesforce Developer Jobs
Scala Jobs
Sharepoint Developer Jobs
Site Reliability Engineer Jobs
Software Engineering Manager Jobs
Solutions Architect Jobs
SQL Developer Jobs
Structural Engineer Jobs
System Engineer Jobs
Test Engineer Jobs
Web Developer Jobs
All Filters
Total selected ()
No Results
No Results


























