Maximum of 25 job preferences reached.
Top Site Reliability Engineer Jobs
Fintech • Financial Services
Lead Site Reliability Engineer responsible for ensuring system reliability, scalability, and performance. Develop automated deployment strategies, maintain monitoring/observability, define SLIs/SLOs, collaborate with cross-functional teams, drive reliability best practices, participate in on-call incident response, and improve delivery through automation and training.
Top Skills:
AWSAzureBashGCPPowershellPython
Fintech • Financial Services
The Director of Splunk Platform Engineering & SRE owns the enterprise Splunk platform, drives incident resolution, optimizes systems, and mentors engineers, focusing on automation and performance.
Top Skills:
AnsibleGitGoJavaKubernetesLinux/UnixMoogPrometheusPythonSplunk
Hardware • Information Technology • Other • Software • Analytics
Architect and operate ML/agent pipelines and infrastructure, deploy and monitor models at scale, pioneer MLOps/Agent Ops best practices, collaborate with domain experts, and test/optimize ML systems for production reliability and cost efficiency.
Top Skills:
Bash ScriptingContainerization (E.G.Docker)Git/GithubLinuxModel VersioningMonitoringNumpyPandasPythonPyTorchScikit-Learn
Healthtech • Telehealth
Lead reliability engineering for mission-critical Azure healthcare workloads. Define and manage SLIs, SLOs, and error budgets; improve observability, scalability, resilience, disaster recovery, and cloud operations. Configure monitoring and alerting, conduct incident response and postmortems, automate remediation, support security and compliance controls, and partner with engineering, product, security, and operations teams. Mentor SRE staff and help establish reliability practices across the organization.
Top Skills:
AnsibleApplication InsightsAWSAzure Chaos StudioAzure Kubernetes Service (Aks)Azure MonitorBicepChaos MeshDatadogDynatraceElastic/ElkGoGrafanaGremlinLog AnalyticsLogicmonitorAzureNew RelicOpentelemetryPowershellPrometheusPythonTerraform
Artificial Intelligence • Computer Vision • Hardware • Robotics • Metaverse
Design, deploy, and optimize on-premises HPC storage infrastructure integrated with cloud environments. Build automation, monitoring, alerting, and self-service tools for large-scale distributed storage systems. Manage enterprise NAS, S3, and parallel filesystems; evaluate technologies; resolve performance bottlenecks; document procedures; and collaborate with engineering teams on infrastructure requirements, application deployment, and resource utilization.
Top Skills:
AWSAzureBashCloudianDistributed StorageDockerElasticsearchEnterprise NasGCPGoGpfsGrafanaHpcInfinibandKibanaKubernetesLsfLustreMesosphere DcosMinioNetappPbsPrometheusPure StoragePythonRdmaRoceS3SlurmSplunkZabbix
Software • Quantum Computing • Metaverse • Infrastructure as a Service (IaaS)
Operate and improve Microsoft Defender services in US Government clouds: lead on-call incident response, automate deployments and monitoring, build scalable reliability tooling, ensure security/compliance, run postmortems, and collaborate across teams to maintain uptime and performance.
Top Skills:
C#Ci/CdCloudGoJavaMicrosoft DefenderPython
Artificial Intelligence • Cloud • Social Impact • Software • Wearables
The Senior Site Reliability Engineer ensures the reliability and performance of cloud-native Kubernetes platforms by building tools, facilitating self-service for engineers, and promoting best practices.
Top Skills:
ArgocdAWSAzureC#Ci/CdGitGoJavaKubernetesPulumiPythonTerraform
Artificial Intelligence • Cloud • Social Impact • Software • Wearables
Build and operate cloud-native Kubernetes platforms and tooling to ensure reliability, performance, and security. Write maintainable code, use CI/CD and IaC, employ observability for debugging, document systems, influence engineering practices, and improve platform reliability and cost efficiency.
Top Skills:
AksApmAWSAzureC#Ci/CdCloud-NativeContainer OrchestrationEksGoInfrastructure As CodeJavaKubernetesLoggingMetricsPulumiPythonTerraform
Artificial Intelligence • Cloud • Social Impact • Software • Wearables
Build and operate cloud-native, production-grade Kubernetes platforms and tooling to improve reliability, operability, and developer experience. Develop IaC and CI/CD automation, use observability to debug distributed systems, document self-service workflows, and influence engineering teams on scalable architectural patterns. Scope and deliver platform projects with focus on security, cost efficiency, and service stability.
Top Skills:
AksApmAWSAzureC#Ci/CdContainer OrchestrationEksGoInfrastructure As CodeJavaKubernetesLoggingMetricsPulumiPythonTerraform
Artificial Intelligence • Cloud • Social Impact • Software • Wearables
The Senior Site Reliability Engineer I will enhance Axon's observability platform, work on distributed tracing, log aggregation, and metrics infrastructure, and develop internal tools while collaborating with engineering teams.
Top Skills:
ArgocdCdkCortexGoGrafanaHelmJaegerJavaLokiOpentelemetryPrometheusPythonTerraform
Software • Quantum Computing • Metaverse • Infrastructure as a Service (IaaS)
Owns reliability and operations for Azure Resource Manager and large-scale distributed cloud services. Responsibilities include on-call incident response, monitoring, observability, automation, performance optimization, service lifecycle ownership, coding, design documentation, and cross-functional delivery. The role requires improving availability, security, efficiency, and customer support across cloud environments while guiding engineers and maintaining service parity with the commercial cloud.
Top Skills:
AzureAzure Resource ManagerJavaScriptPythonService FabricShell Scripting
Software
Define and operationalize reliability for a GPU-accelerated AI platform by creating SLIs, SLOs, error-budget practices, and high-quality alerting. Build APIs exposing reliability state to platform administrators, partner with infrastructure, storage, and networking teams on telemetry, and diagnose observability and performance issues across Kubernetes, bare-metal, and NVIDIA infrastructure in hybrid, edge, and air-gapped environments.
Top Skills:
Cluster ApiGoGrafanaInfinibandK0Rdent AiK0Rdent EnterpriseK0Rdent Observability FrameworkKubernetesNvidia BmcNvlinkOpentelemetryPrometheusPythonRedfishUfmVictorialogsVictoriametrics
New
Track Smarter, Apply Better.
Ditch the spreadsheets. Organize your job search with our freeApplication Tracker.
Use For Free
Insurance
Lead SRE role owning enterprise observability and monitoring strategy across cloud and applications. Architect Dynatrace-to-ServiceNow pipelines, alert correlation, paging integration, CMDB/service mapping, and automation for self-healing. Drive SRE practices (SLIs/SLOs/error budgets), mentor teams, reduce alert noise, improve MTTR, and standardize monitoring and incident workflows.
Top Skills:
ArmBicepCmdbDynatraceGithub ActionsAzureMicrosoft TeamsPagerdutyRunbooksServicenow Itsm/ItomTerraform
Reposted 25 Days AgoSaved
Artificial Intelligence • Machine Learning • Software
Contract SRE to build, operate, and automate colo, on‑prem GPU clusters, and cloud infrastructure. Own provisioning, IaC (Terraform/Ansible), observability (Prometheus/Grafana or DataDog), incident response, runbooks, and customer-facing platform reliability.
Top Skills:
AnsibleBashDatadogGrafanaKubernetesLinuxPrometheusPythonSystemdTerraform
Software • Analytics
Build and lead the SRE function: define SLOs/SLIs/error budgets, own observability and incident response, embed reliability in SDLC, and scale the SRE team for a distributed SaaS platform.
Top Skills:
AnsibleAWSAzureBashDatadogDnsFirewallsGCPGrafanaKubernetesLoad BalancingPrometheusPythonRoutingSplunkSwitchingTcp/IpTerraform
Cloud • Information Technology • Consulting • Cybersecurity
Design, templatize and deploy scalable infrastructure in public clouds (AWS, GCP) using IaC (CloudFormation). Support architects, troubleshoot developer escalations, ensure compliance, and build stable platform services; work within agile teams to create configuration templates and automated deployments.
Top Skills:
AWSAws CloudformationAws EfsEc2GCPPythonRdsRuby
Cloud • Security • Software • Cybersecurity
Lead and mentor SRE teams; partner with engineering, operations and product; apply statistical analysis and networking expertise to diagnose performance and reliability issues; define and implement data feeds; influence technical decisions and investments; build tooling to automate analytical workflows and increase platform reliability.
Top Skills:
CCloudDistributed SystemsDnsEdgeHTTPJavaPerlPythonRSQLTcpTls
Artificial Intelligence • Software • Cybersecurity
Design, build, and maintain scalable, highly available cloud infrastructure for an AI-native cybersecurity platform. Automate deployments and incident response, optimize performance for AI workloads, manage IaC across cloud environments, lead incident management and post-mortems, and collaborate with engineering and security teams to embed reliability.
Top Skills:
AWSAzureDatadogDistributed DatabasesEksElkEvent-Driven SystemsGCPGkeGrafanaKubernetesMicroservicesNetworkingPrometheusPulumiStorageTerraform
Real Estate • Financial Services • PropTech
Support and optimize products migrated to AWS, implement cloud best practices, maintain operational coverage, enhance automation, observability, CI/CD/GitOps, and security. Collaborate with development and platform teams to scale, troubleshoot, and ensure reliable SaaS operations.
Top Skills:
AmisArgocdAWSAws Elastic BeanstalkAws Transfer FamilyAzure DevopsBashCloudwatchCurlDockerEc2EksFluxcdGitGitopsHTTPIstioKubernetesLinkerdLoad BalancerPowershellPythonRdsSQLTerraformWget
Automotive • Logistics
Operate, monitor, and scale on-prem and cloud infrastructure for an autonomous vehicle fleet. Manage data offload systems, CI environments, deployments, ETL/BI pipelines, dashboards, and performance optimization to ensure reliability, security, and operational efficiency.
Top Skills:
AirflowArgoBashCiDockerDocker ComposeETLGrafanaHelmInfluxdbKanikoKubernetesPostgresPythonTimescaledb
Aerospace • Other
Design, operate, and scale on-premise infrastructure for the Starshield satellite constellation. Build automation for Kubernetes cluster deployment and management, operate core infrastructure (databases, monitoring, distributed storage), collaborate with software teams, troubleshoot across the stack, improve service lifecycle, and ensure high availability through monitoring and performance improvements.
Top Skills:
AnsibleBashC++GoKubernetesLinuxOci ContainersPythonTerraform
Aerospace • Other
Design, deploy, and operate on-premises Kubernetes clusters and core infrastructure (databases, monitoring, distributed storage). Build automation, troubleshoot across the Starshield stack, collaborate with software teams to ensure scalable, highly available services, and improve lifecycle processes.
Top Skills:
AnsibleBashBazelC++DatabasesDistributed StorageGoKubernetesLinuxMakefilesMonitoringOci ContainersPythonTcp/IpTerraform
Fintech
The SRE/DevOps Engineer will enhance observability and monitoring tools, improve system reliability, conduct post-incident reviews, and collaborate with developers to optimize workflows and CI/CD processes.
Top Skills:
AWSAzureAzure BicepAzure DevopsChaos MeshCloud FormationDatadogDockerElasticsearchGCPGithub ActionsGitlab Ci/CdGrafanaGremlinJenkinsKafkaKubernetesTerraform
Consulting
Maintain and improve reliability of cloud-based enterprise systems by implementing SRE practices. Participate in design and code reviews, incident management, automation (IaC/CI-CD), monitoring, documentation, and collaboration with cross-functional teams to reduce downtime and improve scalability and security.
Top Skills:
Ansible Automation PlatformArtifactoryAWSAzureBashCi/CdGitlabIacLinuxPackerPowershellPythonTerraformWindows
Information Technology
Lead rearchitecture of a large monolith into cloud-native microservices. Design scalable, highly available systems on AWS and Azure using Kubernetes (AKS/EKS), Terraform, and robust cloud networking. Implement monitoring, logging, and DevOps practices, ensure container and cluster security, evaluate technologies, and mentor engineering teams.
Top Skills:
AksAWSAzureCloud NetworkingEksKubernetesLoad BalancingTerraform
Let Your Resume Do The Work
Upload your resume to be matched with jobs you're a great fit for.
Success! We'll use this to further personalize your experience.
Top Companies Hiring Site Reliability Engineers
See AllPopular Job Searches
All Software Engineer Jobs
.NET Developer Jobs
Aerospace Thermal Engineering Jobs
AI Engineer Jobs
Android Developer Jobs
Automation Engineer Jobs
Backend Developer Jobs
Blockchain Developer Jobs
C# Jobs
C++ Jobs
Cloud Architect Jobs
Cloud Engineer Jobs
Design Engineer Jobs
DevOps Engineer Jobs
Director Of Engineering Jobs
Electrical Engineering Jobs
Embedded Software Engineer Jobs
Engineering Jobs
Engineering Manager Jobs
Environmental Engineering Jobs
Field Engineer Jobs
Front End Developer Jobs
Full Stack Developer Jobs
Game Developer Jobs
Golang Jobs
Hardware Engineer Jobs
Industrial Engineering Jobs
iOS Developer Jobs
Java Developer Jobs
Javascript Developer Jobs
Linux Jobs
Manufacturing Engineer Jobs
Mechanical Engineering Jobs
Network Engineer Jobs
PHP Developer Jobs
Process Engineer Jobs
Project Engineer Jobs
Prompt Engineering Jobs
Python Jobs
QA Jobs
Robotics Engineer Jobs
Ruby on Rails Jobs
Salesforce Administrator Jobs
Salesforce Developer Jobs
Scala Jobs
Sharepoint Developer Jobs
Site Reliability Engineer Jobs
Software Engineering Manager Jobs
Solutions Architect Jobs
SQL Developer Jobs
Structural Engineer Jobs
System Engineer Jobs
Test Engineer Jobs
Web Developer Jobs
All Filters
Total selected ()
No Results
No Results































