Maximum of 25 job preferences reached.
Top Site Reliability Engineer Jobs
Artificial Intelligence • Natural Language Processing • Generative AI
Lead deployment and runtime operations for ML safety systems: configure and verify safeguards across platforms, run canary rollouts and validations, automate launch runbooks into pipelines, maintain a provenance-backed safeguards registry, and participate in on-call and incident response for safety-critical model releases.
Top Skills:
AWSAws BedrockCanary DeploymentsCi/Cd PipelinesClaudeConfig ManagementGCPGcp VertexLlm InferencePythonRustTransformer-Based Models
Blockchain • Fintech • Payments • Financial Services • Cryptocurrency • Web3
Design, build, and operate secure, scalable Kubernetes platforms and infrastructure as code (Terraform). Develop backend services and automation (Go, Python, JS/TS), improve CI/CD and observability, run on-call and incident response, define SLIs/SLOs and disaster recovery, embed security and compliance, mentor team members, and partner with product and engineering to raise reliability, performance, and cost-efficiency across hybrid and public-cloud environments.
Top Skills:
Ci/CdGitopsGoJavaScriptKubernetesPythonTerraformTypescript
Information Technology • Cybersecurity
Lead global SRE team to design, implement, and operate scalable, highly available cloud infrastructure. Drive observability, automation (IaC/CI-CD), cost optimization (FinOps), security hygiene, and post-incident improvements while remaining hands-on and exploring AI tooling for reliability.
Top Skills:
Ai ToolingAlloyAWSAzureCi/CdCloudFormationContainer OrchestrationDatadogEdge ComputingFinopsGCPGrafanaGrafana IrnInfrastructure-As-CodeKubernetesLokiPrometheusPulumiServerlessSplunkSpot InstancesTerraform
Cloud • Security • Software • Cybersecurity
Ensure reliability, scalability, and usability of network infrastructure for Akamai Connected Cloud. Define requirements and SLOs, build automation and CI/CD pipelines, collaborate with dev/QA to improve code and stability, troubleshoot complex network issues (on-call), and mentor teammates while driving architectural standards.
Top Skills:
AnsibleArgocdBashBirdChefFrrGithub ActionsGoGobgpJenkinsLinux NetworkingPuppetPythonSalt Stack
Artificial Intelligence • Marketing Tech
Own and scale Cognitiv’s AWS infrastructure while evaluating architecture, networking, security, scalability, and service management. Lead improvements in deployments, monitoring, disaster recovery, and infrastructure-as-code practices. Support co-located Equinix datacenter deployments and hybrid cloud operations alongside a datacenter-focused SRE. Collaborate with engineering and product teams, provide multi-datacenter coverage, and help establish long-term service management best practices.
Top Skills:
AnsibleAWSBashDatadogEc2EquinixKubernetesPrometheusPythonTerraform
Computer Vision • Hardware • Mobile • Software • Semiconductor
Ensure 24x7 reliability and scalability of AMHS control software in a semiconductor fab. Build and maintain back-end data pipelines, observability (Prometheus/Grafana), and automation tooling. Troubleshoot software, IPC/RPC network issues, and optimize middleware (RabbitMQ, Redis). Modernize infrastructure using Python, Java, Docker, and Kubernetes to reduce toil and prevent production downtime.
Top Skills:
DockerETLGrafanaIpcJavaKubernetesLinux/UnixMesNosql (Nosql Databases)PrometheusPythonRabbitMQRedisRpcSecs/GemSql (Relational Databases)
Artificial Intelligence • Cloud • Fintech • Machine Learning • Mobile • Software
Lead design, development, deployment, and scaling of cloud infrastructure and SRE tooling. Build automation, CI/CD, observability, capacity planning, and reliability improvements; collaborate with product teams to define non-functional requirements and resolve production issues.
Top Skills:
.NetApi GatewayAWSAzureC#Data LakehouseDatabricks DeltaDatadogElasticsearchElkEvent HubsFunctions/ServerlessGitGrafanaJavaJenkinsKafkaKibanaKubernetesLogstashPowershellSnowflakeSqsTeamcityVisual Basic
Software
Lead SRE to define strategy and roadmap for reliability, scalability, observability, and automation across cloud and hybrid environments. Design and operate containerized production workloads, infrastructure-as-code, monitoring/alerting, incident management, and compliance for regulated domains. Mentor SREs, partner with security and product teams, manage cloud costs and capacity, and build a developer platform to improve delivery and production stability.
Top Skills:
AWSAws MarketplaceAzureAzure MarketplaceGCPGoogle Cloud MarketplaceGrafanaKubernetesPrometheusTerraform
Healthtech • Software
Own reliability, security, and continuity of Medgen production environments. Manage Windows and Linux servers, HA configurations, backups, DR, deployments, and CI/CD automation. Coordinate vulnerability assessments, incident response, cloud/hybrid evaluations, and infrastructure modernization.
Top Skills:
.NetAWSAzureBackupsCi/CdDisaster RecoveryDnsFirewallsGCPIisJavaScriptLinuxMicrosoft TfsNetwork SegmentationSQLTcp/IpVpnsWindows Server
Healthtech • Software
Own reliability, security, and continuity of Labgen production: manage Windows/Linux servers, HA, networking, backups/DR, production deployments, CI/CD, monitoring, cloud evaluation, incident response, vendor relations, and compliance.
Top Skills:
.NetApacheAWSAzureBashCentosCertificates/PkiCvsDatadogDnsElkFirewallsGCPGitGrafanaJavaScriptJenkinsLinux (RhelPrometheusPythonSentineloneSql/Relational DatabasesTcp/IpUbuntu)VpnWindows Server
Fintech • Financial Services
Operates and improves high-traffic, business-critical cloud and network systems. Designs scalable network solutions, manages performance and troubleshooting, automates deployment and monitoring, coordinates router and switch installations, and tests redundancy, resilience, and failover. Partners with development teams to improve operability, supports project planning and vendor comparisons, provides training, and participates in on-call coverage.
Top Skills:
Cloud ComputingIpRoutersSwitchesVoip
Healthtech • Software
Manage and optimize a multi-account AWS environment supporting production healthcare applications and analytics platforms. Responsibilities include AWS infrastructure administration, CI/CD and infrastructure automation, observability, incident response, on-call support, disaster recovery validation, HIPAA/HiTrust compliance, IAM and security controls, and support for containerized Java and Python applications. The role also contributes to Kubernetes and EKS modernization initiatives and partners with developers to resolve complex production issues.
Top Skills:
Amazon EksApi GatewayArgocdAuroraAWSAws CloudformationAws CodepipelineAws Security HubBashCloudfrontCloudwatchCortex CloudDatadogDockerEc2EcsFargateGitGuarddutyHelmIamJavaJenkinsKubernetesLambdaLinuxPrismaPythonRdsS3Spring BootUbuntuZabbix
New
Track Smarter, Apply Better.
Ditch the spreadsheets. Organize your job search with our freeApplication Tracker.
Use For Free
Big Data • Fintech • Mobile • Payments • Financial Services • Data Privacy
Designs and matures site reliability engineering capabilities across cloud and on-premises platforms. Defines SLIs, SLOs, error budgets, observability standards, and production readiness requirements. Automates toil reduction, remediation, infrastructure provisioning, CI/CD, and operational workflows. Leads incident triage, root-cause investigations, capacity planning, resiliency initiatives, AIOps adoption, and platform health monitoring. Partners with development, infrastructure, security, and governance teams to improve reliability, scalability, compliance, and operational excellence.
Top Skills:
AiopsAnsibleCi/CdCloud PlatformsContainerizationEvent-Driven RemediationGenaiIamInfrastructure As Code (Iac)Infrastructure AutomationLogsMetricsObservabilityOrchestration PlatformsPolicy-As-CodeSynthetic MonitoringTelemetryTerraformTraces
Financial Services
Define and implement the company’s SRE practice, including customer-focused SLOs, error budgets, burn-rate alerting, reliability patterns, and observability standards. Embed circuit breakers, retries, bulkheads, and load shedding into Ruby, Java, and Elixir services. Extend Prometheus, Honeycomb, and OpenTelemetry monitoring, support scaling on HashiCorp Nomad, mentor engineers, and establish blameless incident review practices.
Top Skills:
ConsulElixirFlow AnalysisGrafanaHashicorp NomadHoneycombJavaLinuxMulticastOpentelemetryPacket CapturePrometheusPythonRubyTcp/IpUdpVault
Legal Tech • Software
Lead observability and incident management efforts: define SLIs/SLOs, build monitoring/alerting, dashboards, logging, and tracing. Drive incident response, postmortems, and reliability improvements to reduce MTTD/MTTR. Integrate observability into CI/CD, maintain AWS and Kubernetes infrastructure, automate operations, and mentor engineers on SRE best practices.
Top Skills:
AWSBashCi/CdDatadogDistributed TracingDynatraceGrafanaKubernetesNew RelicOpentelemetryPowershellPrometheusPython
Legal Tech • Software
Senior technical leader for SRE driving observability, platform infrastructure, SLIs/SLOs, incident response, automation, and self-service platform capabilities. Shapes reliability strategy, mentors engineers, and ensures production-scale operational excellence.
Top Skills:
AiopsBashDatadogGoInfrastructure As CodeKubernetesNew RelicObservabilityPython
Information Technology
Maintain and optimize mission-critical bare-metal infrastructure, build automation for web2/web3 deployments, perform R&D and monitoring for blockchain validator/RPC/operator nodes, ensure SLIs/SLOs, and participate in on-call rotation.
Top Skills:
AnsibleBare-MetalDockerGCPGithub ActionsGoHaproxyKubernetesOracle CloudPostgresPythonTerraform
Artificial Intelligence • Big Data • Information Technology • Security • Software
Design, build, and maintain cloud infrastructure and CI/CD for a high-availability telecommunications product. Define SLOs/SLIs, manage incident response and on-call rotations, implement observability, perform performance and capacity planning, run blameless postmortems, and collaborate with security teams to ensure compliance and access control.
Top Skills:
AnsibleAWSDatadogDockerGCPGitlabHelmJavaJenkinsKubernetesNoSQLTerraform
Information Technology • Logistics • Transportation • Analytics • Business Intelligence • 3PL: Third Party Logistics • Industrial
As a Site Reliability Engineer, you'll enhance reliability for Phenix WMS and automation systems, focusing on incident reduction and system health through observability and automation. Responsibilities include defining SLIs and SLOs, participating in incident response, and testing disaster recovery plans.
Top Skills:
AnsibleAzureBashCi/CdKubernetesPowershellPythonTerraform
Aerospace • Other
The Site Reliability Engineer, GNC at SpaceX oversees mission-critical GNC products, operates servers, maintains HPC clusters, and enhances services and infrastructure to support space operations.
Top Skills:
AnsibleBazelDockerGradleKubernetesLinuxMakeNpmPipPuppetPythonTerraformVagrant
Aerospace • Other
Manage and design server, HPC, storage, and networking infrastructure (including InfiniBand) to support propulsion engineering workflows. Integrate and optimize engineering applications (ANSYS, StarCCM+), automate deployments, troubleshoot performance bottlenecks, and coordinate with IT, facilities, and engineering teams to scale compute resources for rocket engine development.
Top Skills:
AnsibleAnsysBashDockerEnterprise NetworkingHpcInfinibandKubernetesLinuxPuppetPythonStarccm+VirtualizationWindows Server
Artificial Intelligence • Cloud • Software • Infrastructure as a Service (IaaS)
The Site Reliability Engineer will ensure the reliability and performance of AI infrastructure, build core systems, handle incident response, and develop automation tools.
Top Skills:
AWSDatadogElkGCPGithub ActionsGitlab CiGoGrafanaJenkinsKubernetesLinuxPrometheusPulumiPythonRustTerraform
Database
Embed with service teams to define SLIs/SLOs and error budgets, run Operational Readiness Reviews, improve incident-to-improvement pipelines, advise on resilience and architecture, reduce operational toil through automation, and shape org-wide on-call practices and operational maturity.
Top Skills:
AWSCdkGrafanaKubernetesOpentelemetryPostgresPulumiTerraformVictoriametrics
Reposted 24 Days AgoSaved
Fintech • Financial Services
Provide production support and SRE for business-critical digital asset platforms (wallets, blockchain-connected systems, transaction processing). Lead incident management, troubleshoot high-availability distributed systems, perform system-level Linux debugging, use monitoring/observability tools, automate tasks via scripting, investigate data with SQL, and communicate with technical and business stakeholders.
Top Skills:
BlockchainDistributed SystemsElkGrafanaJavaKubernetesLinuxNode OperationsOpenshiftPrometheusPythonSplunkSQLTransaction ProcessingWallets
Fintech • Financial Services
The Site Reliability Engineer will support cloud infrastructure, automate deployments, and ensure operational efficiency and governance across public cloud platforms.
Top Skills:
AnsibleAWSAzureAzure CliAzure FunctionsAzure Kubernetes ServiceCosmodbGCPGitJenkinsKubernetesLinuxPowershellTerraformWindows
Let Your Resume Do The Work
Upload your resume to be matched with jobs you're a great fit for.
Success! We'll use this to further personalize your experience.
Top Companies Hiring Site Reliability Engineers
See AllPopular Job Searches
All Software Engineer Jobs
.NET Developer Jobs
Aerospace Thermal Engineering Jobs
AI Engineer Jobs
Android Developer Jobs
Automation Engineer Jobs
Backend Developer Jobs
Blockchain Developer Jobs
C# Jobs
C++ Jobs
Cloud Architect Jobs
Cloud Engineer Jobs
Design Engineer Jobs
DevOps Engineer Jobs
Director Of Engineering Jobs
Electrical Engineering Jobs
Embedded Software Engineer Jobs
Engineering Jobs
Engineering Manager Jobs
Environmental Engineering Jobs
Field Engineer Jobs
Front End Developer Jobs
Full Stack Developer Jobs
Game Developer Jobs
Golang Jobs
Hardware Engineer Jobs
Industrial Engineering Jobs
iOS Developer Jobs
Java Developer Jobs
Javascript Developer Jobs
Linux Jobs
Manufacturing Engineer Jobs
Mechanical Engineering Jobs
Network Engineer Jobs
PHP Developer Jobs
Process Engineer Jobs
Project Engineer Jobs
Prompt Engineering Jobs
Python Jobs
QA Jobs
Robotics Engineer Jobs
Ruby on Rails Jobs
Salesforce Administrator Jobs
Salesforce Developer Jobs
Scala Jobs
Sharepoint Developer Jobs
Site Reliability Engineer Jobs
Software Engineering Manager Jobs
Solutions Architect Jobs
SQL Developer Jobs
Structural Engineer Jobs
System Engineer Jobs
Test Engineer Jobs
Web Developer Jobs
All Filters
Total selected ()
No Results
No Results































