Maximum of 25 job preferences reached.
Top Site Reliability Engineer Jobs
Aerospace • Defense
Own and operate Loft's cloud and hybrid network infrastructure: design architectures, secure connectivity, manage network-as-code (IaC), define SLOs and observability, automate toil, and participate in SRE incident response and platform reliability.
Top Skills:
ArgocdCi/CdDnsDockerFirewallingFluxcdGCPGitGitopsGrafanaHybrid ConnectivityK8SKubernetesPeeringSdnSite-To-Site RoutingSoftware-Defined NetworkingTerraformVpcVpn
Gaming
The role involves ensuring production quality, owning system reliability, and participating in decision-making. Responsibilities include incident response and lifecycle management in cloud gaming technologies.
Top Skills:
BashC++ElasticsearchGoIstioJavaKafkaKong Api GatewayKubernetesKumaLinkerdMongoDBMySQLPostgresPythonRedisRust
Artificial Intelligence • Other • Sales • Software
The role involves designing and advancing infrastructure for the engineering team, ensuring the reliability of Kubernetes clusters, automating operations, and building machine learning infrastructure.
Top Skills:
ArgoAWSAzureCloudFormationFluxGithub ActionsGoGCPKubernetesPostgresPythonTerraform
Information Technology • Professional Services
Provide senior SRE expertise to improve reliability, scalability, performance, and resilience of a cloud-hosted geospatial platform. Design monitoring/observability, automate deployments, support incident response, optimize capacity and performance, and collaborate across DevSecOps, Kubernetes, database, and support teams in a mission-focused DoD environment.
Top Skills:
Alerting ToolsArcgis EnterpriseAw S Cloud OneAWSCi/CdContainerized SystemsEsriInfrastructure-As-CodeKubernetesLinuxLogging ToolsMonitoring ToolsRmfScriptingStig
Information Technology • Insurance • Software
Own and operate production services end-to-end to ensure reliability, scalability, performance, and operational health. Define SLIs/SLOs, perform incident response and root cause analysis, build automation and self-healing, manage production changes, and collaborate with engineering, product, and operations teams to improve system design and observability.
Top Skills:
.NetAWSC#Ci/CdInfrastructure As CodeJavaKubernetesLinuxPythonReactRelational DatabasesWindows
Artificial Intelligence • Software
Own and scale the GPU compute fleet: build metrics, alerting, observability, and repair automation; design GPU qualification/burn-in pipelines; own Redfish/BMC tooling and firmware telemetry; run incidents, eliminate toil, and deliver end-to-end reliability and orchestration for Kubernetes and bare-metal compute at hyperscale.
Top Skills:
Agentic FrameworksBare MetalBmcCadenceClaude CodeCursorFirmwareGoGpuGrafanaIpmiKubernetesLlm ApisMcp ServersPrometheusPythonRedfishTemporal
Cloud • Security • Software • Cybersecurity
Lead reliability for a serverless AI inference platform: own observability and SLO/SLI frameworks, build automation and tooling, manage incidents and on-call, define deployment safety (canaries, rollbacks), influence architecture with product teams, and mentor other SREs.
Top Skills:
AutoscalingCi/CdContainer OrchestrationContainerizationGoGpu WorkloadsInfrastructure-As-CodeKubernetesModel ServingPythonResource Scheduling
Cloud
Design, build, and maintain secure, air-gapped cloud platform services and CI/CD pipelines for Okta Federal. Operate mission-critical infrastructure, monitor SLOs/SLIs, run incident response and POA&M remediation, and support Authority to Operate activities. Advocate SRE/DevOps practices across teams while working autonomously in secure facilities.
Top Skills:
Air-Gapped EnvironmentsAmazon CloudwatchAws Transit GatewayAws VpcBgpCi/CdEcs FargateEksGrafanaIpsecPythonSplunkTerraformVpc Endpoints
Blockchain • Fintech • Cryptocurrency
Lead design, implementation, and operation of reliable hybrid cloud and on-prem infrastructure. Deliver complex initiatives, build immutable infrastructure with IaC, improve observability and incident response, mentor engineers, and ensure reliability, scalability, and compliance in change-controlled environments.
Top Skills:
AnsibleAWSDatadogGithub ActionsHelmKubernetesKustomizeMakefilesPythonTerraformVirtualization
Artificial Intelligence • Cloud • Software • Cybersecurity
Operate and tune AWS environments to meet SLAs, build observability and alerts, automate infrastructure with IaC and CI/CD, define SLIs/SLOs, support security/compliance within a FISMA Moderate boundary, design resilience and DR plans, and own incident response and post-mortems.
Top Skills:
AnsibleAWSAws CloudwatchAws Trusted AdvisorCi/CdCloudFormationDockerGitlab CiJenkinsNew RelicPythonSplunkTerraform
Fintech • Payments • Financial Services
Build, operate, and scale AWS-based infrastructure using IaC (Terraform), manage EKS and serverless environments, create CI/CD pipelines, implement observability (OpenTelemetry/Prometheus/New Relic), support Postgres/RDS (Aurora), lead incident response and define SRE practices (SLIs/SLOs/error budgets).
Top Skills:
AuroraAWSAws RdsAzureCloudFormationEcsEksGithub ActionsGitlabGoGCPJavaKubernetesNew RelicOpentelemetryOpentofuPostgresPrometheusPythonRubyServerlessTerraformTerragrunt
Reposted 21 Days AgoSaved
Fintech • Analytics
As a Senior Site Reliability Engineer, you'll lead incident recovery, enhance production stability, automate processes, and collaborate with development teams to improve operational efficiency.
Top Skills:
AWSAzureBigpandaCloud-Native ApplicationsDatadogDnsDockerGitHTTPKubernetesShell ScriptingTcp/IpUnix
New
Cut your apply time in half.
Use ourAI Assistantto automatically fill your job applications.
Use For Free
Enterprise Web • Information Technology • Software
As a Platform Engineer, you will enhance reliability and performance, design operational processes, and build monitoring systems while collaborating with a talented team.
Top Skills:
AIAssistantsBackendDeveloper ToolsFrontendInfrastructureMcpsMonitoring SystemsSkills
Software
Join a passionate team to enhance reliability and performance of the AI control plane, manage deployments, and respond to production incidents while ensuring service quality for customers.
Top Skills:
Ai Control PlaneDeveloper ToolsInfrastructure
Legal Tech • Software
Lead Site Reliability Engineer responsible for platform availability and reliability of RelativityOne. Drive SRE best practices, build tools, lead projects, coach SREs, work with stakeholders, support incidents, run postmortems, and improve monitoring, automation, and operational efficiency.
Top Skills:
Ci/CdDevOpsJenkinsJIRAKubernetesAzureMonitoring And AlertingNew RelicNoSQLPowershellRelativity ServerRelativityoneSQLTableau
Automotive • Hardware • Logistics
Builds and supports large-scale, distributed, fault-tolerant systems to improve reliability and automation. Administers networks and databases, monitors system health, configures load and data communications, coordinates equipment and vendor orders, and participates in change management to reduce incidents and support cloud transformations.
Top Skills:
CloudDistributed ComputingInternet SecurityMonitoring ToolsOracle ErpUnixVersion Control SystemsWindows 2000Windows 98Windows Nt
Information Technology • Consulting
As a Senior Staff Site Reliability Engineer, you will lead the SRE team, advocate best practices, ensure resilience in cloud architecture, and mentor team members.
Top Skills:
ArgocdCircleCIGoogle Cloud PlatformKubernetesPulumiTerraformTypescript
Healthtech
The Senior Software Engineer will enhance system reliability, manage Kubernetes and AWS environments, oversee incident responses, and implement observability measures.
Top Skills:
AWSCloudwatchElbGithub ActionsKubernetesObservability ToolingTerraformVpc
Fintech • Financial Services
As a Site Reliability Engineer for AI platform, you will ensure system reliability, develop automation, manage infrastructure, and evaluate new technologies.
Top Skills:
AnsibleApache KafkaAWSAzureCloudFormationDatadogDockerEfkElkGoGCPGpu ClustersGrafanaHelmJavaKubernetesPrometheusPythonSnowflakeSparkTerraform
Artificial Intelligence • Hardware • Robotics • Software
Build, operate, and scale production cloud infrastructure and Kubernetes/EKS clusters on AWS. Own Terraform-defined infra, CI/CD, observability, networking, and reliability. Troubleshoot production incidents, participate in on-call rotations, automate operational work with Python/Go, and expand multi-region deployments.
Top Skills:
Argo CdAWSDatadogEksGithub ActionsGitlab Ci/CdGitopsGoHelmIamJenkinsKubernetesLinuxPostgresPythonSpinnakerTerraformVpc
Software
Front-line SRE covering 8AM–8PM PST monitoring GPU clusters, networking, storage, and sensors. Execute runbooks, perform hardware triage and physical DC tasks (rack, cable, swap), collect diagnostics for escalation, manage tickets (ServiceNow/Jira), update runbooks, and tag incidents to train the AIOps platform.
Top Skills:
AiopsBmcDcgmGpuGrafanaJira Service ManagementLinuxNagiosPrometheusServicenow
Artificial Intelligence • Marketing Tech • Mobile • Software
Lead design and implementation of scalable, reliable platform systems; define SLIs/SLOs and observability; drive cross-team strategic initiatives; mentor engineers; own production standards, incident management, and cost/operational optimization to improve platform reliability and scalability.
Top Skills:
GoJavaPythonTypescript
Financial Services
Design, build, and operate reliable cloud infrastructure and networking (multi-account AWS, VPC, IAM). Implement IaC, CI/CD pipelines, observability (logging/metrics/alerting), automation, and reliability guardrails. Provide production support and incident response, perform root cause analysis, and collaborate with application teams to co-own system design and continuous improvement, using AI-assisted tools where appropriate.
Top Skills:
.NetAi-Assisted Tools (Claude CodeAWSAws OrganizationsBashCi/CdCloudFormationElastic StackGitGithub CopilotIamInfrastructure As CodeJavaJenkinsNode.jsObservabilityOpensearchPowershellPythonTerraformVpcWindsurf)
Artificial Intelligence • Fintech • Payments • Business Intelligence • Financial Services • Generative AI
Lead design and delivery of scalable cloud infrastructure for the Spend product. Embed with development teams to drive reliability, performance, observability, incident response, and automation. Own SLOs, runbooks, DevOps metrics, and collaborate with central DevOps and security teams to ensure compliance and resilience. Lead infrastructure projects including new service launches, data centre migrations, and modernising data pipelines.
Top Skills:
Analytics PipelinesAWSData StreamingDevOpsGCPIncident ResponseKubernetesObservabilitySlosSre
Information Technology
Design, build, and operate a reliable, scalable developer platform and cloud infrastructure (GCP/GKE). Lead SRE practices: SLO/SLI, observability, incident response, on-call, automation, security-by-default, Terraform IaC, CI/CD, capacity planning, mentoring, and platform enablement across teams.
Top Skills:
Ci/CdConfluenceDatadogDockerGCPGitGithub ActionsGkeGoGrafanaGsm (Gcp Secrets Manager)HclHelmHoneycombJIRAKubernetesNew RelicNode.jsOpentelemetryPrometheusPythonSslTerraform
Let Your Resume Do The Work
Upload your resume to be matched with jobs you're a great fit for.
Success! We'll use this to further personalize your experience.
Top Companies Hiring Site Reliability Engineers
See AllPopular Job Searches
All Software Engineer Jobs
.NET Developer Jobs
Aerospace Thermal Engineering Jobs
AI Engineer Jobs
Android Developer Jobs
Automation Engineer Jobs
Backend Developer Jobs
Blockchain Developer Jobs
C# Jobs
C++ Jobs
Cloud Architect Jobs
Cloud Engineer Jobs
Design Engineer Jobs
DevOps Engineer Jobs
Director Of Engineering Jobs
Electrical Engineering Jobs
Embedded Software Engineer Jobs
Engineering Jobs
Engineering Manager Jobs
Environmental Engineering Jobs
Field Engineer Jobs
Front End Developer Jobs
Full Stack Developer Jobs
Game Developer Jobs
Golang Jobs
Hardware Engineer Jobs
Industrial Engineering Jobs
iOS Developer Jobs
Java Developer Jobs
Javascript Developer Jobs
Linux Jobs
Manufacturing Engineer Jobs
Mechanical Engineering Jobs
Network Engineer Jobs
PHP Developer Jobs
Process Engineer Jobs
Project Engineer Jobs
Prompt Engineering Jobs
Python Jobs
QA Jobs
Robotics Engineer Jobs
Ruby on Rails Jobs
Salesforce Administrator Jobs
Salesforce Developer Jobs
Scala Jobs
Sharepoint Developer Jobs
Site Reliability Engineer Jobs
Software Engineering Manager Jobs
Solutions Architect Jobs
SQL Developer Jobs
Structural Engineer Jobs
System Engineer Jobs
Test Engineer Jobs
Web Developer Jobs
All Filters
Total selected ()
No Results
No Results










.jpg)


.png)



















