Maximum of 25 job preferences reached.
Top Site Reliability Engineer Jobs
3D Printing • Artificial Intelligence • Software • Design
Lead design and operation of scalable, multi-tenant spatial streaming platforms. Build Terraform-based cloud infrastructure, optimize CDN/content delivery, implement observability (SLI/SLO), run incident response/on-call, conduct post-mortems, enforce compliance and security practices, and mentor DevOps engineers to improve reliability and production readiness.
Top Skills:
Aws FargateCdnCoreweaveGrafanaKubernetesPrometheusTerraform
Software
As an AI Support Engineer, you'll manage support requests, resolve user issues, optimize ML models, and contribute to product development.
Top Skills:
Tensorrt
Artificial Intelligence • Information Technology • Software
Build and operate the core infrastructure for Arena's online evaluation systems: design low-latency, high-reliability APIs and gateways, implement enterprise-grade features (rate limiting, auth, metering, audit logging), instrument observability (tracing, latency, usage tracking), integrate with LLM providers and the evaluation platform, and collaborate with research and product teams to scale and harden systems for bursty, unpredictable traffic.
Top Skills:
Anthropic ApiAWSDistributed TracingGCPGoGoogle Llm ApisKubernetesOpenai ApiPostgresRedisRustTerraform
Artificial Intelligence • Cloud • Information Technology • Software
Build and operate production-grade AI infrastructure using Kubernetes, ensuring high availability, reliability, and performance. Develop custom operators and implement automation for efficient operations and monitoring.
Top Skills:
AnsibleBashElk StackEnterprise Storage SystemsGrafanaHigh-Performance NetworkingKubernetesLinuxNvidia Gpu TechnologiesPrometheusPythonTerraform
Reposted 23 Days AgoSaved
Fintech • Analytics
As a Site Reliability Engineer, you will ensure the reliability and performance of a FX trading platform, develop automation, improve system health, and manage SLOs while collaborating with development teams.
Top Skills:
AWSAzureBashC#JavaKubernetesPythonSQL
Gaming • Software
The Site Reliability Engineer will manage infrastructure stability and scalability, lead cloud migrations, and optimize performance across systems while mentoring team members.
Top Skills:
AnsibleAWSAzureBashChefCloudFormationDatadogDockerElk StackGCPGoGrafanaKubernetesPrometheusPuppetPythonTerraformUnix/Linux
Cloud
Design, build, and maintain cloud platform services for sensitive federal missions. Operate and monitor air-gapped, secure environments; build CI/CD pipelines without internet; maintain SLOs/SLIs; own runbooks and incident response; support POA&M and Authority to Operate activities; and deliver internal platform enablement while ensuring strict compliance and security.
Top Skills:
Aws VpcsBgpCi/CdCloudwatchContainersEcs FargateEksGrafanaIpsecLinuxPythonSplunkTerraformTgwsVpc Endpoints
Healthtech • Insurance
Owner of enterprise observability and SRE practices: define SLOs/SLA measurement, drive MTTR reduction, lead incident response, maintain service dependency maps and reliability dashboards, and leverage AI/AIOps to automate triage, root cause analysis, and self-healing remediation across vendor and internal platforms.
Top Skills:
Ai/AiopsBashChaos EngineeringCi/CdCmdbDashboardingData ModelingDistributed TracingInfrastructure-As-CodeItsm/Ticketing SystemsLog AggregationMonitoring PlatformsObservability PlatformsPowershellPythonSIEMTelemetry
Fintech
Lead adoption of SRE practices to improve reliability, observability, automation, and incident response. Implement and maintain observability tooling, instrumentation, CI/CD, and infrastructure-as-code. Partner with developers, participate in on-call rotations, drive postmortems, and reduce operational overhead through automation.
Top Skills:
AnthropicAWSAws EcsAws EksAzureC#DockerGitlab CiGrafanaLinuxOpenaiPrometheusPuppetPythonSplunkTerraformTypescriptWindows
Consumer Web • Information Technology • Mobile • Other • Software • App development
Build, maintain, and automate onX's infrastructure platform and deployment pipeline using IaC. Manage Kubernetes/GKE, Terraform, GCP services, observability, and incident response. Improve performance, availability, cost, and developer path to production while participating in on-call rotations and collaborating on architecture.
Top Skills:
AirflowBigQueryBigtableChecklyClaude CodeCloud RunCloud SqlCockroachdbGCPGkeGoogle Cloud MonitoringGoogle Cloud StorageGoogle ComposerIamKubernetesNoSQLOpentelemetryOpentofuPrometheusPub/SubRootlySQLTerraform
Fintech
The Site Reliability Engineer will manage AWS infrastructures, oversee application deployments, and ensure system reliability and security while collaborating with teams.
Top Skills:
AWSBashCodebuildCodedeployCodepipelineEc2IamPythonRdsRoute 53S3TerraformVpc
Aerospace • Hardware • Software • Defense • Manufacturing
As a Site Reliability Engineer, you'll ensure robotics system reliability, build telemetry integration, and develop tools for diagnostics and automation, collaborating with engineering teams for enhanced production reliability.
Top Skills:
C++DatadogGoKubernetesOpentelemetryPrometheusPythonRos2TelegrafTypescript
New
Cut your apply time in half.
Use ourAI Assistantto automatically fill your job applications.
Use For Free
Cloud • Fintech • HR Tech
Operate and improve service reliability by partnering with development and infrastructure teams. Provide on-call support, troubleshoot distributed systems and cloud infrastructure, implement observability, perform capacity planning and performance optimization, automate operational tasks and infrastructure provisioning, and lead incident management and post-incident analysis.
Software
As a Site Reliability Engineer, you'll enhance system reliability, collaborate on production readiness, define SLIs/SLOs, and improve incident response.
Top Skills:
AWSDatadogGrafanaKubernetesOpentelemetryPrometheusTypescript
Transportation
Design and develop Waabi's observability stack, optimize performance, build automation tooling, and support application requirements while leading projects and mentoring teams.
Top Skills:
AWSC/C++DockerGoGrafanaJavaKubernetesOpentelemetryPythonRust
Events
The Site Reliability Engineer II designs and maintains scalable systems, focusing on automation, monitoring, incident response, and collaboration with developers to enhance operational practices and efficiency.
Top Skills:
BashCloud Service OperationsContainersContinuous DeliveryContinuous IntegrationGoInfrastructure As CodeOrchestration PlatformsPython
Edtech • Kids + Family • Sports
Audit infrastructure, deployment pipelines, monitoring, alerting, incident response, on-call practices, and internal tools. Produce actionable audit reports, implement code and configuration fixes, improve SLOs and reliability practices, advise on scalable architecture, and partner with engineers on implementation and handoff. The role is a fully remote, part-time consulting engagement with potential for full-time conversion.
Top Skills:
Ai Coding ToolsAWSCi/CdDatadogGrafanaPrometheus
Software
Lead end-to-end architecture of AI data center networks, DCI, and global backbone for large GPU clusters. Design underlay and overlay networks, IP/VLAN/VXLAN planning, congestion control (PFC/ECN/INT), and optical transport. Produce HLD/LLD, topology diagrams, standards, and SOPs; engage vendors and drive risk mitigation and network improvements.
Top Skills:
BgpClosDwdmEcmpEcnEvpnGpu ClustersIn-Band Network Telemetry (Int)InfinibandNvidiaOptical Transport NetworksOspfPfcRocev2Sdn ControllersSegment Routing (Sr-Mpls)Spine-LeafSrv6VlanVxlan
Information Technology • Productivity • Software • Manufacturing
Senior individual contributor SRE responsible for reliability, observability, automation, and incident command across multi-account AWS SaaS. Design Terraform modules, CI/CD pipelines, autoscaling/self-healing, OpenTelemetry-based monitoring, and security controls; mentor engineers and integrate validated AI tooling.
Top Skills:
Aws Ec2CloudwatchEcs FargateEksGithub ActionsIamLambdaOpenobserveOpentelemetryPagerdutyPythonRds PostgresqlS3TerraformVpc
Cloud • Information Technology
Maintain and improve production service reliability by building automation, monitoring and alerting, participating in on-call incident response, and partnering with engineering and operations to embed reliability practices and reduce operational toil.
Top Skills:
AnsibleAWSAzureBashCatchpointCi/CdDockerElkGCPGoGrafanaJenkinsKubernetesPrometheusPythonTerraform
Cloud • Information Technology
Act as first responder for customer-impacting incidents, monitor and respond to Zabbix alerts, ensure pod and server farm health, run filesystem checks, support Vault deployments and migrations, troubleshoot pod/Ansible/network issues, participate in on-call rotation, document and automate daily tasks, and coordinate escalations with DC techs and management.
Top Skills:
AnsibleHashicorp VaultKubernetesLinuxNetworkingZabbix
Blockchain • Fintech • Payments • Financial Services • Cryptocurrency • Web3
Design, build, and operate secure, scalable Kubernetes platforms and infrastructure as code (Terraform). Develop backend services and automation (Go, Python, JS/TS), improve CI/CD and observability, run on-call and incident response, define SLIs/SLOs and disaster recovery, embed security and compliance, mentor team members, and partner with product and engineering to raise reliability, performance, and cost-efficiency across hybrid and public-cloud environments.
Top Skills:
Ci/CdGitopsGoJavaScriptKubernetesPythonTerraformTypescript
Information Technology • Cybersecurity
Lead global SRE team to design, implement, and operate scalable, highly available cloud infrastructure. Drive observability, automation (IaC/CI-CD), cost optimization (FinOps), security hygiene, and post-incident improvements while remaining hands-on and exploring AI tooling for reliability.
Top Skills:
Ai ToolingAlloyAWSAzureCi/CdCloudFormationContainer OrchestrationDatadogEdge ComputingFinopsGCPGrafanaGrafana IrnInfrastructure-As-CodeKubernetesLokiPrometheusPulumiServerlessSplunkSpot InstancesTerraform
Cloud • Security • Software • Cybersecurity
Ensure reliability, scalability, and usability of network infrastructure for Akamai Connected Cloud. Define requirements and SLOs, build automation and CI/CD pipelines, collaborate with dev/QA to improve code and stability, troubleshoot complex network issues (on-call), and mentor teammates while driving architectural standards.
Top Skills:
AnsibleArgocdBashBirdChefFrrGithub ActionsGoGobgpJenkinsLinux NetworkingPuppetPythonSalt Stack
Artificial Intelligence • Marketing Tech
Own and scale Cognitiv’s AWS infrastructure while evaluating architecture, networking, security, scalability, and service management. Lead improvements in deployments, monitoring, disaster recovery, and infrastructure-as-code practices. Support co-located Equinix datacenter deployments and hybrid cloud operations alongside a datacenter-focused SRE. Collaborate with engineering and product teams, provide multi-datacenter coverage, and help establish long-term service management best practices.
Top Skills:
AnsibleAWSBashDatadogEc2EquinixKubernetesPrometheusPythonTerraform
Let Your Resume Do The Work
Upload your resume to be matched with jobs you're a great fit for.
Success! We'll use this to further personalize your experience.
Top Companies Hiring Site Reliability Engineers
See AllPopular Job Searches
All Software Engineer Jobs
.NET Developer Jobs
Aerospace Thermal Engineering Jobs
AI Engineer Jobs
Android Developer Jobs
Automation Engineer Jobs
Backend Developer Jobs
Blockchain Developer Jobs
C# Jobs
C++ Jobs
Cloud Architect Jobs
Cloud Engineer Jobs
Design Engineer Jobs
DevOps Engineer Jobs
Director Of Engineering Jobs
Electrical Engineering Jobs
Embedded Software Engineer Jobs
Engineering Jobs
Engineering Manager Jobs
Environmental Engineering Jobs
Field Engineer Jobs
Front End Developer Jobs
Full Stack Developer Jobs
Game Developer Jobs
Golang Jobs
Hardware Engineer Jobs
Industrial Engineering Jobs
iOS Developer Jobs
Java Developer Jobs
Javascript Developer Jobs
Linux Jobs
Manufacturing Engineer Jobs
Mechanical Engineering Jobs
Network Engineer Jobs
PHP Developer Jobs
Process Engineer Jobs
Project Engineer Jobs
Prompt Engineering Jobs
Python Jobs
QA Jobs
Robotics Engineer Jobs
Ruby on Rails Jobs
Salesforce Administrator Jobs
Salesforce Developer Jobs
Scala Jobs
Sharepoint Developer Jobs
Site Reliability Engineer Jobs
Software Engineering Manager Jobs
Solutions Architect Jobs
SQL Developer Jobs
Structural Engineer Jobs
System Engineer Jobs
Test Engineer Jobs
Web Developer Jobs
All Filters
Total selected ()
No Results
No Results




.jpg)





























