Top Site Reliability Engineer Jobs

Reposted 8 Days AgoSaved
Hybrid
Austin, TX, USA
Senior level
Senior level
Logistics • Mobile • Productivity • Software • Transportation
The Senior Site Reliability Engineer will manage the reliability of Zello's data tier, contribute to monitoring and incident response while improving cloud infrastructure and database performance.
Top Skills: BashDockerElasticsearchGoKubernetesLokiMongoDBMySQLPrometheusPythonRedisScylladbTempo
Reposted 8 Days AgoSaved
Hybrid
O'Fallon, MO, USA
96K-163K Annually
Senior level
96K-163K Annually
Senior level
Blockchain • Fintech • Payments • Consulting • Cryptocurrency • Cybersecurity • Quantum Computing
Lead performance and reliability engineering for ABU services: design and run automated performance tests, analyze bottlenecks, tune systems (including JVM/GC), support capacity planning and chaos testing, automate repeatable tasks, review incidents, and mentor engineers to improve scalability, resiliency, and engineering velocity.
Top Skills: Ai-Assisted ToolsBitbucketBlazemeterCloud-Native StacksGarbage Collection TuningGatlingJavaJenkinsJvmLarge Language ModelsLoadrunnerMavenPythonScala
3 Days AgoSaved
In-Office
2 Locations
147K-234K Annually
Senior level
147K-234K Annually
Senior level
Fintech • Payments • Financial Services
Leads the design and operation of highly available cloud systems, primarily on AWS, using Terraform, GitLab CI/CD, containers, and observability tools. Responsibilities include reliability engineering, incident response, disaster recovery, automation, monitoring, security integration, vulnerability management, and cost optimization. The role builds internal tools, guides architecture, mentors SRE engineers, documents operational processes, and partners with software engineering teams to improve production reliability and compliance.
Top Skills: Agentic ApplicationsAmazon Api GatewayAmazon AuroraAmazon CloudfrontAmazon CloudwatchAmazon DynamodbAmazon EbsAmazon Ec2Amazon EcsAmazon EfsAmazon RdsAmazon Route 53Amazon S3Amazon VpcAWSAws FargateAws LambdaAws X-RayDastDatadogDockerGitlabGrafanaIamJavaKubernetesLlmsNew RelicNode.jsOwasp Top 10PythonSastSplunkTerraform
4 Days AgoSaved
In-Office
Boston, MA, USA
131K-200K Annually
Expert/Leader
131K-200K Annually
Expert/Leader
Information Technology • Internet of Things • Software • Virtual Reality
Lead reliability, scalability, and operational excellence for large-scale Java distributed systems. Drive architecture improvements, incident management, observability, SLO adoption, automation, infrastructure strategy, and performance optimization. Influence cross-team technical decisions, mentor engineers, establish reliability standards, and guide incident response and durable remediation. The role requires hybrid work in Boston and participation in an on-call rotation.
Top Skills: AWSCi/CdInfrastructure As CodeJavaMongoDBRabbitMQZookeeper
Reposted 4 Days AgoSaved
In-Office
2 Locations
93K-204K Annually
Senior level
93K-204K Annually
Senior level
Fitness • Healthtech • Retail • Pharmaceutical
Ensure reliability, scalability, and performance of distributed retail and pharmacy systems. Implement observability, monitoring, SLOs, incident response, and reliability improvements. Lead cloud-native microservices and Kubernetes/OpenShift deployments, CI/CD automation, performance testing, and collaborate with engineering and operations. Participate in on-call rotation.
Top Skills: Ai/AiopsApigeeAWSBitbucketDatadogDatapowerDockerDynatraceGitGCPGrafanaJavaJenkinsKubernetesAzureOpenshiftPrometheusPythonRancherSplunkVordelWeb Apis
Reposted 4 Days AgoSaved
In-Office
12 Locations
149K-282K Annually
Senior level
149K-282K Annually
Senior level
Cloud • Information Technology • Internet of Things • Professional Services • Software
Designs, deploys, and operates reliable, scalable cloud and network infrastructure; automates platforms; monitors SLOs/SLAs; leads on-call incident response, postmortems, DR planning, and security controls to improve service reliability.
Top Skills: AutomationCloud InfrastructureDisaster RecoveryFedramp HighIl-5Monitoring And AlertingNetworkingOn-Call Incident ManagementSecurity ControlsSre PracticesStorage Systems
4 Days AgoSaved
In-Office
Chapel Hill, NC, USA
100K-150K Annually
Senior level
100K-150K Annually
Senior level
Artificial Intelligence • Information Technology • Software • Consulting
Operate and improve the reliability, availability, and performance of large-scale distributed systems. Build automation and tooling, manage Linux and Kubernetes environments, design CI/CD pipelines, implement observability, troubleshoot production systems, and reduce operational toil. Preferred work includes defining SLOs and error budgets, chaos engineering, cloud operations, capacity planning, load testing, and service mesh management.
Top Skills: AWSAzureChaos MonkeyCi/CdConsulDockerEfkElkGCPGoGrafanaGremlinIstioJavaKubernetesLinkerdLinuxLitmusOpentelemetryPrometheusPython
4 Days AgoSaved
In-Office or Remote
Santa Clara, CA, USA
200K-322K Annually
Senior level
200K-322K Annually
Senior level
Artificial Intelligence • Computer Vision • Hardware • Robotics • Metaverse
Lead the transformation, design, deployment, and operation of globally scaled on-premises and cloud compute infrastructure. Build reliable core services including DNS, NTP/PTP, DHCP, and LDAP; develop automation, monitoring, capacity planning, and lifecycle management. Define performance metrics, optimize systems using technologies such as SR-IOV and DPUs, and develop data analysis and visualization tools. Partner with engineering, product, program, and company leadership to deliver scalable IT services.
Top Skills: AnycastBare MetalBgpConfiguration Management ToolsContainersDhcpDnsDpuEbpfGoInfrastructure As CodeLdapLinuxLinux KernelMicroservicesNtpPtpPythonSdnSr-IovTerraformVlanVxlanXdp
4 Days AgoSaved
Remote
United States
102K-219K Annually
Junior
102K-219K Annually
Junior
Software • Quantum Computing • Metaverse • Infrastructure as a Service (IaaS)
Own reliability and operational health for Microsoft Substrate cloud services in regulated environments. Participate in on-call rotations, diagnose and resolve production incidents, build automation to reduce toil, maintain production code, develop monitoring and telemetry supporting SLOs, lead post-incident reviews, and collaborate with software engineering teams to improve service reliability, scalability, security, and operability.
Top Skills: Exchange OnlineMicrosoft 365 CopilotMicrosoft CloudMicrosoft Substrate
4 Days AgoSaved
Remote
United States
90K-100K Annually
Mid level
90K-100K Annually
Mid level
Digital Media
Build, maintain, and operate Ookla’s globally distributed infrastructure platform at massive scale. Responsibilities include managing cloud instances, containers, serverless applications, databases, streaming systems, and big-data tooling; supporting 24/7 production operations and on-call rotations; implementing security programs; improving deployment pipelines, monitoring, observability, and reliability; and guiding software and data engineering teams on operational best practices and troubleshooting.
Top Skills: Amazon AuroraAmazon RdsAnsibleSparkAWSChefCloudFormationDockerDynamoDBGitGitGoIds/IpsJavaKafkaKinesisKubernetesLinuxMongoDBMySQLPHPPostgresPythonRubySQLTerraformTypescript
Reposted 9 Days AgoSaved
Remote
United States
180K-220K Annually
Senior level
180K-220K Annually
Senior level
Software • Defense
Work as an SRE embedded with product teams to improve reliability by fixing application code (primarily TypeScript), building observability (Prometheus, Loki, Grafana, Alloy), defining SLIs/SLOs, leading incident response and postmortems, automating toil, and supporting deployments across on‑prem DoD and AWS environments.
Top Skills: AlloyAWSBashContainersDockerGithub ActionsGitlab Ci/CdGoGrafanaJenkinsKubectlKubernetesLokiNode.jsPrometheusPythonTypescript
4 Days AgoSaved
Hybrid
Holmdel, NJ, USA
135K-160K Annually
Mid level
135K-160K Annually
Mid level
Healthtech • Software
Support reliable deployment and daily operations across AWS accounts, Linux workloads, AWS data services, networking, and Snowflake connectivity. Investigate infrastructure incidents, deployment failures, outages, and access issues. Improve automation, configuration management, monitoring, alerting, CI/CD, infrastructure-as-code practices, deployment consistency, and operational documentation. Collaborate with application engineering, data engineering, security, and other stakeholders to maintain secure, available, and dependable environments.
Top Skills: Amazon KinesisAmazon S3AWSAws GlueAws PrivatelinkBashCi/CdDnsInfrastructure As CodeLinuxPythonSnowflakeTcp/IpTls
New

Track Smarter, Apply Better.

Ditch the spreadsheets. Organize your job search with our freeApplication Tracker.

Use For Free
Application Tracker Preview
Reposted 4 Days AgoSaved
In-Office
New York, NY, USA
120K-175K Annually
Senior level
120K-175K Annually
Senior level
Fintech • Financial Services
Design, build, and maintain reliable, scalable virtual desktop infrastructure (VDI) and supporting platforms. Lead incident response, automate deployments and operations with IaC and CI/CD, implement secure configurations, monitor system health, collaborate cross-functionally, and drive continuous improvement and operational excellence.
Top Skills: Active DirectoryAnsibleArm/BicepAzure DevopsCitrix CloudCitrix GatewayCvadDnsDscGithub ActionsGitlab CiGposJenkinsPowershellSsl/Tls CertificatesTerraformVdi Profile ManagementWindows 11 Multi-SessionWindows Server
4 Days AgoSaved
Remote
United States
125K-150K Annually
Mid level
125K-150K Annually
Mid level
Aerospace • Manufacturing
Build and lead a centralized observability platform for satellite, ground-station, and distributed network systems. Responsibilities include scaling metrics, logging, and tracing infrastructure; defining SLOs, SLIs, and error budgets; enabling application instrumentation; automating deployments with Terraform and ArgoCD; monitoring Kubernetes, GCP, and AWS environments; and developing incident response, alerting, and reliability practices. The role includes on-call responsibilities and requires an active Top Secret/SCI clearance.
Top Skills: ArgocdAWSC++ElkGitlab CiGoGoogle Cloud PlatformGrafanaHoneycombIstioJaegerJavaKubernetesLinkerdLokiOpentelemetryPrometheusPythonTempoTerraform
4 Days AgoSaved
In-Office
Santa Barbara, CA, USA
175K-230K Annually
Expert/Leader
175K-230K Annually
Expert/Leader
Payments
Own platform reliability, security, automation, uptime, and incident response for payment-processing systems. Build and evolve AWS infrastructure with Terraform, manage infrastructure as code across cloud and physical environments, containerize legacy services, administer Linux servers, and embed PCI DSS compliance through patching, vulnerability management, logging, and monitoring. Provide technical leadership and support secure infrastructure for analytics and BI systems.
Top Skills: AnsibleAWSContainersLinuxOpenvoxPci DssPuppetTerraform
4 Days AgoSaved
Hybrid
5 Locations
194K-267K Annually
Senior level
194K-267K Annually
Senior level
Cloud
Build and manage highly available Kubernetes platforms on AWS. Responsibilities include Kubernetes and Helm platform creation, Terraform infrastructure automation, Karpenter scaling, Istio service mesh management, CI/CD enablement, incident response, security, compliance, monitoring, cost optimization, troubleshooting, and documentation. The role supports multi-region cloud environments and requires federal-environment access eligibility, U.S. Person documentation, and occasional travel for in-person onboarding.
Top Skills: AnsibleAWSBashCi/CdCircleCICloudFormationCloudwatchDockerEc2EcsEksElk StackGitlabGoGrafanaHelmIamIstioJenkinsKarpenterKubernetesPrometheusPythonRdsS3SpinnakerTerraformVpc
5 Days AgoSaved
In-Office
2 Locations
106K-145K Annually
Senior level
106K-145K Annually
Senior level
Hardware • Information Technology • Other • Software • Analytics
Configure, optimize, and support AWS and Azure cloud infrastructure, especially AWS GovCloud environments supporting FedRAMP-compliant services. Improve availability, performance, security, scalability, and capacity; manage monitoring, application operations, incident response, runbooks, architecture documentation, and compliance. Collaborate with development teams and international stakeholders while using scripting, infrastructure-as-code, container, CI/CD, and observability tools.
Top Skills: AnsibleAWSAws GovcloudAzureBashEcsEksFedrampGitGrafanaJenkinsJIRAKubernetesPerlPowershellPrtgPythonSumo LogicTerraform
5 Days AgoSaved
In-Office
2 Locations
91K-125K Annually
Mid level
91K-125K Annually
Mid level
Hardware • Information Technology • Other • Software • Analytics
Design, build, and maintain scalable cloud infrastructure across AWS GovCloud and Azure. Develop automated deployment, monitoring, observability, incident response, and resilience testing systems. Establish SLOs, manage error budgets, troubleshoot application operations, and enforce FedRAMP security practices. Collaborate with architecture and software engineering teams to improve availability, performance, and automation across mission-critical SaaS and PaaS platforms.
Top Skills: Active DirectoryAnsibleAWSAws GovcloudAws SdkAzureBashDockerEcsEksFedrampGrafanaJira Service ManagementKubernetesNoSQLPerlPowershellPrtgPythonRdbmsSAMLSumoTerraform
Senior level
Financial Services
Lead application support and SRE activities for mission-critical systems, improving reliability, observability, resilience, and operational efficiency. Responsibilities include resolving production incidents, managing incident and problem processes, conducting root-cause reviews, supporting releases and disaster recovery testing, maintaining operational documentation, optimizing alerts, automating workflows, and strengthening risk and control practices. The role partners with engineering, infrastructure, operations, and global teams in a highly regulated financial services environment.
Top Skills: ActivemqAmazon Rds AuroraAutosysAws Ec2Aws IamAws LambdaAws S3Aws SqsBashCicsCobolDb2Db2 Stored ProceduresDynatraceFile-AidGrafanaIbm MqJavaScriptJclJIRAKafkaLinuxOpenshiftOracleOracle AqPerlPostgresPythonRabbitMQRubySeleniumServicenowShell ScriptingSnowflakeSplunkSpufiSQLWindows
5 Days AgoSaved
In-Office
3 Locations
90K-110K Annually
Senior level
90K-110K Annually
Senior level
Mining Operations
Leads reliability improvements for steel-producing equipment and facilities, reducing downtime through predictive and preventive maintenance programs. Uses Maximo and equipment databases, analyzes failure data, and supports criticality assessments, spare-parts analysis, root cause analysis, and FMEA. Provides troubleshooting and reliability expertise, collaborates on new and modified installations, updates engineering standards, and develops best practices through periodic travel to other company sites.
Top Skills: CmmsElectrical SystemsFmeaHydraulicsIbm Maximo EamMechanical SystemsPneumaticsReliability-Centered Maintenance
5 Days AgoSaved
Remote or Hybrid
USA
136K-181K Annually
Entry level
136K-181K Annually
Entry level
Cloud • Security • Software
Design, deploy, and maintain resilient cloud infrastructure for Ping Identity’s mission-critical services. Build and optimize automated CI/CD pipelines, support cloud security and observability, evaluate technologies, participate in planning and on-call rotations, and help improve engineering practices. Collaborate across development and operations teams while sharing expertise and supporting distributed production systems.
Top Skills: Ci/CdCloud PlatformsDistributed SystemsDockerGitGoIdentity And Access ManagementKubernetesNetworking
5 Days AgoSaved
In-Office
Alpharetta, GA, USA
Senior level
Senior level
Gaming • Mobile
Enhance production-system stability, performance, scalability, and reliability through observability, automation, Kubernetes operations, and AWS infrastructure management. Define and monitor SLIs, SLOs, and SLAs; manage incidents, on-call response, troubleshooting, root-cause analysis, and post-incident remediation. Improve CI/CD production readiness, secrets management, and application resilience while documenting runbooks and operational procedures.
Top Skills: .NetAmazon Ec2Amazon EksAmazon Route 53Amazon S3Argo CdAWSAws IamBashGithub ActionsGitlab Ci/CdGraylogHashicorp VaultHelmKubernetesNew RelicPackerPythonRancherTerraform
5 Days AgoSaved
In-Office or Remote
2 Locations
144K-230K Annually
Senior level
144K-230K Annually
Senior level
Artificial Intelligence • Computer Vision • Hardware • Robotics • Metaverse
Build and deploy AI-powered tools for GeForce NOW SRE operations, transforming production signals, metrics, logs, and other data into actionable intelligence. Develop LLM- and agent-based systems for automated incident root-cause analysis and service trend prediction. Own large-scale data workflows, LLM pipelines, monitoring, cloud deployments, and AI framework architecture decisions using Kubernetes and AWS.
Top Skills: Ai/MlAWSContainer OrchestrationData PipelinesGoGrafanaKubernetesLlmsPython
5 Days AgoSaved
Hybrid
Atlanta, GA, USA
115K-140K Annually
Senior level
115K-140K Annually
Senior level
Energy
Build and support scalable, resilient cloud-native platforms and applications. Responsibilities include reliability engineering, observability, incident response, platform engineering, cloud infrastructure automation, Infrastructure as Code, CI/CD enablement, disaster recovery, chaos engineering, and AI-driven operations. The role leads reliability improvements, develops self-service tooling, partners cross-functionally on platform modernization, and mentors junior engineers.
Top Skills: APIsAWSAzureBashChaos EngineeringCi/CdCloudFormationDatadogDistributed SystemsDockerGitGitopsGrafanaHelmInfrastructure As CodeKubernetesLinuxMicroservicesNew RelicOpentelemetryPrometheusPythonSplunkTerraform
Reposted 5 Days AgoSaved
In-Office
Orlando, FL, USA
175K-215K Annually
Senior level
175K-215K Annually
Senior level
Digital Media • Gaming • News + Entertainment • Sports
Lead multiple SRE teams to ensure reliability, scalability, and security across cloud and on-prem systems. Drive observability, automation, CI/CD, infrastructure-as-code, and operational excellence; set strategy, manage resources, mentor leaders, and influence stakeholders for commerce platforms.
Top Skills: Ai/MlAnsibleAWSAzureCi/CdCloudFormationGCPGitlabHarnessInfrastructure-As-CodeKubernetesServerlessTerraform
All Filters
JobType
New Jobs
Job Category
Experience
Industry
Company Name
Company Size

Sign up now Access later

Create Free Account