Maximum of 25 job preferences reached.
Top Site Reliability Engineer Jobs
Logistics • Mobile • Productivity • Software • Transportation
The Senior Site Reliability Engineer will manage the reliability of Zello's data tier, contribute to monitoring and incident response while improving cloud infrastructure and database performance.
Top Skills:
BashDockerElasticsearchGoKubernetesLokiMongoDBMySQLPrometheusPythonRedisScylladbTempo
Blockchain • Fintech • Payments • Consulting • Cryptocurrency • Cybersecurity • Quantum Computing
Lead performance and reliability engineering for ABU services: design and run automated performance tests, analyze bottlenecks, tune systems (including JVM/GC), support capacity planning and chaos testing, automate repeatable tasks, review incidents, and mentor engineers to improve scalability, resiliency, and engineering velocity.
Top Skills:
Ai-Assisted ToolsBitbucketBlazemeterCloud-Native StacksGarbage Collection TuningGatlingJavaJenkinsJvmLarge Language ModelsLoadrunnerMavenPythonScala
Fintech • Payments • Financial Services
Leads the design and operation of highly available cloud systems, primarily on AWS, using Terraform, GitLab CI/CD, containers, and observability tools. Responsibilities include reliability engineering, incident response, disaster recovery, automation, monitoring, security integration, vulnerability management, and cost optimization. The role builds internal tools, guides architecture, mentors SRE engineers, documents operational processes, and partners with software engineering teams to improve production reliability and compliance.
Top Skills:
Agentic ApplicationsAmazon Api GatewayAmazon AuroraAmazon CloudfrontAmazon CloudwatchAmazon DynamodbAmazon EbsAmazon Ec2Amazon EcsAmazon EfsAmazon RdsAmazon Route 53Amazon S3Amazon VpcAWSAws FargateAws LambdaAws X-RayDastDatadogDockerGitlabGrafanaIamJavaKubernetesLlmsNew RelicNode.jsOwasp Top 10PythonSastSplunkTerraform
Information Technology • Internet of Things • Software • Virtual Reality
Lead reliability, scalability, and operational excellence for large-scale Java distributed systems. Drive architecture improvements, incident management, observability, SLO adoption, automation, infrastructure strategy, and performance optimization. Influence cross-team technical decisions, mentor engineers, establish reliability standards, and guide incident response and durable remediation. The role requires hybrid work in Boston and participation in an on-call rotation.
Top Skills:
AWSCi/CdInfrastructure As CodeJavaMongoDBRabbitMQZookeeper
Fitness • Healthtech • Retail • Pharmaceutical
Ensure reliability, scalability, and performance of distributed retail and pharmacy systems. Implement observability, monitoring, SLOs, incident response, and reliability improvements. Lead cloud-native microservices and Kubernetes/OpenShift deployments, CI/CD automation, performance testing, and collaborate with engineering and operations. Participate in on-call rotation.
Top Skills:
Ai/AiopsApigeeAWSBitbucketDatadogDatapowerDockerDynatraceGitGCPGrafanaJavaJenkinsKubernetesAzureOpenshiftPrometheusPythonRancherSplunkVordelWeb Apis
Cloud • Information Technology • Internet of Things • Professional Services • Software
Designs, deploys, and operates reliable, scalable cloud and network infrastructure; automates platforms; monitors SLOs/SLAs; leads on-call incident response, postmortems, DR planning, and security controls to improve service reliability.
Top Skills:
AutomationCloud InfrastructureDisaster RecoveryFedramp HighIl-5Monitoring And AlertingNetworkingOn-Call Incident ManagementSecurity ControlsSre PracticesStorage Systems
Artificial Intelligence • Information Technology • Software • Consulting
Operate and improve the reliability, availability, and performance of large-scale distributed systems. Build automation and tooling, manage Linux and Kubernetes environments, design CI/CD pipelines, implement observability, troubleshoot production systems, and reduce operational toil. Preferred work includes defining SLOs and error budgets, chaos engineering, cloud operations, capacity planning, load testing, and service mesh management.
Top Skills:
AWSAzureChaos MonkeyCi/CdConsulDockerEfkElkGCPGoGrafanaGremlinIstioJavaKubernetesLinkerdLinuxLitmusOpentelemetryPrometheusPython
Artificial Intelligence • Computer Vision • Hardware • Robotics • Metaverse
Lead the transformation, design, deployment, and operation of globally scaled on-premises and cloud compute infrastructure. Build reliable core services including DNS, NTP/PTP, DHCP, and LDAP; develop automation, monitoring, capacity planning, and lifecycle management. Define performance metrics, optimize systems using technologies such as SR-IOV and DPUs, and develop data analysis and visualization tools. Partner with engineering, product, program, and company leadership to deliver scalable IT services.
Top Skills:
AnycastBare MetalBgpConfiguration Management ToolsContainersDhcpDnsDpuEbpfGoInfrastructure As CodeLdapLinuxLinux KernelMicroservicesNtpPtpPythonSdnSr-IovTerraformVlanVxlanXdp
Software • Quantum Computing • Metaverse • Infrastructure as a Service (IaaS)
Own reliability and operational health for Microsoft Substrate cloud services in regulated environments. Participate in on-call rotations, diagnose and resolve production incidents, build automation to reduce toil, maintain production code, develop monitoring and telemetry supporting SLOs, lead post-incident reviews, and collaborate with software engineering teams to improve service reliability, scalability, security, and operability.
Top Skills:
Exchange OnlineMicrosoft 365 CopilotMicrosoft CloudMicrosoft Substrate
Digital Media
Build, maintain, and operate Ookla’s globally distributed infrastructure platform at massive scale. Responsibilities include managing cloud instances, containers, serverless applications, databases, streaming systems, and big-data tooling; supporting 24/7 production operations and on-call rotations; implementing security programs; improving deployment pipelines, monitoring, observability, and reliability; and guiding software and data engineering teams on operational best practices and troubleshooting.
Top Skills:
Amazon AuroraAmazon RdsAnsibleSparkAWSChefCloudFormationDockerDynamoDBGitGitGoIds/IpsJavaKafkaKinesisKubernetesLinuxMongoDBMySQLPHPPostgresPythonRubySQLTerraformTypescript
Reposted 9 Days AgoSaved
Software • Defense
Work as an SRE embedded with product teams to improve reliability by fixing application code (primarily TypeScript), building observability (Prometheus, Loki, Grafana, Alloy), defining SLIs/SLOs, leading incident response and postmortems, automating toil, and supporting deployments across on‑prem DoD and AWS environments.
Top Skills:
AlloyAWSBashContainersDockerGithub ActionsGitlab Ci/CdGoGrafanaJenkinsKubectlKubernetesLokiNode.jsPrometheusPythonTypescript
Healthtech • Software
Support reliable deployment and daily operations across AWS accounts, Linux workloads, AWS data services, networking, and Snowflake connectivity. Investigate infrastructure incidents, deployment failures, outages, and access issues. Improve automation, configuration management, monitoring, alerting, CI/CD, infrastructure-as-code practices, deployment consistency, and operational documentation. Collaborate with application engineering, data engineering, security, and other stakeholders to maintain secure, available, and dependable environments.
Top Skills:
Amazon KinesisAmazon S3AWSAws GlueAws PrivatelinkBashCi/CdDnsInfrastructure As CodeLinuxPythonSnowflakeTcp/IpTls
New
Track Smarter, Apply Better.
Ditch the spreadsheets. Organize your job search with our freeApplication Tracker.
Use For Free
Fintech • Financial Services
Design, build, and maintain reliable, scalable virtual desktop infrastructure (VDI) and supporting platforms. Lead incident response, automate deployments and operations with IaC and CI/CD, implement secure configurations, monitor system health, collaborate cross-functionally, and drive continuous improvement and operational excellence.
Top Skills:
Active DirectoryAnsibleArm/BicepAzure DevopsCitrix CloudCitrix GatewayCvadDnsDscGithub ActionsGitlab CiGposJenkinsPowershellSsl/Tls CertificatesTerraformVdi Profile ManagementWindows 11 Multi-SessionWindows Server
Aerospace • Manufacturing
Build and lead a centralized observability platform for satellite, ground-station, and distributed network systems. Responsibilities include scaling metrics, logging, and tracing infrastructure; defining SLOs, SLIs, and error budgets; enabling application instrumentation; automating deployments with Terraform and ArgoCD; monitoring Kubernetes, GCP, and AWS environments; and developing incident response, alerting, and reliability practices. The role includes on-call responsibilities and requires an active Top Secret/SCI clearance.
Top Skills:
ArgocdAWSC++ElkGitlab CiGoGoogle Cloud PlatformGrafanaHoneycombIstioJaegerJavaKubernetesLinkerdLokiOpentelemetryPrometheusPythonTempoTerraform
Payments
Own platform reliability, security, automation, uptime, and incident response for payment-processing systems. Build and evolve AWS infrastructure with Terraform, manage infrastructure as code across cloud and physical environments, containerize legacy services, administer Linux servers, and embed PCI DSS compliance through patching, vulnerability management, logging, and monitoring. Provide technical leadership and support secure infrastructure for analytics and BI systems.
Top Skills:
AnsibleAWSContainersLinuxOpenvoxPci DssPuppetTerraform
Cloud
Build and manage highly available Kubernetes platforms on AWS. Responsibilities include Kubernetes and Helm platform creation, Terraform infrastructure automation, Karpenter scaling, Istio service mesh management, CI/CD enablement, incident response, security, compliance, monitoring, cost optimization, troubleshooting, and documentation. The role supports multi-region cloud environments and requires federal-environment access eligibility, U.S. Person documentation, and occasional travel for in-person onboarding.
Top Skills:
AnsibleAWSBashCi/CdCircleCICloudFormationCloudwatchDockerEc2EcsEksElk StackGitlabGoGrafanaHelmIamIstioJenkinsKarpenterKubernetesPrometheusPythonRdsS3SpinnakerTerraformVpc
Hardware • Information Technology • Other • Software • Analytics
Configure, optimize, and support AWS and Azure cloud infrastructure, especially AWS GovCloud environments supporting FedRAMP-compliant services. Improve availability, performance, security, scalability, and capacity; manage monitoring, application operations, incident response, runbooks, architecture documentation, and compliance. Collaborate with development teams and international stakeholders while using scripting, infrastructure-as-code, container, CI/CD, and observability tools.
Top Skills:
AnsibleAWSAws GovcloudAzureBashEcsEksFedrampGitGrafanaJenkinsJIRAKubernetesPerlPowershellPrtgPythonSumo LogicTerraform
Hardware • Information Technology • Other • Software • Analytics
Design, build, and maintain scalable cloud infrastructure across AWS GovCloud and Azure. Develop automated deployment, monitoring, observability, incident response, and resilience testing systems. Establish SLOs, manage error budgets, troubleshoot application operations, and enforce FedRAMP security practices. Collaborate with architecture and software engineering teams to improve availability, performance, and automation across mission-critical SaaS and PaaS platforms.
Top Skills:
Active DirectoryAnsibleAWSAws GovcloudAws SdkAzureBashDockerEcsEksFedrampGrafanaJira Service ManagementKubernetesNoSQLPerlPowershellPrtgPythonRdbmsSAMLSumoTerraform
5 Days AgoSaved
Financial Services
Lead application support and SRE activities for mission-critical systems, improving reliability, observability, resilience, and operational efficiency. Responsibilities include resolving production incidents, managing incident and problem processes, conducting root-cause reviews, supporting releases and disaster recovery testing, maintaining operational documentation, optimizing alerts, automating workflows, and strengthening risk and control practices. The role partners with engineering, infrastructure, operations, and global teams in a highly regulated financial services environment.
Top Skills:
ActivemqAmazon Rds AuroraAutosysAws Ec2Aws IamAws LambdaAws S3Aws SqsBashCicsCobolDb2Db2 Stored ProceduresDynatraceFile-AidGrafanaIbm MqJavaScriptJclJIRAKafkaLinuxOpenshiftOracleOracle AqPerlPostgresPythonRabbitMQRubySeleniumServicenowShell ScriptingSnowflakeSplunkSpufiSQLWindows
Mining Operations
Leads reliability improvements for steel-producing equipment and facilities, reducing downtime through predictive and preventive maintenance programs. Uses Maximo and equipment databases, analyzes failure data, and supports criticality assessments, spare-parts analysis, root cause analysis, and FMEA. Provides troubleshooting and reliability expertise, collaborates on new and modified installations, updates engineering standards, and develops best practices through periodic travel to other company sites.
Top Skills:
CmmsElectrical SystemsFmeaHydraulicsIbm Maximo EamMechanical SystemsPneumaticsReliability-Centered Maintenance
Cloud • Security • Software
Design, deploy, and maintain resilient cloud infrastructure for Ping Identity’s mission-critical services. Build and optimize automated CI/CD pipelines, support cloud security and observability, evaluate technologies, participate in planning and on-call rotations, and help improve engineering practices. Collaborate across development and operations teams while sharing expertise and supporting distributed production systems.
Top Skills:
Ci/CdCloud PlatformsDistributed SystemsDockerGitGoIdentity And Access ManagementKubernetesNetworking
Gaming • Mobile
Enhance production-system stability, performance, scalability, and reliability through observability, automation, Kubernetes operations, and AWS infrastructure management. Define and monitor SLIs, SLOs, and SLAs; manage incidents, on-call response, troubleshooting, root-cause analysis, and post-incident remediation. Improve CI/CD production readiness, secrets management, and application resilience while documenting runbooks and operational procedures.
Top Skills:
.NetAmazon Ec2Amazon EksAmazon Route 53Amazon S3Argo CdAWSAws IamBashGithub ActionsGitlab Ci/CdGraylogHashicorp VaultHelmKubernetesNew RelicPackerPythonRancherTerraform
Artificial Intelligence • Computer Vision • Hardware • Robotics • Metaverse
Build and deploy AI-powered tools for GeForce NOW SRE operations, transforming production signals, metrics, logs, and other data into actionable intelligence. Develop LLM- and agent-based systems for automated incident root-cause analysis and service trend prediction. Own large-scale data workflows, LLM pipelines, monitoring, cloud deployments, and AI framework architecture decisions using Kubernetes and AWS.
Top Skills:
Ai/MlAWSContainer OrchestrationData PipelinesGoGrafanaKubernetesLlmsPython
Energy
Build and support scalable, resilient cloud-native platforms and applications. Responsibilities include reliability engineering, observability, incident response, platform engineering, cloud infrastructure automation, Infrastructure as Code, CI/CD enablement, disaster recovery, chaos engineering, and AI-driven operations. The role leads reliability improvements, develops self-service tooling, partners cross-functionally on platform modernization, and mentors junior engineers.
Top Skills:
APIsAWSAzureBashChaos EngineeringCi/CdCloudFormationDatadogDistributed SystemsDockerGitGitopsGrafanaHelmInfrastructure As CodeKubernetesLinuxMicroservicesNew RelicOpentelemetryPrometheusPythonSplunkTerraform
Digital Media • Gaming • News + Entertainment • Sports
Lead multiple SRE teams to ensure reliability, scalability, and security across cloud and on-prem systems. Drive observability, automation, CI/CD, infrastructure-as-code, and operational excellence; set strategy, manage resources, mentor leaders, and influence stakeholders for commerce platforms.
Top Skills:
Ai/MlAnsibleAWSAzureCi/CdCloudFormationGCPGitlabHarnessInfrastructure-As-CodeKubernetesServerlessTerraform
Let Your Resume Do The Work
Upload your resume to be matched with jobs you're a great fit for.
Success! We'll use this to further personalize your experience.
Top Companies Hiring Site Reliability Engineers
See AllPopular Job Searches
All Software Engineer Jobs
.NET Developer Jobs
Aerospace Thermal Engineering Jobs
AI Engineer Jobs
Android Developer Jobs
Automation Engineer Jobs
Backend Developer Jobs
Blockchain Developer Jobs
C# Jobs
C++ Jobs
Cloud Architect Jobs
Cloud Engineer Jobs
Design Engineer Jobs
DevOps Engineer Jobs
Director Of Engineering Jobs
Electrical Engineering Jobs
Embedded Software Engineer Jobs
Engineering Jobs
Engineering Manager Jobs
Environmental Engineering Jobs
Field Engineer Jobs
Front End Developer Jobs
Full Stack Developer Jobs
Game Developer Jobs
Golang Jobs
Hardware Engineer Jobs
Industrial Engineering Jobs
iOS Developer Jobs
Java Developer Jobs
Javascript Developer Jobs
Linux Jobs
Manufacturing Engineer Jobs
Mechanical Engineering Jobs
Network Engineer Jobs
PHP Developer Jobs
Process Engineer Jobs
Project Engineer Jobs
Prompt Engineering Jobs
Python Jobs
QA Jobs
Robotics Engineer Jobs
Ruby on Rails Jobs
Salesforce Administrator Jobs
Salesforce Developer Jobs
Scala Jobs
Sharepoint Developer Jobs
Site Reliability Engineer Jobs
Software Engineering Manager Jobs
Solutions Architect Jobs
SQL Developer Jobs
Structural Engineer Jobs
System Engineer Jobs
Test Engineer Jobs
Web Developer Jobs
All Filters
Total selected ()
No Results
No Results
















.jpeg)















