Top Site Reliability Engineer Jobs

Reposted 26 Days AgoSaved
In-Office
Tyson's Corner, VA, USA
159K-230K Annually
Senior level
159K-230K Annually
Senior level
Artificial Intelligence • Big Data • Machine Learning • Software
The role involves designing and implementing custom installations of the C3 AI Platform for Federal customers, ensuring uptime, and automating system processes while collaborating with cross-functional teams.
Top Skills: AnsibleAWSAzureBashKubernetesLinuxPuppetPythonRubyTerraform
Reposted 26 Days AgoSaved
In-Office or Remote
San Francisco, CA, USA
114K-235K Annually
Mid level
114K-235K Annually
Mid level
Social Media
Operate, scale, and improve a cloud-native platform on AWS and Kubernetes. Manage GitOps deployments with ArgoCD and Helm, provision infra with Terraform/Terragrunt, build CI/CD automation, enhance observability, respond to incidents, reduce operational toil through scripting, and collaborate with security and application teams to improve reliability and platform guardrails.
Top Skills: ArgocdAWSBashContainersEksGithub ActionsGitopsHelmIamKubernetesLinuxPythonTerraformTerragrunt
Reposted 26 Days AgoSaved
In-Office
New York, NY, USA
131K-164K Annually
Expert/Leader
131K-164K Annually
Expert/Leader
Software
Design, deploy, and automate VMware-based private cloud infrastructure across global datacenters. Administer Linux and Windows Server platforms, integrate Active Directory, manage storage, networking, ADCs (F5/AVI), and ensure availability, security, and compliance. Build automation (PowerCLI/Ansible/Python), participate in on-call rotations, document systems, and mentor junior engineers while driving infrastructure modernization and reliability improvements.
Top Skills: Active DirectoryAnsibleAvi (Nsx Advanced Load Balancer)CentosCi/CdDnsF5 Big-IpGitNasPowercliPowershellPythonRhelSanTcp/IpUbuntuVcenter)Vmware Vsphere (EsxiVpnWindows Server
Reposted 26 Days AgoSaved
Remote
Texas, USA
Mid level
Mid level
Blockchain
The Blockchain Site Reliability Engineer is responsible for maintaining blockchain nodes' reliability, monitoring, incident response, and building automation tools to enhance operations.
Top Skills: DockerElkGoGrafanaJavaScriptKubernetesLinuxPrometheusPythonRustShell
27 Days AgoSaved
In-Office
Scottsdale, AZ, USA
Mid level
Mid level
Information Technology • Professional Services • Software • Consulting
Responsible for operating and improving reliability for large-scale hybrid (on‑prem/cloud) applications: build automation, dashboards, observability (OTEL), maintain containerized workloads (GKE/RKE/AKE), support cloud migrations (GCP/Rancher), troubleshoot networking and databases, and work with GraphQL frameworks.
Top Skills: AkeApolloApplication Performance ManagementAutomation ScriptingClickhouseCloud InfrastructureContainerizationDashboardsDnsGCPGkeGoGraphQLHasuraHTTPJavaKubernetesLoad BalancingMongoDBObservabilityOpentelemetry (Otel)OraclePostgresPrismaPythonRancherRedisRkeRustService MeshSQL ServerTcp/IpTime-Series Databases
27 Days AgoSaved
Remote
2 Locations
Mid level
Mid level
Artificial Intelligence • Hardware • Software • Semiconductor
Operate and scale production AI inference infrastructure, run releases and capacity changes, build self-service CD pipelines and automation, extend telemetry and observability, collaborate on SLOs, post-mortems, and capacity planning to reduce operational toil.
Top Skills: Argo CdBazelFluxGitopsGoGrafanaInfluxdbKubernetesPrometheusPython
Senior level
Artificial Intelligence • Hardware • Software • Semiconductor
Lead automation and platform engineering to eliminate toil and deliver self-service GitOps-driven CD, capacity provisioning, and observability for large-scale inference clusters. Define SLOs/SLIs, mentor SREs, support incident escalation, implement reliability practices, and measure impact via deployment velocity, SLO compliance, MTTR, and adoption of self-service workflows.
Top Skills: Argo CdBazelCapacity PlanningChaos EngineeringCi/CdGitopsLokiMimirPredictive AutoscalingPrometheusTempoWafer-Scale Engine
27 Days AgoSaved
Hybrid
Sunnyvale, CA, USA
Expert/Leader
Expert/Leader
Artificial Intelligence • Hardware • Software • Semiconductor
Architect and drive a unified, self-service inference control plane and reliability practices for large-scale multi-datacenter and cloud inference fleets. Build capacity orchestration, rollout safety, observability, SLO-based reliability, incident response, and automation; mentor senior SREs and measure impact through reduced toil, faster deployments, and SLO compliance.
Top Skills: BazelCapacity ManagementChaos EngineeringGpu OrchestrationModel Serving RuntimesObservability PlatformsOrchestration SystemsSchedulersWafer-Scale Engine (Wse)
Reposted 12 Days AgoSaved
In-Office or Remote
7 Locations
Senior level
Senior level
Cloud • Software
The Senior Site Reliability Engineer will automate operations using Python, manage Kubernetes and OpenStack clusters, and ensure high availability for enterprise infrastructures.
Top Skills: KubernetesLinuxOpenstackPython
27 Days AgoSaved
In-Office
2 Locations
85K-110K Annually
Mid level
85K-110K Annually
Mid level
Healthtech • Database
Design, implement, and maintain observability and performance solutions for production systems. Monitor, analyze, and optimize performance, conduct capacity planning, automate operational tasks, collaborate with development and security teams, run performance testing, document systems, and drive continuous reliability and performance improvements.
Top Skills: .NetAnsibleApm ToolsAWSAzureBashCC++Ci/CdCloudFormationConfluenceDockerDynatraceGCPGoGrafanaHTMLIisJavaJavaScriptJbossJIRAJmeterKubernetesNeoloadPerlPrometheusPythonServicenowSplunkTerraformVersion Control
Reposted 27 Days AgoSaved
In-Office
Secaucus, NJ, USA
90K-120K Annually
Mid level
90K-120K Annually
Mid level
Healthtech • Database
Responsible for reliability engineering, monitoring system performance, automating processes, and collaborating with development teams to enhance operational efficiency.
Top Skills: AWSAzureBashCi/CdCloudFormationDockerDynatraceGCPGoJmeterKubernetesNeoloadPythonSplunkTerraform
Reposted 28 Days AgoSaved
In-Office
2 Locations
Senior level
Senior level
Artificial Intelligence • Software
The Site Reliability Engineer ensures the reliability and performance of products Devin and Windsurf, managing incident response, CI/CD pipelines, infrastructure as code, and fostering a reliability culture within the engineering team.
Top Skills: AWSAzureCi/CdGCPKubernetesTerraform
New

Track Smarter, Apply Better.

Ditch the spreadsheets. Organize your job search with our freeApplication Tracker.

Use For Free
Application Tracker Preview
28 Days AgoSaved
In-Office
Edison, NJ, USA
Senior level
Senior level
Healthtech
Design, build, and maintain multi-account cloud infrastructure and CI/CD pipelines; create and manage IaC and configuration management; monitor, troubleshoot, and resolve production incidents; automate operational tasks; implement observability, dashboards, and alerts; support a customer-facing SaaS platform while meeting security and compliance requirements.
Top Skills: AnsibleAWSDockerGCPGitGroovyLinuxMySQLPHPTerraform
6 Days AgoSaved
Remote
United States
145K-193K Annually
Senior level
145K-193K Annually
Senior level
Gaming
Own and operate large-scale infrastructure for sports betting and media platforms across cloud and production environments. Lead infrastructure migrations, build Kubernetes platform tooling and CI/CD automation, improve observability and alerting, support development teams, and participate in incident response. The role requires strong distributed-systems expertise, production troubleshooting, cross-team project leadership, technical communication, and mentoring.
Top Skills: ArgocdAWSBashCephCiliumDatadogGCPGithub ActionsGoHelmIstioKubernetesLinuxPgbouncerPostgresPythonTalos OsTerraform
Reposted 28 Days AgoSaved
In-Office
Columbus, OH, USA
Senior level
Senior level
Fitness • Retail • Sports • Manufacturing
Design, implement, and maintain highly available, scalable infrastructure and CI/CD for applications. Automate deployments, monitor performance, troubleshoot incidents, manage IaC with Terraform, support disaster recovery, and collaborate with dev and ops teams to improve reliability and security.
Top Skills: Application InsightsAzureAzure DevopsBashBitbucketDockerGCPGcp Cloud MonitoringGitGrafanaHelmJenkinsKubernetesPowershellPrometheusTerraform
Reposted 28 Days AgoSaved
Remote
US
101K-161K Annually
Senior level
101K-161K Annually
Senior level
Cloud • Software • Analytics
Join Arista Networks as a Site Reliability Engineer to manage CloudVision service reliability, scalability, and stability in a FedRAMP environment, focusing on areas like architecture, security, and performance optimization.
Top Skills: AnsibleBashGCPGkeGoKubernetesPulumiPython
Reposted 28 Days AgoSaved
Hybrid
New York, NY, USA
Mid level
Mid level
Cryptocurrency
Own production reliability, availability, and performance for cloud-native systems. Operate and scale Kubernetes (EKS) clusters, manage AWS infrastructure, implement IaC with Terraform and Helm, improve CI/CD, build observability with Prometheus/Grafana/EFK, lead incident response and RCA, participate in on-call rotations, and support security and compliance.
Top Skills: AirflowAws BatchAws Ec2Aws LambdaAws OrganizationsBashClickhouseCloudwatchDatabricksDockerDynamoDBEfk (ElasticsearchEksElasticacheEmrFluentdGitlab Ci/CdGitopsGrafanaHelmHpaKafkaKarpenterKedaKibana)KubernetesLoad BalancingNatPostgresPrometheusPythonRdsRedisS3SnowflakeSparkSqsTerraformTlsVpcVpn
Reposted 28 Days AgoSaved
In-Office
Scottsdale, AZ, USA
194K-237K Annually
Expert/Leader
194K-237K Annually
Expert/Leader
Fintech
The Principal Site Reliability Engineer designs , improves software and tools for performance, scalability, and availability, while leading incident management and collaborating with development teams.
Top Skills: AuroraAWSChefDockerDynamo DbGitGoJavaJenkinsJmsKafkaKubernetesMavenMemcachedOraclePythonRedisSqsSwarm
Reposted 28 Days AgoSaved
In-Office or Remote
5 Locations
160K-260K Annually
Senior level
160K-260K Annually
Senior level
Artificial Intelligence • Machine Learning • Natural Language Processing • Software • Generative AI
The Site Reliability Engineer will develop, deploy, and operate AI infrastructure, focusing on high-performance and scalable machine learning systems using Kubernetes and cloud platforms.
Top Skills: AWSAzureC++GCPGoKubernetesOci
29 Days AgoSaved
In-Office
30328, Atlanta, GA, USA
100K-115K Annually
Senior level
100K-115K Annually
Senior level
Fintech • Financial Services
Provide 24x7 production support and on-call rotation for cloud-native Java applications on AWS/EKS. Monitor systems with Datadog, Splunk, and CloudWatch; respond to incidents, perform RCA, manage deployments via CI/CD, automate operational tasks with Python/Bash, and improve reliability through observability, automation, SLIs/SLOs, and runbooks. Collaborate with Dev, DevOps, DB, Network, and Security teams.
Top Skills: AlbAmazon EksArgo CdAWSAws Auto ScalingBashCloudwatchCluster AutoscalerDatadogDnsDockerEc2Github ActionsGitopsGrafanaHelmHTTPHttpsIamJavaJenkinsJvmKarpenterKubernetesLinuxMySQLNlbOpentelemetryOraclePrometheusPythonRdsRoute 53S3SplunkTcp/IpTerraformTlsVpc
29 Days AgoSaved
In-Office
Norfolk, VA, USA
108K-195K Annually
Senior level
108K-195K Annually
Senior level
Information Technology • Software
Administer and maintain distributed Splunk Enterprise and Cloud platforms for Navy cyber operations. Monitor ingestion, performance, onboarding, and compliance (STIG); troubleshoot incidents; build dashboards and alerts; automate tasks and integrate with CI/CD; participate in incident response, upgrades, modernization, documentation, and on-call support.
Top Skills: AnsibleBashChefCi/CdDeployment ServerGlass TablesHeavy ForwarderIndexerJenkinsPowershellPythonRest ApiRhel 8Rhel 9RmfSearch HeadService AnalyzerSIEMSplunk CloudSplunk EnterpriseSplunk ItsiSslStigSyslogTcp/UdpTerraformUniversal Forwarder
29 Days AgoSaved
Remote
USA
140K-170K Annually
Mid level
140K-170K Annually
Mid level
Artificial Intelligence • Software • Generative AI • Automation
Operate and harden Blitzy's self-hosted, Kubernetes-based AI platform inside customer-controlled secure cloud environments. Own deployments, upgrades, capacity planning, observability, incident response, and customer-facing technical coordination while championing security and feeding operational learnings back into the product roadmap.
Top Skills: Alerting)BashCloud (Aws/Gcp/Azure)Container OrchestrationGoInfrastructure-As-CodeKubernetesMetricsObservability (LoggingPulumiPythonTerraformTracing
29 Days AgoSaved
In-Office
Buffalo, NY, USA
140K-233K Annually
Senior level
140K-233K Annually
Senior level
Fintech
Lead design and implementation of highly available, fault-tolerant platform architectures. Define SLOs/SLAs, drive incident and problem management, build observability and automation, mentor engineers, and partner with stakeholders to improve reliability, performance, and compliance.
Top Skills: AWSAzureCi/CdDevOpsIncident Management ToolingMonitoringObservability (LoggingSlo/Sli FrameworksTracing)
29 Days AgoSaved
In-Office
Tempe, AZ, USA
181K-237K Annually
Senior level
181K-237K Annually
Senior level
Healthtech • Insurance
Design, build, and operate scalable, resilient cloud infrastructure and SRE tooling. Lead cross-team technical projects, mentor engineers, define SLOs and incident management, automate CI/CD and IaC, reduce failure domains, and own medium-to-large infrastructure features from design through delivery.
Top Skills: ArgocdAWSCi/CdGCPGithub ActionsGrafanaIamIstioKubernetesPrometheusTerraformVpc Peering
29 Days AgoSaved
In-Office
San Francisco, CA, USA
181K-237K Annually
Senior level
181K-237K Annually
Senior level
Healthtech • Insurance
Lead design, development, and operation of cloud infrastructure and SRE-focused systems. Own medium-to-large infrastructure projects, build resilient platforms, and drive cross-team technical delivery. Mentor engineers, define SLOs, reduce failure domains, and build tooling for automated, secure CI/CD and production reliability.
Top Skills: ArgocdAWSGCPGithub ActionsGrafanaIamIstioKubernetesPrometheusTerraformVpc Peering
All Filters
JobType
New Jobs
Job Category
Experience
Industry
Company Name
Company Size

Sign up now Access later

Create Free Account