Maximum of 25 job preferences reached.
Top Site Reliability Engineer Jobs
Artificial Intelligence • Big Data • Machine Learning • Software
The role involves designing and implementing custom installations of the C3 AI Platform for Federal customers, ensuring uptime, and automating system processes while collaborating with cross-functional teams.
Top Skills:
AnsibleAWSAzureBashKubernetesLinuxPuppetPythonRubyTerraform
Social Media
Operate, scale, and improve a cloud-native platform on AWS and Kubernetes. Manage GitOps deployments with ArgoCD and Helm, provision infra with Terraform/Terragrunt, build CI/CD automation, enhance observability, respond to incidents, reduce operational toil through scripting, and collaborate with security and application teams to improve reliability and platform guardrails.
Top Skills:
ArgocdAWSBashContainersEksGithub ActionsGitopsHelmIamKubernetesLinuxPythonTerraformTerragrunt
Software
Design, deploy, and automate VMware-based private cloud infrastructure across global datacenters. Administer Linux and Windows Server platforms, integrate Active Directory, manage storage, networking, ADCs (F5/AVI), and ensure availability, security, and compliance. Build automation (PowerCLI/Ansible/Python), participate in on-call rotations, document systems, and mentor junior engineers while driving infrastructure modernization and reliability improvements.
Top Skills:
Active DirectoryAnsibleAvi (Nsx Advanced Load Balancer)CentosCi/CdDnsF5 Big-IpGitNasPowercliPowershellPythonRhelSanTcp/IpUbuntuVcenter)Vmware Vsphere (EsxiVpnWindows Server
Blockchain
The Blockchain Site Reliability Engineer is responsible for maintaining blockchain nodes' reliability, monitoring, incident response, and building automation tools to enhance operations.
Top Skills:
DockerElkGoGrafanaJavaScriptKubernetesLinuxPrometheusPythonRustShell
Information Technology • Professional Services • Software • Consulting
Responsible for operating and improving reliability for large-scale hybrid (on‑prem/cloud) applications: build automation, dashboards, observability (OTEL), maintain containerized workloads (GKE/RKE/AKE), support cloud migrations (GCP/Rancher), troubleshoot networking and databases, and work with GraphQL frameworks.
Top Skills:
AkeApolloApplication Performance ManagementAutomation ScriptingClickhouseCloud InfrastructureContainerizationDashboardsDnsGCPGkeGoGraphQLHasuraHTTPJavaKubernetesLoad BalancingMongoDBObservabilityOpentelemetry (Otel)OraclePostgresPrismaPythonRancherRedisRkeRustService MeshSQL ServerTcp/IpTime-Series Databases
Artificial Intelligence • Hardware • Software • Semiconductor
Operate and scale production AI inference infrastructure, run releases and capacity changes, build self-service CD pipelines and automation, extend telemetry and observability, collaborate on SLOs, post-mortems, and capacity planning to reduce operational toil.
Top Skills:
Argo CdBazelFluxGitopsGoGrafanaInfluxdbKubernetesPrometheusPython
Artificial Intelligence • Hardware • Software • Semiconductor
Lead automation and platform engineering to eliminate toil and deliver self-service GitOps-driven CD, capacity provisioning, and observability for large-scale inference clusters. Define SLOs/SLIs, mentor SREs, support incident escalation, implement reliability practices, and measure impact via deployment velocity, SLO compliance, MTTR, and adoption of self-service workflows.
Top Skills:
Argo CdBazelCapacity PlanningChaos EngineeringCi/CdGitopsLokiMimirPredictive AutoscalingPrometheusTempoWafer-Scale Engine
Artificial Intelligence • Hardware • Software • Semiconductor
Architect and drive a unified, self-service inference control plane and reliability practices for large-scale multi-datacenter and cloud inference fleets. Build capacity orchestration, rollout safety, observability, SLO-based reliability, incident response, and automation; mentor senior SREs and measure impact through reduced toil, faster deployments, and SLO compliance.
Top Skills:
BazelCapacity ManagementChaos EngineeringGpu OrchestrationModel Serving RuntimesObservability PlatformsOrchestration SystemsSchedulersWafer-Scale Engine (Wse)
Cloud • Software
The Senior Site Reliability Engineer will automate operations using Python, manage Kubernetes and OpenStack clusters, and ensure high availability for enterprise infrastructures.
Top Skills:
KubernetesLinuxOpenstackPython
Healthtech • Database
Design, implement, and maintain observability and performance solutions for production systems. Monitor, analyze, and optimize performance, conduct capacity planning, automate operational tasks, collaborate with development and security teams, run performance testing, document systems, and drive continuous reliability and performance improvements.
Top Skills:
.NetAnsibleApm ToolsAWSAzureBashCC++Ci/CdCloudFormationConfluenceDockerDynatraceGCPGoGrafanaHTMLIisJavaJavaScriptJbossJIRAJmeterKubernetesNeoloadPerlPrometheusPythonServicenowSplunkTerraformVersion Control
Healthtech • Database
Responsible for reliability engineering, monitoring system performance, automating processes, and collaborating with development teams to enhance operational efficiency.
Top Skills:
AWSAzureBashCi/CdCloudFormationDockerDynatraceGCPGoJmeterKubernetesNeoloadPythonSplunkTerraform
Artificial Intelligence • Software
The Site Reliability Engineer ensures the reliability and performance of products Devin and Windsurf, managing incident response, CI/CD pipelines, infrastructure as code, and fostering a reliability culture within the engineering team.
Top Skills:
AWSAzureCi/CdGCPKubernetesTerraform
New
Track Smarter, Apply Better.
Ditch the spreadsheets. Organize your job search with our freeApplication Tracker.
Use For Free
Healthtech
Design, build, and maintain multi-account cloud infrastructure and CI/CD pipelines; create and manage IaC and configuration management; monitor, troubleshoot, and resolve production incidents; automate operational tasks; implement observability, dashboards, and alerts; support a customer-facing SaaS platform while meeting security and compliance requirements.
Top Skills:
AnsibleAWSDockerGCPGitGroovyLinuxMySQLPHPTerraform
Gaming
Own and operate large-scale infrastructure for sports betting and media platforms across cloud and production environments. Lead infrastructure migrations, build Kubernetes platform tooling and CI/CD automation, improve observability and alerting, support development teams, and participate in incident response. The role requires strong distributed-systems expertise, production troubleshooting, cross-team project leadership, technical communication, and mentoring.
Top Skills:
ArgocdAWSBashCephCiliumDatadogGCPGithub ActionsGoHelmIstioKubernetesLinuxPgbouncerPostgresPythonTalos OsTerraform
Fitness • Retail • Sports • Manufacturing
Design, implement, and maintain highly available, scalable infrastructure and CI/CD for applications. Automate deployments, monitor performance, troubleshoot incidents, manage IaC with Terraform, support disaster recovery, and collaborate with dev and ops teams to improve reliability and security.
Top Skills:
Application InsightsAzureAzure DevopsBashBitbucketDockerGCPGcp Cloud MonitoringGitGrafanaHelmJenkinsKubernetesPowershellPrometheusTerraform
Cloud • Software • Analytics
Join Arista Networks as a Site Reliability Engineer to manage CloudVision service reliability, scalability, and stability in a FedRAMP environment, focusing on areas like architecture, security, and performance optimization.
Top Skills:
AnsibleBashGCPGkeGoKubernetesPulumiPython
Cryptocurrency
Own production reliability, availability, and performance for cloud-native systems. Operate and scale Kubernetes (EKS) clusters, manage AWS infrastructure, implement IaC with Terraform and Helm, improve CI/CD, build observability with Prometheus/Grafana/EFK, lead incident response and RCA, participate in on-call rotations, and support security and compliance.
Top Skills:
AirflowAws BatchAws Ec2Aws LambdaAws OrganizationsBashClickhouseCloudwatchDatabricksDockerDynamoDBEfk (ElasticsearchEksElasticacheEmrFluentdGitlab Ci/CdGitopsGrafanaHelmHpaKafkaKarpenterKedaKibana)KubernetesLoad BalancingNatPostgresPrometheusPythonRdsRedisS3SnowflakeSparkSqsTerraformTlsVpcVpn
Fintech
The Principal Site Reliability Engineer designs , improves software and tools for performance, scalability, and availability, while leading incident management and collaborating with development teams.
Top Skills:
AuroraAWSChefDockerDynamo DbGitGoJavaJenkinsJmsKafkaKubernetesMavenMemcachedOraclePythonRedisSqsSwarm
Artificial Intelligence • Machine Learning • Natural Language Processing • Software • Generative AI
The Site Reliability Engineer will develop, deploy, and operate AI infrastructure, focusing on high-performance and scalable machine learning systems using Kubernetes and cloud platforms.
Top Skills:
AWSAzureC++GCPGoKubernetesOci
Fintech • Financial Services
Provide 24x7 production support and on-call rotation for cloud-native Java applications on AWS/EKS. Monitor systems with Datadog, Splunk, and CloudWatch; respond to incidents, perform RCA, manage deployments via CI/CD, automate operational tasks with Python/Bash, and improve reliability through observability, automation, SLIs/SLOs, and runbooks. Collaborate with Dev, DevOps, DB, Network, and Security teams.
Top Skills:
AlbAmazon EksArgo CdAWSAws Auto ScalingBashCloudwatchCluster AutoscalerDatadogDnsDockerEc2Github ActionsGitopsGrafanaHelmHTTPHttpsIamJavaJenkinsJvmKarpenterKubernetesLinuxMySQLNlbOpentelemetryOraclePrometheusPythonRdsRoute 53S3SplunkTcp/IpTerraformTlsVpc
Information Technology • Software
Administer and maintain distributed Splunk Enterprise and Cloud platforms for Navy cyber operations. Monitor ingestion, performance, onboarding, and compliance (STIG); troubleshoot incidents; build dashboards and alerts; automate tasks and integrate with CI/CD; participate in incident response, upgrades, modernization, documentation, and on-call support.
Top Skills:
AnsibleBashChefCi/CdDeployment ServerGlass TablesHeavy ForwarderIndexerJenkinsPowershellPythonRest ApiRhel 8Rhel 9RmfSearch HeadService AnalyzerSIEMSplunk CloudSplunk EnterpriseSplunk ItsiSslStigSyslogTcp/UdpTerraformUniversal Forwarder
Artificial Intelligence • Software • Generative AI • Automation
Operate and harden Blitzy's self-hosted, Kubernetes-based AI platform inside customer-controlled secure cloud environments. Own deployments, upgrades, capacity planning, observability, incident response, and customer-facing technical coordination while championing security and feeding operational learnings back into the product roadmap.
Top Skills:
Alerting)BashCloud (Aws/Gcp/Azure)Container OrchestrationGoInfrastructure-As-CodeKubernetesMetricsObservability (LoggingPulumiPythonTerraformTracing
Fintech
Lead design and implementation of highly available, fault-tolerant platform architectures. Define SLOs/SLAs, drive incident and problem management, build observability and automation, mentor engineers, and partner with stakeholders to improve reliability, performance, and compliance.
Top Skills:
AWSAzureCi/CdDevOpsIncident Management ToolingMonitoringObservability (LoggingSlo/Sli FrameworksTracing)
Healthtech • Insurance
Design, build, and operate scalable, resilient cloud infrastructure and SRE tooling. Lead cross-team technical projects, mentor engineers, define SLOs and incident management, automate CI/CD and IaC, reduce failure domains, and own medium-to-large infrastructure features from design through delivery.
Top Skills:
ArgocdAWSCi/CdGCPGithub ActionsGrafanaIamIstioKubernetesPrometheusTerraformVpc Peering
Healthtech • Insurance
Lead design, development, and operation of cloud infrastructure and SRE-focused systems. Own medium-to-large infrastructure projects, build resilient platforms, and drive cross-team technical delivery. Mentor engineers, define SLOs, reduce failure domains, and build tooling for automated, secure CI/CD and production reliability.
Top Skills:
ArgocdAWSGCPGithub ActionsGrafanaIamIstioKubernetesPrometheusTerraformVpc Peering
Let Your Resume Do The Work
Upload your resume to be matched with jobs you're a great fit for.
Success! We'll use this to further personalize your experience.
Popular Job Searches
All Software Engineer Jobs
.NET Developer Jobs
Aerospace Thermal Engineering Jobs
AI Engineer Jobs
Android Developer Jobs
Automation Engineer Jobs
Backend Developer Jobs
Blockchain Developer Jobs
C# Jobs
C++ Jobs
Cloud Architect Jobs
Cloud Engineer Jobs
Design Engineer Jobs
DevOps Engineer Jobs
Director Of Engineering Jobs
Electrical Engineering Jobs
Embedded Software Engineer Jobs
Engineering Jobs
Engineering Manager Jobs
Environmental Engineering Jobs
Field Engineer Jobs
Front End Developer Jobs
Full Stack Developer Jobs
Game Developer Jobs
Golang Jobs
Hardware Engineer Jobs
Industrial Engineering Jobs
iOS Developer Jobs
Java Developer Jobs
Javascript Developer Jobs
Linux Jobs
Manufacturing Engineer Jobs
Mechanical Engineering Jobs
Network Engineer Jobs
PHP Developer Jobs
Process Engineer Jobs
Project Engineer Jobs
Prompt Engineering Jobs
Python Jobs
QA Jobs
Robotics Engineer Jobs
Ruby on Rails Jobs
Salesforce Administrator Jobs
Salesforce Developer Jobs
Scala Jobs
Sharepoint Developer Jobs
Site Reliability Engineer Jobs
Software Engineering Manager Jobs
Solutions Architect Jobs
SQL Developer Jobs
Structural Engineer Jobs
System Engineer Jobs
Test Engineer Jobs
Web Developer Jobs
All Filters
Total selected ()
No Results
No Results


.png)





























