Maximum of 25 job preferences reached.
Top Site Reliability Engineer Jobs
Energy
Build, transition, operate, and improve a multi-cloud production platform across Azure, GCP, AWS, and Kubernetes. Responsibilities include infrastructure as code, GitOps and CI/CD, managed Kubernetes operations, troubleshooting, incident response, observability, reliability improvements, disaster recovery, automation, and cost control. The role partners with architects and global teams to ensure platforms are secure, scalable, recoverable, and supportable.
Top Skills:
AksAnsibleAWSAzureCi/CdEksGitopsGkeGoGoogle Cloud PlatformKubernetesLinuxMongoDBMySQLNetworkingOpentofuPostgresPythonShell ScriptingTerraformTerragrunt
Artificial Intelligence • Healthtech • Information Technology • Software
As a Site Reliability Engineer, you will manage the production environment, focusing on infrastructure design, automation, and optimizing deployment pipelines to ensure high availability.
Top Skills:
HelmKafkaKubernetesPostgresPythonRedisTerraformTypescript
Agency • Marketing Tech • Software • Consulting
Lead and maintain performance, security, and reliability of client hosting environments across multi-cloud platforms. Architect resilient infrastructure, manage IaC and CI/CD, administer Windows/IIS and WP Engine environments, handle SSL/DNS/SSO, participate in on-call rotations, and engage with clients as senior escalation and trusted advisor.
Top Skills:
App ServicesApplication InsightsAWSAzureAzure DevopsAzure SqlCi/CdDnsEc2IamIisKey VaultsPowershellRdsS3Ssl/TlsSsoTerraformVulnerability ManagementWindows ServerWp Engine
Digital Media • Software
Deploy, operate, validate, and improve Ateme video delivery platforms in customer environments, including critical live production systems. Responsibilities include Linux administration, networking, video technology integration, troubleshooting, performance investigations, documentation, customer training, technical collaboration with R&D, and ensuring service availability. The role supports live production events and participates in a 24/7 on-call rotation.
Top Skills:
ArqAutomationAv1AvcCdnCloud InfrastructureCmafDistributed SystemsDrmHevcHlsJpeg-XsKubernetesLinuxMpeg-TsNetworkingObservabilityScte-104Scte-224Scte-30Scte-35Smpte 2110UnixVideo Encoding
eCommerce • Fintech • Information Technology • Payments • Financial Services
Manage, deploy, and architect highly scalable, redundant infrastructure across AWS and Kubernetes environments. Build CI/CD pipelines with GitHub Actions, automate infrastructure using Terraform, manage databases and document storage, and implement observability with monitoring and tracing tools. Lead technical design reviews, mentor teams on SRE principles, and improve software performance through infrastructure architecture.
Top Skills:
AWSBashDatadogDnsDockerDocument StorageDynatraceGitGithub ActionsKubernetesLinuxLoad BalancingNew RelicNode.jsPythonRdbmsRuby On RailsTerraformUnixVirtual Networking
Healthtech • Software
Provides technical leadership for site reliability and operational development. Designs automated, repeatable cloud and infrastructure solutions; monitors service objectives and business metrics; improves application resiliency, performance, efficiency, and cost. Builds monitoring, deployment, testing, and vulnerability-response automation while supporting mission-critical production systems. Collaborates with software, security, product, and business teams, conducts technical training and resilience exercises, and promotes sound development, change-management, and operational practices.
Top Skills:
Amazon EcsAnsibleAzureCC++ChefDockerGoJavaKubernetesLinuxPerlPuppetPythonRubyTerraformWindows
Artificial Intelligence • Cloud • Software
The Senior Site Reliability Engineer will automate operations, optimize workflows for teams, manage secure infrastructure, and participate in on-call duties.
Top Skills:
AristaAWSBashCephChefCifsCiscoDnsDockerElk StackFortinetHpHTTPIcmpIscsiJenkinsKubernetesLinux/Debian Family/UbuntuMesosphereNfsNode.jsPivotal GreenplumPostgresPythonRabbitMQRubyS3ScyllaSshSslSupermicroTcpTls
Hardware • Information Technology • Other • Software • Analytics
Serve as primary technical escalation for enterprise cloud customers, design and automate deployment pipelines, ensure Azure integrations, implement Terraform IaC, manage CI/CD with GitHub Actions/Azure DevOps, and use New Relic and AI-driven monitoring to improve reliability and performance.
Top Skills:
Azure DevopsBashCi/CdGithub ActionsAzureNew RelicPowershellTerraform
Artificial Intelligence • Machine Learning • Security • Database • Analytics • Big Data Analytics
As a Site Reliability Engineer, you'll ensure the availability and performance of AI applications, maintain infrastructure, automate tasks, and troubleshoot issues in high-scale environments.
Top Skills:
AnsibleAWSAzureBashCircleCICloudFormationDatadogDockerDynatraceEc2Elk StackGCPGitlab CiGoGrafanaJenkinsKubernetesLambdaLinuxPrometheusPythonS3TerraformUnix
Artificial Intelligence • Robotics • Automation • Manufacturing
Lead the Platform & SRE team to design and operate a unified deployment platform spanning cloud and on-premise. Architect Pulumi/Kubernetes-based deployments, package software for industrial hardware, build GitHub Actions CI/CD (including HIL), and define observability with Prometheus, Grafana, and OIDC.
Top Skills:
AWSAzureCertificate ManagementCnisGCPGithub ActionsGoGrafanaHelmHilKubernetesLinux NetworkingMulti-Cluster ManagementOidcPrometheusPulumiTerraform
Artificial Intelligence • Information Technology • Cybersecurity • Defense
The Forward Deployed Site Reliability Engineer ensures the reliability of a mission-critical platform, manages incident response, defines SLIs and SLOs, and liaises between engineering and government customers.
Top Skills:
AWSBashDockerGrafanaLokiMimirPrometheusPythonTerraform
Artificial Intelligence • Machine Learning • Biotech • Generative AI
The Site Reliability Engineer will manage digital infrastructure, ensuring access to compute resources, automating processes, and maintaining resource visibility for researchers.
Top Skills:
AnsibleDockerGrafanaKubernetesPrometheusPythonTailscaleTalos Linux
New
Cut your apply time in half.
Use ourAI Assistantto automatically fill your job applications.
Use For Free
Big Data
Lead and operate cloud infrastructure and SRE practices, focusing on IaC, CI/CD, observability, incident response, and production-grade agentic AI/LLM systems. Architect, build, and run scalable platforms, drive reliability and SLOs, author runbooks/ADRs, mentor engineers, participate in on-call rotations, and reduce toil via automation and platform tooling.
Top Skills:
Agent OrchestrationAgentic AiAWSAzureDockerFluxcdGCPGithub ActionsGitopsGoGrafanaJavaJenkinsKubernetesLinuxLlmLlm GatewayLokiOpentofuPrometheusPythonSentrySignozTcp/IpTerraform
Reposted 29 Days AgoSaved
Artificial Intelligence
The Deployment Engineer will build and operate AI inference clusters, ensure scalable deployments, optimize allocation, and maintain infrastructure. Responsibilities include software updates, telemetry development, and collaborative improvements with teams.
Top Skills:
DockerGrafanaInfluxdbK8SLinuxPrometheusPython
AdTech • Big Data • Internet of Things • Marketing Tech • Mobile • Software • Analytics
Senior Site Reliability Engineer responsible for improving platform availability, scalability, resilience, and observability. The role manages AWS and EKS infrastructure with Terraform, builds CI/CD pipelines using GitHub Actions, expands GitOps through ArgoCD and Helm, strengthens disaster recovery, supports incident response, and develops monitoring, alerting, dashboards, and SLOs. The engineer partners with product teams, leads infrastructure initiatives, troubleshoots Kubernetes, and supports high-availability distributed systems.
Top Skills:
Amazon EksAmazon RdsAmazon VpcArgocdAWSAws IamCluster AutoscalerCrossplaneDatadogGithub ActionsGoHelmKarpenterKubernetesOpentofuPrometheusPythonShell ScriptingTerraform
Cloud • Security • Software • Cybersecurity
Build and operate the Veeam Data Cloud GOV environment: map systems, write runbooks, define SLIs/SLOs, run incident response, close observability gaps, automate deployments and support fleet management while working across security and compliance constraints.
Top Skills:
Api ManagementApplication InsightsArgocdAws CloudformationAzureAzure Arm TemplatesAzure DevopsAzure FunctionsAzure GovernmentAzure MonitorAzure StorageBitbucketC#Cosmos DbDaggerElastic StackElkEntra IdFluxcdGitGithub ActionsGitlab CiGoGrafanaJavaJavaScriptKubernetesMicrosoft TfsOpentelemetryPrometheusPulumiServerless FrameworkTerraformTerragruntTypescript
Artificial Intelligence • Cloud • Software
The Senior Site Reliability Engineer will automate operations, improve workflows, manage secure infrastructure, and participate in on-call rotation for an AI-driven company.
Top Skills:
AristaAWSBashCephChefCifsCiscoDnsDockerElk StackFortinetHpHTTPIcmpIpIscsiJenkinsKubernetesLinux/DebianMesosphereNfsNode.jsPivotal GreenplumPostgresPythonRabbitMQRaidRubyS3ScyllaSshSslSupermicroTcpTlsUbuntu
Big Data • Healthtech • Information Technology • Analytics
Serve as an SRE-focused AI enablement partner: evaluate AI architectures, coach engineering teams on LLM integrations and agentic/RAG patterns, advise on observability, SLOs, governance, and operational readiness, and develop internal standards and documentation.
Top Skills:
Agentic FrameworksAnthropic ClaudeAWSAzureAzure Ai FoundryCi/CdDatabricksDatadogDockerGrafanaKubernetesLangchainLlamaindexLlm ApiOpentelemetryRagSemantic Kernel
Cloud
Build and operate reliable, scalable, secure infrastructure for SaaS security and Snowflake data systems. Automate infrastructure provisioning, deployments, and incident response using Terraform, Spinnaker, Kubernetes, and related tools. Ensure security and compliance, participate in on-call rotations, lead incident response and root-cause analysis, and collaborate with development, data science, and security teams on architectural decisions and service implementation.
Top Skills:
Artificial IntelligenceCi/CdFlywayInfrastructure As CodeKubernetesMachine LearningSnowflakeSpinnakerTerraform
Cloud
Designs, builds, and operates reliable, scalable infrastructure for security SaaS and Snowflake data systems. Automates infrastructure provisioning, deployments, and incident response using Terraform, Spinnaker, Kubernetes, and Flyway. Partners with security, development, and data science teams on secure, compliant architecture. Participates in on-call rotations, leads critical incident response and root-cause analysis, and implements preventative improvements. In-person onboarding and travel to the Toronto office are required during the first employment week.
Top Skills:
Ci/CdContainerizationFlywayInfrastructure As CodeKubernetesSnowflakeSpinnakerTerraform
Information Technology • Security • Cybersecurity
Own production operations and highly available cloud infrastructure for regulated government environments. Lead incident response, root-cause analysis, disaster recovery testing, compliance operationalization, audit readiness, vulnerability management, and continuous monitoring. Build automation, secure CI/CD pipelines, infrastructure-as-code, observability, and compliance tooling across Kubernetes, Linux, containers, and cloud platforms. Partner with security, compliance, and engineering teams to improve reliability, deployment safety, and regulatory sustainment.
Top Skills:
Aws GovcloudBashCi/CdDod Il4Dod Il5FedrampGitopsGoGrafanaKubernetesLinuxNist 800-53PrometheusPythonStigTerraformUnixZero Trust
Social Media
Operate, scale, and harden an AWS- and Kubernetes-based platform using GitOps. Build CI/CD and infrastructure-as-code (Terraform/Terragrunt), manage ArgoCD/Helm deployments, improve observability, automate toil reduction, lead incident response and post-incident remediation, and partner with application, security, and platform teams to improve reliability and delivery.
Top Skills:
ArgocdAWSBashContainersEksGithub ActionsGitopsHelmIamKubernetesLinuxPythonRbacTerraformTerragrunt
Cloud • Information Technology • Internet of Things • Software • Consulting • Infrastructure as a Service (IaaS) • Automation
Design, build, automate, and operate Red Hat Hybrid OpenShift platforms across cloud and on-prem. Implement GitOps, CI/CD, monitoring, and SRE practices; develop operators and tooling; lead incident response, on-call duties, and postmortems; mentor peers and improve platform reliability and self-service.
Top Skills:
ArgocdAWSAzureCatchpointCi/CdDatadogDnsFedoraGitlabGitopsGoGoGCPGrafanaHttp/TlsKubernetesKubernetes OperatorsLdapLinuxOpenshiftOpenshift PipelinesOperator SdkPrometheusPythonRhelSplunkSplunk ImTcp/IpTekton
Reposted One Month AgoSaved
Automotive • Cloud • Hardware • Software
Design, build, and operate developer platform infrastructure to support firmware build pipelines. Implement IaC patterns, GitOps, Kubernetes operators, CI optimization, cloud reliability, security policy-as-code, and developer tooling while mentoring engineers and collaborating on architecture and risk mitigation.
Top Skills:
ArgocdAWSAzureBashCiFluxGCPGitopsGoKubernetesPythonService MeshTerraform
Information Technology • Automation
The SRE/Infrastructure Engineer will architect and manage secure, scalable systems for automated penetration testing, optimizing reliability, and enhancing infrastructure based on customer demand. Responsibilities include maintaining production environments, leading technical discussions, and promoting high coding standards.
Top Skills:
AWSAzureCloudFormationElkGCPNew RelicOpentelemetryPostgresPrometheusTerraform
Let Your Resume Do The Work
Upload your resume to be matched with jobs you're a great fit for.
Success! We'll use this to further personalize your experience.
Popular Job Searches
All Software Engineer Jobs
.NET Developer Jobs
Aerospace Thermal Engineering Jobs
AI Engineer Jobs
Android Developer Jobs
Automation Engineer Jobs
Backend Developer Jobs
Blockchain Developer Jobs
C# Jobs
C++ Jobs
Cloud Architect Jobs
Cloud Engineer Jobs
Design Engineer Jobs
DevOps Engineer Jobs
Director Of Engineering Jobs
Electrical Engineering Jobs
Embedded Software Engineer Jobs
Engineering Jobs
Engineering Manager Jobs
Environmental Engineering Jobs
Field Engineer Jobs
Front End Developer Jobs
Full Stack Developer Jobs
Game Developer Jobs
Golang Jobs
Hardware Engineer Jobs
Industrial Engineering Jobs
iOS Developer Jobs
Java Developer Jobs
Javascript Developer Jobs
Linux Jobs
Manufacturing Engineer Jobs
Mechanical Engineering Jobs
Network Engineer Jobs
PHP Developer Jobs
Process Engineer Jobs
Project Engineer Jobs
Prompt Engineering Jobs
Python Jobs
QA Jobs
Robotics Engineer Jobs
Ruby on Rails Jobs
Salesforce Administrator Jobs
Salesforce Developer Jobs
Scala Jobs
Sharepoint Developer Jobs
Site Reliability Engineer Jobs
Software Engineering Manager Jobs
Solutions Architect Jobs
SQL Developer Jobs
Structural Engineer Jobs
System Engineer Jobs
Test Engineer Jobs
Web Developer Jobs
All Filters
Total selected ()
No Results
No Results





















.png)











