Top Site Reliability Engineer Jobs

Reposted 11 Days AgoSaved
In-Office
Golden, CO, USA
103K-136K Annually
Senior level
103K-136K Annually
Senior level
Manufacturing
The Site Reliability Engineer will ensure the reliability, security, and support of Databricks applications while collaborating with various teams to optimize data workflows and incident management.
Top Skills: AzureCi/CdDatabricksDelta LakePysparkPythonSQLUnity Catalog
Reposted 11 Days AgoSaved
Remote
United States
Senior level
Senior level
Automotive
Design and implement scalable cloud infrastructure, monitor performance, automate processes, ensure security and compliance, and lead a DevOps team.
Top Skills: AWSBashCi/CdDockerElk StackGCPGrafanaKubernetesPrometheusPythonTerraform
Reposted 11 Days AgoSaved
Hybrid
San Francisco, CA, USA
200K-240K Annually
Mid level
200K-240K Annually
Mid level
Artificial Intelligence • Logistics • Software
The Site Reliability Engineer will enhance operational resilience, ensuring system stability, observability, and debugging workflows for complex failures while improving developer focus and uptime.
Top Skills: DatadogGoPrometheusPythonSentry
Reposted 11 Days AgoSaved
In-Office
Arlington, VA, USA
192K-455K Annually
Senior level
192K-455K Annually
Senior level
Artificial Intelligence • Information Technology • Cybersecurity • Defense
As a Site Reliability Engineer, you'll ensure system reliability in a government environment, manage incidents, and collaborate with engineering teams on operational tasks and improvements while maintaining security compliance.
Top Skills: AWSBashDockerDocker ComposeGrafanaLinux/UnixLokiMimirPrometheusPythonTerraform
Reposted 11 Days AgoSaved
In-Office or Remote
4 Locations
110K-120K Annually
Mid level
110K-120K Annually
Mid level
Fintech • Software
Operate and improve the reliability, availability, and security of a FedRAMP High cloud platform as an SRE/L3 escalation point. Monitor systems, troubleshoot incidents, lead incident response and RCA, build runbooks, enhance observability and automation, support deployments, and ensure compliance while partnering with engineering teams to reduce toil and improve resilience.
Top Skills: AWSAws BackupBashCi/CdCloudwatchConfiguration ManagementDatadogEksFedramp HighGoGrafanaIamInfrastructure As CodeIstioJira Service ManagementKubernetesLinuxOpentelemetryPagerdutyPowershellPrometheusPythonRdsRoute 53SplunkVpc
Reposted 11 Days AgoSaved
In-Office
San Francisco, CA, USA
230K-390K Annually
Senior level
230K-390K Annually
Senior level
Artificial Intelligence • Software
As a Software Engineer on the Site Reliability team, you'll ensure system reliability, scalability, and observability while partnering with engineering teams and improving incident management processes.
Top Skills: AWSCi/Cd ToolingContainer OrchestrationDatadogGrafanaPrometheusTerraform
Reposted 11 Days AgoSaved
Hybrid
Arlington, TX, USA
Mid level
Mid level
Fintech • Financial Services
The Site Reliability Engineer I will support cloud infrastructure and assist in cloud transformation initiatives, focusing on performance and delivery of public cloud solutions, primarily in Azure. Responsibilities include troubleshooting, monitoring, automation, and contributing to operational readiness practices for cloud services.
Top Skills: .NetAnsibleAWSAzureAzure CliGCPJenkinsKubernetesLinuxPowershellTerraformWindows
Reposted 11 Days AgoSaved
Hybrid
2 Locations
Senior level
Senior level
Fintech • Financial Services
The role involves shaping release engineering practices, implementing AI-driven solutions, and ensuring software reliability through collaboration and automation.
Top Skills: Ai-Powered ToolsAzureBashC#Github CopilotJavaPowershell
Reposted 11 Days AgoSaved
In-Office
Annapolis Junction, MD, USA
165K-230K Annually
Expert/Leader
165K-230K Annually
Expert/Leader
Information Technology • Software • Automation
The Senior Site Reliability Engineer will manage AWS environments, develop Infrastructure as Code, and automate operational tasks to ensure high availability in cloud systems.
Top Skills: Amazon Web Services (Aws)AnsibleAws Certified Developer-AssociateAws Certified Solutions Architect-AssociateAws Certified Solutions Architect-ProfessionalAws Certified Sysops Administrator-AssociateCertified Kubernetes Administrator (Ckad)Ci/CdDockerElastic Certified EngineerElastic Certified Observability EngineerKubernetesTerraform
Reposted 11 Days AgoSaved
In-Office
Seattle, WA, USA
Senior level
Senior level
Other
In this role, you will manage day-to-day operations of Internet-based enterprise systems, identify operational issues, develop tools for maintenance, and collaborate on infrastructure documentation and project execution.
Top Skills: .NetAnsibleApacheAzureChefIisJbossPerlPowershellPuppetPythonRubyTomcat
Reposted 11 Days AgoSaved
In-Office
Seattle, WA, USA
Senior level
Senior level
Other
Responsible for monitoring, provisioning, and customer interactions, with a focus on maintaining high availability in complex web environments.
Top Skills: .NetAnsibleApacheCfengineChefDyanatraceGoIisJavaJbossNasNew RelicPerlPowershellPuppetPythonRaidRubySanSplunkSumo LogicTomcatWindows
Reposted 11 Days AgoSaved
In-Office
New York, NY, USA
220K-260K Annually
Expert/Leader
220K-260K Annually
Expert/Leader
Payments • Software • Automation
Lead platform and infrastructure direction on AWS, evolve CI/CD and ephemeral environments, set observability and SLO standards, drive incident response and postmortems, mentor engineers, and build automation to reduce operational risk.
Top Skills: AWSCi/CdDistributed SystemsEcsEphemeral Environments/Preview DeploysFargateGithub ActionsLogsObservability (MetricsSlos/Slis/Error BudgetsTracing)
New

Cut your apply time in half.

Use ourAI Assistantto automatically fill your job applications.

Use For Free
Application Tracker Preview
Reposted 11 Days AgoSaved
In-Office
Plantation, FL, USA
Mid level
Mid level
Healthtech
Respond to production incidents, monitor systems with Datadog, troubleshoot Kubernetes/Docker environments, support CI/CD and automation, participate in on-call rotations, and collaborate with engineering teams to maintain availability and compliance in regulated healthcare/fintech systems.
Top Skills: AWSAzureBashC#Ci/CdDatadogDockerGCPHelmJavaKubernetesMySQLNoSQLPowershellPythonSQL
Reposted 11 Days AgoSaved
In-Office
Redmond, WA, USA
125K-175K Annually
Junior
125K-175K Annually
Junior
Aerospace • Other
Design, operate, scale, and automate HPC clusters and services for silicon design workflows. Manage infrastructure-as-code, CI/CD pipelines, observability, and storage automation. Collaborate with cross-functional teams to eliminate performance bottlenecks and accelerate simulation and regression turnaround times.
Top Skills: AnsibleAnsysBambooBashCadenceClaude CodeDockerGrafanaGrokJenkinsKeysightKubernetesLinuxLsfMySQLNetapp OntapNfsPostgresPrometheusPuppetPythonRest ApiSiemensSlurmSqliteSynopsysTcp/IpTerraform
Reposted 11 Days AgoSaved
In-Office
San Francisco, CA, USA
238K-290K Annually
Expert/Leader
238K-290K Annually
Expert/Leader
Artificial Intelligence • Legal Tech • Professional Services • Software
As a Staff Software Engineer in Site Reliability, you'll manage infrastructure for reliability and scalability, lead incident management, and automate operational tasks.
Top Skills: AWSAzureBashCloudFormationDatadogGCPGoIncidentioPagerdutyPulumiPythonSentryTerraform
Reposted 11 Days AgoSaved
In-Office or Remote
7 Locations
200K-200K Annually
Mid level
200K-200K Annually
Mid level
Cloud • Software
The Site Reliability Engineer will ensure reliable cloud operations by applying Python for infrastructure automation, managing OpenStack and Kubernetes, and practicing devsecops in a fast-paced environment.
Top Skills: KubernetesLinuxOpenstackPython
Reposted 11 Days AgoSaved
In-Office
San Francisco, CA, USA
200K-260K Annually
Mid level
200K-260K Annually
Mid level
Artificial Intelligence • Legal Tech • Professional Services • Software
As a Software Engineer in Site Reliability, you will ensure the reliability and performance of our AI platform through automation and strategic infrastructure management.
Top Skills: AWSAzureBashCloudFormationDatadogGCPGoKubernetesPagerdutyPythonSentryTerraform
Reposted 11 Days AgoSaved
In-Office
San Francisco, CA, USA
Mid level
Mid level
Information Technology • Software • Big Data Analytics
The Site Reliability Engineer will design, analyze, and troubleshoot large-scale distributed systems, focusing on operating systems and performance tuning.
Top Skills: ApacheJava
Reposted 11 Days AgoSaved
In-Office or Remote
2 Locations
95K-171K Annually
Junior
95K-171K Annually
Junior
Cloud • Security • Software • Cybersecurity
As a Site Reliability Engineer II, you'll automate tasks, monitor AI workloads, enhance dashboards, support CI/CD processes, and collaborate with engineering teams on complex issues while participating in on-call rotations.
Top Skills: GoGrafanaKubernetesLinuxPrometheusPythonSaltstackTerraform
Reposted 11 Days AgoSaved
In-Office
2 Locations
Mid level
Mid level
Artificial Intelligence
The Deployment Engineer will manage AI inference clusters, optimizing deployment, capacity allocation, and ensuring reliability of pipeline operations across datacenters.
Top Skills: DockerGrafanaInfluxdbK8SLinuxPrometheusPython
Reposted 11 Days AgoSaved
In-Office
Pittsburgh, PA, USA
146K-162K Annually
Senior level
146K-162K Annually
Senior level
Financial Services
The Lead Site Reliability Engineer will establish the SRE operating model, implement AI-enabled reliability use cases, manage reliability metrics, and oversee operational readiness while collaborating with teams and mentoring engineers.
Top Skills: Ai/MlAnsibleAzure DevopsDockerGithub ActionsGitlab CiJenkinsKubernetesTerraformVMware
Reposted 11 Days AgoSaved
In-Office
Reston, VA, USA
136K-184K Annually
Senior level
136K-184K Annually
Senior level
Information Technology • Software
The Systems Engineer manages Linux systems, designs CI/CD pipelines, administers application security platforms, and ensures compliance with security standards.
Top Skills: AnsibleBashCloudbees JenkinsDockerElkGitGithub ActionsGithub Advanced SecurityJfrog ArtifactoryJfrog XrayJIRAKubernetesLinuxNagiosNexus IqPrometheusPythonTerraform
Reposted 11 Days AgoSaved
In-Office or Remote
2 Locations
165K-215K Annually
Senior level
165K-215K Annually
Senior level
Software • Cybersecurity
This role involves managing Kubernetes clusters, cloud infrastructure, and CI/CD pipelines. The engineer will enhance system reliability and efficiency while troubleshooting production issues.
Top Skills: AlertmanagerAWSAzureBashCi/CdDockerElastic StackElasticsearchGCPGoGrafanaHelmKafkaKubernetesLokiMongoDBOciPrometheusPythonRedisSparkTerraform
Reposted 11 Days AgoSaved
In-Office
Owings Mills, MD, USA
159K-339K Annually
Senior level
159K-339K Annually
Senior level
Financial Services
As a Principal Site Reliability Engineer, you'll lead a team focusing on observability and automating solutions for cloud and on-prem infrastructures, enhancing reliability and incident response across T. Rowe Price's tech ecosystem.
Top Skills: .Net CoreAmazon AwsAnsibleElastic StackGoGrafanaJavaMySQLNew RelicNode.jsPostgresPrometheusPythonSolarwinds DpaSplunkSQL ServerTerraformVagrantVault
Reposted 11 Days AgoSaved
In-Office
Cambridge, MA, USA
Mid level
Mid level
Cloud • Information Technology • Biotech
The Site Reliability Engineer will build and deploy Linux servers, research technologies, monitor system performance, and resolve technical incidents.
Top Skills: Infrastructure-As-CodeLinuxNetworkingVirtualization
All Filters
JobType
New Jobs
Job Category
Experience
Industry
Company Name
Company Size

Sign up now Access later

Create Free Account