Top Site Reliability Engineer Jobs

One Month AgoSaved
In-Office
Annapolis Junction, MD, USA
112K-222K Annually
Senior level
112K-222K Annually
Senior level
Artificial Intelligence • Information Technology • Consulting • Cybersecurity
Support and develop data pipelines and telemetry capture for air-gapped cloud environments; collaborate with hardware, firmware, and data science teams; build queries and dashboards; create and standardize deployment solutions; provide operational support and learn new technologies as needed.
Top Skills: SparkAzureAzure Blob StorageAzure Data ExplorerAzure Data FactoryAzure SynapseC#GitJSONKusto Query Language (Kql)PowershellPublic Key Infrastructure (Pki)PythonSQL
One Month AgoSaved
In-Office
Washington, DC, USA
174K-238K Annually
Senior level
174K-238K Annually
Senior level
Cloud
Lead design, automation, and operation of large-scale cloud services for federal customers. Improve reliability via SRE practices (SLIs/SLOs, incident response), build automation and self-service platforms using Go/Python, Terraform, and Kubernetes, and ensure compliance with FedRAMP/IL6. Mentor teams, drive reliability initiatives, and modernize deployments with CI/CD and GitOps.
Top Skills: ArgocdAWSCassandraCi/CdDnsGCPGitopsGoHelmIamKubernetesLoad BalancingMySQLOpensearchPostgresPythonRedisSecrets ManagementTerraformTls
One Month AgoSaved
In-Office or Remote
Jupiter, FL, USA
Mid level
Mid level
Cloud • Other
Maintain and monitor Rocket.net hosting platform reliability and performance. Provide advanced escalation-level technical support for WordPress VIP customers, troubleshoot Linux-based production environments, web servers, databases, caching, DNS, and networking. Participate in incident response, root cause analysis, automation, documentation, and collaborate with support and engineering teams to improve platform stability.
Top Skills: ApacheBashCachingCdnCloudflareDatadogDnfDnsHTTPHttpsLinuxMariadbMySQLNedataNginxPhp-FpmRedisSshSsl/TlsWafWordpressYum
One Month AgoSaved
Remote
USA
75K-90K Annually
Mid level
75K-90K Annually
Mid level
3PL: Third Party Logistics
Own and improve uptime for backend services, APIs, workers, and ML pipelines; build monitoring, alerting, auto-remediation, optimize GCP infrastructure and Postgres performance, run on-call rotation and blameless postmortems.
Top Skills: Apollo ServerBashCi/CdCloud RunCloud SqlDatadogDockerExpressGCPGcsGrafanaNext.JsNode.jsPostgresPrometheusPythonReactTerraformTypescriptZabbix
Reposted One Month AgoSaved
In-Office
Wacker, IL, USA
132K-220K Annually
Expert/Leader
132K-220K Annually
Expert/Leader
Financial Services
The Staff Site Reliability Engineer will lead Platform Engineering's SRE efforts by defining technical strategy, overseeing architecture, and enhancing operational excellence through mentorship and governance.
Top Skills: ArgocdGCPGkeGoKafkaNode.jsPythonTerraform
One Month AgoSaved
In-Office
2 Locations
150K-170K Annually
Mid level
150K-170K Annually
Mid level
Artificial Intelligence • Logistics • Robotics • Software
Own reliability across cloud, edge, and on-site deployments. Build observability, monitoring, and alerting. Define incident response and on-call processes, improve deployment workflows, diagnose infra/network/distributed-system issues, and make deployments repeatable and scalable.
Top Skills: AWSAzureGCPGrafanaKafkaKubernetesLinuxOpentelemetryPrometheusRtspSecure TunnelsVpnWebrtc
One Month AgoSaved
In-Office
Reston, VA, USA
112K-222K Annually
Senior level
112K-222K Annually
Senior level
Artificial Intelligence • Information Technology • Consulting • Cybersecurity
Design, build, and support telemetry data pipelines and reporting for air-gapped Azure environments. Collaborate with hardware, firmware, and data teams to ingest, transform, query (KQL/SQL), and dashboard telemetry. Create standardized queries and deployment solutions, support air-gapped cloud reporting, and participate in on-call troubleshooting as needed.
Top Skills: SparkAzureAzure Blob StorageAzure Data ExplorerAzure Data FactoryAzure SynapseC#GitJSONKusto Query Language (Kql)PowershellPublic Key Infrastructure (Pki)PythonSQL
Reposted 10 Days AgoSaved
In-Office
Palo Alto, CA, USA
165K-280K Annually
Senior level
165K-280K Annually
Senior level
Aerospace • Other
Design, deploy, and operate highly available, sharded, geo-redundant distributed systems and multi-region infrastructure. Manage petabyte-scale bare-metal clusters, improve performance, and advance deployment, monitoring, and alerting. Collaborate across teams through full software lifecycle to build scalable, operable services for Starlink.
Top Skills: AlertingApache FlinkApache KafkaSparkBare Metal Compute ClustersC#Continuous IntegrationGoHbaseHdfsIstioJavaKubernetesLinuxMonitoringPythonScalaVersion Control
Reposted 10 Days AgoSaved
In-Office
Hawthorne, CA, USA
165K-265K Annually
Senior level
165K-265K Annually
Senior level
Aerospace • Other
Design, upgrade, and operate large-scale distributed systems for Starlink. Improve sharding, geo-redundancy, multi-region deployment, monitoring, and performance. Manage petabyte-scale bare-metal clusters and collaborate across teams through the full software lifecycle.
Top Skills: Apache KafkaSparkC#FlinkGoHbaseHdfsIstioJavaKubernetesLinuxPythonScala
Reposted 10 Days AgoSaved
In-Office
Redmond, WA, USA
165K-270K Annually
Senior level
165K-270K Annually
Senior level
Aerospace • Other
Design, deploy, and operate sharded, geo-redundant distributed systems and multi-region infrastructure. Manage petabyte-scale bare-metal clusters, improve deployment/monitoring/alerting, collaborate across teams, and optimize performance throughout the software lifecycle.
Top Skills: Apache KafkaSparkBare MetalC#FlinkGoHbaseHdfsIstioJavaKubernetesLinuxPythonScala
Reposted One Month AgoSaved
In-Office
El Segundo, CA, USA
170K-195K Annually
Mid level
170K-195K Annually
Mid level
Information Technology
Owner of production reliability across cloud and edge: define and drive SLIs/SLOs, build observability (Grafana/Prometheus/Loki/OpenTelemetry), participate in on-call and incident response, encode reliability in infrastructure-as-code (Terraform/OpenTofu), manage Kubernetes clusters, AWS hardening, HA databases, and oversee IoT/edge device fleet operations.
Top Skills: AWSGitGrafanaIamIotKubernetesLokiNebulaNvidia JetsonOpentelemetryOpentofuPrometheusPyrraSlothStatefulsetsTailscaleTerraformWireguard
Reposted One Month AgoSaved
In-Office or Remote
17 Locations
Senior level
Senior level
Fintech • Information Technology • Software • Financial Services
Design, build, and maintain real-time, secure distributed systems and observability UIs/APIs. Implement CI/CD, containerized deployments (Docker/Kubernetes/OpenShift), integrate observability stack (Elasticsearch/Logstash/Grafana), and apply secure coding and API security standards to ensure reliability, performance, and incident automation. Collaborate in Agile teams and explore AI to improve resiliency.
Top Skills: Agentic AiCi/CdDockerElasticsearchGrafanaJava Spring BootKafkaKubernetesLogstashMariadbNode.jsOauth2OpenshiftReactSecrets Management
New

Cut your apply time in half.

Use ourAI Assistantto automatically fill your job applications.

Use For Free
Application Tracker Preview
Reposted One Month AgoSaved
In-Office
Alpharetta, GA, USA
Senior level
Senior level
Information Technology • Consulting
Design, build, customize, and support Oracle E-Business Suite and related financial applications. Gather requirements, implement EBS customizations and integrations, optimize performance and SQL, perform testing and upgrades, and provide tier-3 support and issue resolution with Oracle Support.
Top Skills: Invoice AutomationOracle CloudOracle E-Business SuiteOracle Ebs ApisOracle FinancialsOracle Integration CloudOracle Supplier Portal CloudSQLWebcenter Content Imaging
Reposted One Month AgoSaved
In-Office
Wacker, IL, USA
101K-168K Annually
Senior level
101K-168K Annually
Senior level
Financial Services
The Site Reliability Engineer III designs secure, scalable technology solutions, ensures operational resiliency, and collaborates with teams to maintain high availability across environments.
Top Skills: AutomicAWSAzureBambooBigQueryDockerGitGoogle Cloud PlatformGrafanaJavaJIRAKubernetesLinuxOpentelemetryOraclePostgresPrometheusPythonSplunkUc4Unix
Reposted One Month AgoSaved
In-Office
Saratoga, CA, USA
100K-165K Annually
Senior level
100K-165K Annually
Senior level
Other
As a Platform Engineer/Dev Ops, you will expand cloud infrastructure, implement monitoring systems, manage databases, and leverage CI/CD tools, working collaboratively with various teams.
Top Skills: AWSAzureBashDatadogElk StackKubernetesOpentofuPrometheusPythonTerraform
Reposted One Month AgoSaved
In-Office
Cambridge, MA, USA
160K-205K Annually
Senior level
160K-205K Annually
Senior level
Software
Design, build, and operate multi-account cloud infrastructure using IaC. Automate customer deployments, manage CI/CD, troubleshoot production across infra/data/app layers, and handle networking, security, and compliance for regulated environments while collaborating with platform and professional services teams.
Top Skills: AirflowAuth0AWSAzureDbtDockerEcsGCPGithub ActionsLlmsOktaPackerPostgresSnowflakeTailscaleTerraformWireguard
Reposted One Month AgoSaved
In-Office
San Francisco, CA, USA
181K-263K Annually
Senior level
181K-263K Annually
Senior level
Big Data • Cloud • Marketing Tech • Social Impact • Software
The Senior Staff Site Reliability Engineer at LiveRamp will define the SRE strategy, oversee critical automation, and lead operational excellence in a global infrastructure, influencing architectural decisions and mentoring teams.
Top Skills: Aws)CassandraCircleCICloud Security (GcpDynamoDBGoJenkinsKubernetesPythonScylladbSinglestoreTerraform
Reposted 11 Days AgoSaved
In-Office
2 Locations
Senior level
Senior level
Energy
The Senior Site Reliability Engineer improves infrastructure reliability and scalability, partners with various teams, implements IaC and CI/CD, and ensures business continuity through effective BCP/DR planning.
Top Skills: AWSBashCloudFormationDatadogElkGithub ActionsGitlab CiGoGrafanaJenkinsKubernetesOpensearchPrometheusPythonTerraform
11 Days AgoSaved
In-Office or Remote
2 Locations
147K-264K Annually
Senior level
147K-264K Annually
Senior level
Cloud • Security • Software • Cybersecurity
Deploy and operate scalable, highly available cloud systems; improve application and network security, reliability, speed, and capacity; automate cloud deployments; monitor and troubleshoot services to meet SLAs; analyze logs and events; and resolve infrastructure issues across large-scale distributed systems.
Top Skills: Cloud ComputingDistributed SystemsHTTPJavaLinux/UnixPerlPythonTcp/IpTls/Ssl
Reposted 11 Days AgoSaved
In-Office
Seattle, WA, USA
Senior level
Senior level
Other
The Sr. Site Reliability Engineer will maintain and administer enterprise systems, troubleshoot operational issues, and develop scripts. This role requires collaboration across teams and participation in project planning and execution.
Top Skills: AnsibleApacheAzureC#ChefIisJavaJbossPerlPowershellPuppetPythonRubyTomcat
Reposted 11 Days AgoSaved
In-Office
Santa Clara, CA, USA
140K-180K Annually
Senior level
140K-180K Annually
Senior level
Software
Lead the modernization of AWS cloud infrastructure, implement automation, ensure system reliability, and manage performance with a focus on security and incident response.
Top Skills: AngularApexAWSC#ElasticacheNew RelicNode.jsNpmPm2PythonRedisShell ScriptingTerraform
Reposted 11 Days AgoSaved
In-Office
Austin, TX, USA
Senior level
Senior level
Financial Services
The Senior Site Reliability Engineer will own the operational reliability of developer tooling ecosystems and improve developer productivity through efficient processes and automation.
Top Skills: .NetBashPowershellPython
Reposted 11 Days AgoSaved
In-Office
Washington, DC, USA
Senior level
Senior level
Big Data • Analytics • Business Intelligence • Big Data Analytics
Seeking a Site Reliability Engineer to manage AI platform reliability, automate tasks, optimize ML pipelines, and lead incident response in a hybrid engineering role.
Top Skills: ArgocdBigQueryCloud BuildDockerDvcGithub ActionsGoGrafanaKubeflowKubernetesMlflowPrometheusPub/SubPythonTerraformVertex Ai
Reposted 11 Days AgoSaved
In-Office
New York, NY, USA
141K-217K Annually
Senior level
141K-217K Annually
Senior level
Artificial Intelligence • Cloud • Social Impact • Software • Wearables
Own reliability, observability, and operational excellence for the Unified Call (911) platform. Design monitoring, alerting, incident response, deployment automation, and dashboards. Improve system resiliency, analyze architecture for operational risks, and build tooling and practices to enable reliable production operations across a Kubernetes-based cloud environment.
Top Skills: AWSDatadogKafkaKubernetesRabbitMQ
Reposted 12 Days AgoSaved
In-Office
Reston, VA, USA
137K-244K Annually
Senior level
137K-244K Annually
Senior level
Cloud • Fintech • HR Tech
Responsible for designing, building, automating, and maintaining a secure, highly available Kubernetes-based analytics platform for federal deployments. Improve CI/CD, observability, provisioning (Terraform, Argo CD), troubleshooting, on-call incident response, and collaborate across teams to deliver Workday Prism Analytics in GovCloud.
Top Skills: SparkArgo CdAWSCi/CdDockerGoGovcloudGrafanaKubernetesObservabilityPrism AnalyticsPrometheusPythonTerraformTracing
All Filters
JobType
New Jobs
Job Category
Experience
Industry
Company Name
Company Size

Sign up now Access later

Create Free Account