Maximum of 25 job preferences reached.
Top Site Reliability Engineer Jobs
News + Entertainment
As a Senior Machine Learning Engineer, you will develop advanced machine learning and deep learning models and platforms for optimizing advertising performance and conduct complex experiments.
Top Skills:
AIControl SystemsDeep LearningMachine LearningReinforcement LearningStatistical Techniques
News + Entertainment
Design, operate, and scale cloud-native ML infrastructure across GCP and AWS (GPU/TPU), build CI/CD for models, maintain low-latency real-time inference systems, define observability and monitoring for ML models, participate in on-call incident response, and partner with data scientists to improve MLOps and platform usability.
Top Skills:
AerospikeApache AirflowApache FlinkSparkAWSChrononDatadogEksGCPGitlab RunnerGkeGpuGrafanaJavaJenkinsKafkaKubernetesKv StoreMlflowPrometheusPythonRayScalaTerraformTpuVector Database
Aerospace • Manufacturing
As a Site Reliability Engineer, you'll build and manage observability platforms for satellite communications, define SLOs/SLIs, and collaborate on incident response and deployment automation.
Top Skills:
ArgocdAWSElkGCPGoGrafanaIstioJaegerKubernetesLinkerdLokiOpentelemetryPrometheusPythonTempoTerraform
Other
Design, deploy, and operate a high-performance SaaS manufacturing platform on cloud infrastructure. Ensure scalability, reliability, automation, monitoring, and KPIs. Integrate AI into dev/ops workflows, lead impactful projects, and improve uptime, latency, and efficiency. Work cross-functionally in a fast-paced startup while meeting U.S. government data access citizenship requirements.
Top Skills:
Ai ToolsApm ToolsAWSContainerizationKubernetesLinuxLogging ToolsMonitoring ToolsShellTerraform
Aerospace • Other
Design, deploy, and scale on-prem Kubernetes clusters and core infrastructure for Starlink. Build automation, manage databases, monitoring, and distributed storage. Collaborate with engineers to improve service lifecycle, availability, and performance; troubleshoot across the Starlink stack and drive reliability improvements.
Top Skills:
AnsibleBashBazelC++GoKubernetesLinuxMakefilesOci ContainersPythonTcp/IpTerraform
Aerospace • Other
Design, deploy, and scale on‑premise compute and core infrastructure for Starlink. Develop automation, manage databases/monitoring/distributed storage, collaborate with software teams, troubleshoot end-to-end, and improve deployment and developer velocity.
Top Skills:
AnsibleBashCC++DatabasesDistributed StorageDockerGoHypervisor TechnologiesKubernetesLinuxMonitoringPythonTcp/IpTerraformVirtualization
Cybersecurity
Ensure reliability, scalability, observability, and cost efficiency of a customer-facing SaaS security platform. Manage Kubernetes/Helm deployments, CI/CD (GitLab/ArgoCD), monitoring, and service verification. Embed with engineering teams, optimize developer CI/CD workflows, monitor and debug production on AWS/GCP, and participate in a 24/7 on-call rotation.
Top Skills:
ArgocdAWSGCPGitlab Ci/CdGrafanaHelmKubernetesMicroservicesPrometheus
Fintech • Payments • Financial Services
Operate and maintain AWS infrastructure and CI/CD pipelines, deploy containerized applications, write Terraform IaC, implement monitoring/observability, troubleshoot incidents, resolve security vulnerabilities, participate in on-call rotation, and produce operational/runbook documentation to ensure resilient cloud platform delivery.
Top Skills:
AgileAlbAmazon Web Services (Aws)Aws IamAws LambdaCloudbeaverCloudwatchDevsecopsDirect ConnectDnsEc2EcsEfsEksElbGitlab CiGlueGrafanaHelmJavaJbossJIRANginxOktaOracle RdsPythonRds PostgresRoute53S3ScalaSnsTerraformTomcatTransit GatewayVpc
Artificial Intelligence • Software • Generative AI
The Lead Site Reliability Engineer will drive technical strategy, ensure high service availability, manage cloud infrastructure, and lead a team to optimize systems and automate processes.
Top Skills:
AWSAzureDockerGoogle Cloud PlatformKubernetesTerraform
Cybersecurity
Assist with monitoring system performance, uptime, and reliability; support incident response, troubleshooting, and root cause analysis; monitor SRE alerts in Slack and report issues; learn cloud platforms; and collaborate with development and DevOps teams.
Top Skills:
AWSAzureDevOpsDevsecopsGCPGrafanaKubernetesPrometheusSlack
Reposted 15 Days AgoSaved
Cloud • Software
Responsible for maintaining FedRAMP-compliant infrastructure, collaborating with software engineers, and ensuring system availability and security. Duties include infrastructure design, automation, monitoring, and incident response.
Top Skills:
AWSGoKubernetesPuppetPythonTerraform
Artificial Intelligence • Information Technology • Software • Consulting
The Senior Site Reliability Engineer will improve the reliability, scalability, performance, and resilience of enterprise systems. Responsibilities include developing automation and self-healing solutions, reducing operational toil, managing cloud infrastructure, Kubernetes, containers, and microservices, and implementing observability through monitoring, logging, alerting, and distributed tracing. The role also handles incident management, root cause analysis, performance tuning, and capacity planning while collaborating across teams. This is a hybrid position requiring 50% onsite work in Albany or Manhattan, with the first week onsite in Albany.
Top Skills:
AWSAzureContainersDynatraceElastic StackElkGCPKubernetesMicroservicesPythonSplunk
New
Cut your apply time in half.
Use ourAI Assistantto automatically fill your job applications.
Use For Free
Blockchain • eCommerce • Fintech • Payments • Software • Financial Services • Cryptocurrency
As a Senior Site Reliability Engineer, you will enhance platform reliability, lead incident management, and drive AI-driven improvements in operational workflows.
Top Skills:
Amazon Web ServicesDatadogDynamoDBEnvoyEvent Driven ArchitecturesGrpcHTTPIstioJSONKotlinKubernetesLaunchdarklyModern JavaMySQLProtocol BuffersTerraformVitess
Blockchain • eCommerce • Fintech • Payments • Software • Financial Services • Cryptocurrency
The Senior Site Reliability Engineer will enhance reliability of Block's platform, improve incident response using AI tools, and coordinate incident management. Responsibilities include building reliable systems, standardizing tools, and leading high-severity incidents during on-call rotations.
Top Skills:
Amazon Web ServicesDatadogDynamoDBGrpcHTTPIstioJavaJSONKotlinKubernetesLaunchdarklyMySQLProtocol BuffersTerraformVitess
Enterprise Web • Hardware • Internet of Things • Software
Lead observability and reliability efforts: mentor teams on SLIs/SLOs, maintain triage/remediation workflows, perform incident response, debug production systems, and design core infrastructure and tooling for engineering teams.
Top Skills:
AlloyClaude SkillsGemini GemsGoGrafanaKubernetesLokiMimirMongoDBOpentelemetryPostgresPrometheusPromqlTempoTypescript
Fintech • Software
Lead SRE efforts for DFIN SaaS: ensure availability, performance, scalability, and automation. Implement monitoring, CI/CD, IaC, container orchestration, AI-enhanced observability, incident response, RCA, and runbook automation while collaborating across engineering teams.
Top Skills:
.NetAiopsAksAnsibleAppdynamicsAWSAzureAzure DevopsBashC#Ci/CdCloud Ai ServicesContainersCosmosDatadogDynatraceEksFirewallHarnessIdera Sql Diagnostic ManagerInfrastructure As Code (Iac)JavaJenkinsKubernetesLinuxLoad BalancingNew RelicPowershellPythonRedgate Sql MonitorSolarwinds Database Performance AnalyzerSQLTerraformWindows
Reposted 10 Days AgoSaved
Automotive • Cloud • Hardware • Software
Lead and operate ProdOps as senior SRE: coordinate and communicate during high-severity incidents, run blameless post-incident reviews, prioritize and drive systemic action items, verify fixes, build automation and observability (including AI-agent workflows), and grow a lean team while mentoring engineers and setting technical standards.
Top Skills:
Ai AgentsCloudDatadogGoInstrumentationLlmsLoggingMetricsObservabilityPythonSliSloTracing
Cloud • Fintech • HR Tech
Design, implement, test, deploy, and maintain automation and configuration management for containerized services. Automate deployments and scaling on Kubernetes, build self-service platforms for developers, monitor and triage production incidents, run retrospectives, and participate in infrequent on-call rotations. Collaborate across teams, use observability tools to debug, mentor engineers, and drive continuous improvement and operational efficiency.
Top Skills:
AnsibleArgocdAWSBashGCPGoGrafanaJenkinsKubernetesPrometheusPythonRubySplunk
Fintech • Software • Financial Services
Own AWS and Kubernetes reliability, deployments, observability, incident response, infrastructure as code, CI/CD, and compliance controls for SOC 2 and PCI-DSS. Build scalable platform foundations, AI infrastructure, and greenfield product infrastructure while improving cost efficiency, security, vendor flexibility, and deployment consistency. Participate in on-call support and partner with security and engineering teams across backend, mobile, data, ML, and AI.
Top Skills:
Ai/Llm ToolingAmazon EksAWSAws CdkCi/CdDatadogKotlinKubernetesMcpNode.jsOpentelemetryOpentofuPci-DssPythonSoc 2TerraformTypescript
Cloud • Security • Software • Cybersecurity
Deploy and operate scalable, highly available cloud systems; improve application and network security, reliability, speed, and capacity; automate cloud deployments; monitor and troubleshoot services to meet SLAs; analyze logs and events; and resolve infrastructure issues across large-scale cloud infrastructure.
Top Skills:
AnsibleAWSBashBitbucketDatabricksDatadogDigicertGradleJenkinsNode.jsPythonSplunkTenableTerraformThreatmetrix
Artificial Intelligence • Information Technology
Lead the reliability, stability, observability, and scalability of Seekr’s software platform. Design and implement systems, networks, and services that meet SLAs; develop monitoring, automation, and failure-detection solutions; conduct load testing; manage observability tools; troubleshoot production and deployment issues; and support incident response. Collaborate closely with software engineering teams across hybrid cloud and on-premises environments.
Top Skills:
AerospikeAnsibleArgocdBashChefDockerElasticsearchElkGitGitlabGrafanaInfluxdbJavaKafkaKubernetesLinuxPrometheusPuppetPythonRubyTerraform
Artificial Intelligence • Fintech • Machine Learning • Natural Language Processing • Business Intelligence
Lead architecture and implementation of reliability platforms and SRE practices for a production SaaS. Build self-service reliability tooling, drive AIOps automation, advance observability (monitoring, tracing, profiling), lead incident response and postmortems, mentor engineers, and embed production readiness across teams to achieve 99.99% uptime.
Top Skills:
AWSAzureContinuous ProfilingDatadogDnsElkGCPGoGrafanaHttp/SKubernetesLoad BalancingOpentelemetryPrometheusPythonTcp/Ip
11 Days AgoSaved
Big Data • Fintech • Mobile • Payments • Financial Services • Data Privacy
Leads reliability and operational excellence for an enterprise Kubernetes container platform. Responsibilities include incident response, root cause analysis, L3 support, upgrades, observability, automation, capacity planning, performance tuning, security controls, compliance, resilience testing, and runbook development. The role partners with engineering, architecture, security, infrastructure, product, and operations teams while mentoring SRE resources and improving platform adoption, developer experience, and production reliability.
Top Skills:
AnsibleArgocdBashBitbucketCi/CdDnsDynatraceElkGitGitlabGitopsGoGrafanaHelmInfrastructure As CodeJenkinsKubernetesLinuxLoad BalancingNetwork PoliciesOpenshiftOpenshift VirtualizationOpentelemetryPrometheusPythonRancherRbacService MeshSplunkTerraformVcfVksVMware
Artificial Intelligence • Fintech • Machine Learning • Software • App development • Conversational AI • Generative AI
Own and improve production infrastructure reliability, deployments, Infrastructure-as-Code, Kubernetes environments, automation, CI/CD, monitoring, alerting, and observability. Investigate incidents, optimize system performance, maintain documentation and runbooks, and support DNS, WAF, CDN, and caching infrastructure. The role requires strong Linux administration, Bash scripting, networking, Git, and containerization skills, with independent ownership and collaboration across development and operations teams.
Top Skills:
AkamaiAmqpAnsibleAWSBashCdnCloudflareDnsDockerGCPGitGitlab CiGrafanaHttp/HttpsKubernetesLinuxPodmanPrometheusPythonRabbitMQTerraformVictoriametricsWafZabbix
Manufacturing
The Site Reliability Engineer will ensure the reliability, security, and support of Databricks applications while collaborating with various teams to optimize data workflows and incident management.
Top Skills:
AzureCi/CdDatabricksDelta LakePysparkPythonSQLUnity Catalog
Let Your Resume Do The Work
Upload your resume to be matched with jobs you're a great fit for.
Success! We'll use this to further personalize your experience.
Top Companies Hiring Site Reliability Engineers
See AllPopular Job Searches
All Software Engineer Jobs
.NET Developer Jobs
Aerospace Thermal Engineering Jobs
AI Engineer Jobs
Android Developer Jobs
Automation Engineer Jobs
Backend Developer Jobs
Blockchain Developer Jobs
C# Jobs
C++ Jobs
Cloud Architect Jobs
Cloud Engineer Jobs
Design Engineer Jobs
DevOps Engineer Jobs
Director Of Engineering Jobs
Electrical Engineering Jobs
Embedded Software Engineer Jobs
Engineering Jobs
Engineering Manager Jobs
Environmental Engineering Jobs
Field Engineer Jobs
Front End Developer Jobs
Full Stack Developer Jobs
Game Developer Jobs
Golang Jobs
Hardware Engineer Jobs
Industrial Engineering Jobs
iOS Developer Jobs
Java Developer Jobs
Javascript Developer Jobs
Linux Jobs
Manufacturing Engineer Jobs
Mechanical Engineering Jobs
Network Engineer Jobs
PHP Developer Jobs
Process Engineer Jobs
Project Engineer Jobs
Prompt Engineering Jobs
Python Jobs
QA Jobs
Robotics Engineer Jobs
Ruby on Rails Jobs
Salesforce Administrator Jobs
Salesforce Developer Jobs
Scala Jobs
Sharepoint Developer Jobs
Site Reliability Engineer Jobs
Software Engineering Manager Jobs
Solutions Architect Jobs
SQL Developer Jobs
Structural Engineer Jobs
System Engineer Jobs
Test Engineer Jobs
Web Developer Jobs
All Filters
Total selected ()
No Results
No Results




























