Top Site Reliability Engineer Jobs

Reposted 15 Days AgoSaved
Hybrid
Palo Alto, CA, USA
Senior level
Senior level
Financial Services
Lead design and implementation of SRE practices, observability, reliability and AI-assisted operational workflows. Mentor engineers, define NFRs/SLOs, build automation, logging/metrics/tracing pipelines, containerized CI/CD/GitOps, and integrate AI agents for production-grade reliability.
Top Skills: AutogenCassandraChaos MonkeyChromaClaudeCrewaiDatadogDockerDynamoDBDynatraceFlinkFluentdGithub CopilotGitopsGo (Golang)GrafanaGremlinHadoopInfluxdbJavaKafkaKubernetesLangchainLanggraphLitmuschaosLogstashModel Context Protocol (Mcp)MongoDBNeo4JPineconePrometheusPythonPyTorchRabbitMQScikit-LearnSparkSplunkSqsTensorFlowTerraformTigergraphTimescaledbVectorWeaviate
Reposted 15 Days AgoSaved
Hybrid
New York, NY, USA
Mid level
Mid level
Financial Services
Operate and improve production reliability for critical services by building automation, monitoring, and runbooks. Triage incidents, reduce MTTR, improve observability, partner with engineering for root-cause fixes, and apply validated enterprise AI-assisted tools to support SRE workflows and reduce toil.
Top Skills: .NetAWSCi/CdDatadogEnterprise-Authorized AiJavaKafkaMqPythonSpring Boot
Reposted 15 Days AgoSaved
Hybrid
Orem, UT, USA
Mid level
Mid level
Financial Services
Design, implement, and operate reliable, scalable application and platform infrastructure. Build IaC, CI/CD pipelines, observability, and automation; collaborate with engineers to improve availability and SRE best practices.
Top Skills: .NetAzureDatadogDockerDynatraceEcsGitlabGrafanaJavaJenkinsKubernetesPrometheusPythonSplunkSpring BootTerraform
Reposted 15 Days AgoSaved
Hybrid
2 Locations
Senior level
Senior level
Financial Services
Lead design and delivery of reliability, observability, and SRE practices for large-scale data and AI/ML platforms. Define NFRs, SLI/SLOs, incident response, and implement reliable, secure, scalable infrastructure and automation while mentoring teams and driving safe, auditable AI-assisted operations.
Top Skills: AWSAws GlueCi/CdDatabricksDatadogDockerDynatraceGrafanaKubernetesMapreducePrometheusPythonSparkSplunkTerraform
Reposted 15 Days AgoSaved
Hybrid
Plano, TX, USA
Senior level
Senior level
Financial Services
Lead SRE responsible for driving reliability, scalability, observability, and resilience of web hosting platforms. Lead design reviews, mentor engineers, define SLOs/error budgets, use data-driven analytics and enterprise AI to improve incident triage and SDLC/toolchain reliability, and champion SRE practices and guardrails across teams.
Top Skills: Ai/MlAnsibleAWSAzureCi/CdCloudwatchDynatraceEnterprise AiGrafanaMonitoringObservabilityPrometheusPythonSdlcSplunkTelemetry
Reposted 15 Days AgoSaved
Hybrid
Houston, TX, USA
Mid level
Mid level
Financial Services
Designs, implements, and optimizes infrastructure and deployments to improve application availability, reliability, and scalability. Implements IaC and network-as-code, builds CI/CD pipelines, and enhances observability (monitoring, telemetry, SLOs). Collaborates cross-functionally, uses enterprise-authorized AI to accelerate incident triage and analysis, and drives iterative reliability improvements.
Top Skills: .NetAi (Enterprise-Authorized)Ci/CdCloudContainer OrchestrationContainerizationInfrastructure As CodeJavaNetwork As CodeObservabilityPythonService Level ObjectivesSpring BootTelemetry
Reposted 15 Days AgoSaved
Hybrid
2 Locations
Mid level
Mid level
Financial Services
Design, implement, monitor, and optimize production infrastructure and applications. Build infrastructure-as-code, CI/CD pipelines, and observability. Use enterprise-authorized AI/ML to accelerate incident triage and automate remediation. Collaborate cross-functionally on reliability, SLI/SLO/SLA, incident management, and resilience improvements.
Top Skills: .NetAWSCi/CdDatabricksDatadogDynatraceGrafanaJavaKubernetesPrometheusPysparkPythonSnowflakeSplunkSpring Boot
Reposted 15 Days AgoSaved
Hybrid
2 Locations
Senior level
Senior level
Financial Services
Lead SRE/DevOps engineer responsible for designing and delivering resilient, scalable Branch systems. Collaborates with product, architecture, security, and operations to embed reliability best practices across the SDLC, write and review production Java/Python code, implement IaC on AWS, use observability tooling, drive CI/CD and automation, and lead engineering communities of practice.
Top Skills: AWSCi/CdDatadogDynatraceEnvironment As Code (Eac)GrafanaInfrastructure As Code (Iac)JavaObservabilityPythonSplunkTerraform
Reposted 15 Days AgoSaved
Hybrid
Wilmington, DE, USA
Senior level
Senior level
Financial Services
Lead design, develop, and troubleshoot production-grade software with emphasis on reliability, scalability, security, and automation. Serve as technical lead during incidents, implement SRE best practices, and adopt enterprise-authorized AI tools to improve code quality, testing, and operational decisioning. Mentor engineers, drive observability and CI/CD automation, and ensure responsible, auditable AI usage across the SDLC.
Top Skills: .NetDatadogDynatraceEnterprise-Authorized Ai-Assisted Development ToolsGrafanaJava Spring BootPrometheusPythonSplunk
Reposted 16 Days AgoSaved
Hybrid
New York, NY, USA
112K-159K Annually
Senior level
112K-159K Annually
Senior level
Artificial Intelligence • Cloud • Fintech • Information Technology • Insurance • Financial Services • Big Data Analytics
Lead application readiness and remediation coordination for AWS, EOL, and vulnerability patches. Validate impacts, define smoke and regression tests, drive automation, resolve dependencies, escalate blockers, and secure production sign-off to ensure audit-ready closure.
Top Skills: AmiApi TestingAWSCertificatesCi/Cd PipelinesContainerizationDastDatabasesDockerEc2EksLibrariesMiddlewareNew RelicNew Relic MonitorsObservability ToolingRegression AutomationRuntimesSastScaService DashboardsSmoke TestingSyntheticsTerraform
Reposted 16 Days AgoSaved
Hybrid
O'Fallon, MO, USA
76K-127K Annually
Mid level
76K-127K Annually
Mid level
Blockchain • Fintech • Payments • Consulting • Cryptocurrency • Cybersecurity • Quantum Computing
The BizOps Engineer II is responsible for enhancing system reliability, automating workflows, and ensuring operational excellence across technology services at Mastercard, including application monitoring and CI/CD processes.
Top Skills: ArtifactoryBitbucketChefGitJavaJenkinsLinuxMainframeMaven
Reposted 16 Days AgoSaved
Hybrid
New York, NY, USA
112K-159K Annually
Senior level
112K-159K Annually
Senior level
Artificial Intelligence • Cloud • Fintech • Information Technology • Insurance • Financial Services • Big Data Analytics
Own and operate endpoint patch deployment and remediation for workstations and user devices: manage pilot rings and rollback groups, monitor patch success and endpoint health, validate post-patch functionality, coordinate incident triage with cross-functional teams, and improve automation, reporting, and evidence capture for vulnerability remediation.
Top Skills: DashboardingEdrEndpoint Management PlatformsItsmMecmMicrosoft IntunePowershellSccmTaniumVpn
New

Track Smarter, Apply Better.

Ditch the spreadsheets. Organize your job search with our freeApplication Tracker.

Use For Free
Application Tracker Preview
Reposted 16 Days AgoSaved
Easy Apply
Remote
United States
Easy Apply
100K-110K Annually
Mid level
100K-110K Annually
Mid level
Healthtech • Software
Operate and maintain AWS-hosted MERN applications and large-scale data workflows. Manage serverless and Spark-based pipelines, perform incident response and on-call duties, engineer automation to eliminate operational toil, ensure HIPAA/SOC2/HITRUST compliance, build observability and lead blameless post-mortems.
Top Skills: Amazon EcsAmazon EksAmazon EmrAthenaAws GlueAws LambdaAws SnsAws SqsCloudwatchEc2IamJavaScriptMernMySQLNode.jsOpentofuPysparkPythonRabbitMQTerraformTypescriptVpc
Reposted 16 Days AgoSaved
Remote or Hybrid
Orlando, FL, USA
Expert/Leader
Expert/Leader
AdTech • Cloud • Digital Media • Information Technology • News + Entertainment • App development
The Staff Site Reliability Engineer is responsible for ensuring the reliability, performance, and security of workplace collaboration services, focusing on automation, incident management, and operational excellence while providing technical leadership and mentoring to engineers.
Top Skills: Ai EngineeringAzure Virtual DesktopDefender For Office 365Exchange OnlineGraph ApiIntuneJamf ProMicrosoft 365Microsoft Entra IdMicrosoft PurviewOnedrivePowershellSharepoint OnlineTeams
17 Days AgoSaved
Remote or Hybrid
USA
140K-215K Annually
Senior level
140K-215K Annually
Senior level
Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity
Senior SRE owning availability, automation, and observability for CI/CD platform services. Build and operate infrastructure, run on-call, lead incident response, mentor engineers, drive design/capacity planning, integrate AI-assisted workflows, and improve cross-team reliability.
Top Skills: Active DirectoryAnsibleApache AirflowSparkAWSAzureBashBazelBitbucketCassandraChefDatadogDnsFirewall RulesGCPGitGithub ActionsGitlabGitlab CiGoGrafanaHoneycombHumio/LogscaleJenkinsKafkaKubernetesLoad BalancersMongoDBMySQLNasNew RelicNfsObject StorageOpensearchOraclePostgresPowershellPrometheusPulsarPuppetPythonRabbitMQRedis/ValkeyRedpandaRoutingSaltSanSplunkTerraformVarnishVipsWindows Server
Reposted 18 Days AgoSaved
Easy Apply
Hybrid
New York City, NY, USA
Easy Apply
151K-191K Annually
Senior level
151K-191K Annually
Senior level
Fintech • Information Technology • Software • Financial Services
The role involves designing and automating infrastructure management, improving reliability, building internal tools, and contributing to architectural decisions. Responsibilities include working with Kubernetes and managing large-scale infrastructure, while participating in on-call rotations to prevent incidents.
Top Skills: CloudwatchDatadogDockerEfkElkGoJavaScriptKubernetesPythonTerraform
Reposted 18 Days AgoSaved
In-Office or Remote
New York, NY, USA
150K-250K Annually
Mid level
150K-250K Annually
Mid level
Mobile • Software
Site Reliability Engineers will work on production infrastructure, focusing on AWS and Kubernetes while ensuring high availability and customer satisfaction.
Top Skills: AirflowAWSCircleCICloudwatchEksGrafanaMongoDBPagerdutyPingdomRustScala SparkTerraformTypescript
Reposted 18 Days AgoSaved
Easy Apply
Remote or Hybrid
6 Locations
Easy Apply
126K-248K Annually
Senior level
126K-248K Annually
Senior level
Big Data • Cloud • Software • Database
The Senior Site Reliability Engineer will develop and support distributed storage services, ensuring reliability and operational safety, with a focus on automation and efficiency.
Top Skills: AWSAzureDnsGoGoogle Cloud PlatformKubernetesLinuxPythonTcp/IpTls
Reposted 19 Days AgoSaved
Easy Apply
In-Office
Chicago, IL, USA
Easy Apply
Senior level
Senior level
AdTech
Design, build, and scale cloud-native infrastructure with automation-first approach. Develop Terraform modules, Helm charts, Istio routing, observability (Prometheus/Grafana/Datadog), maintain GCP databases, improve CI/CD, and use AI agents to automate and operationalize reliability and developer experience.
Top Skills: Amazon KinesisAWSAws LambdaAws SnsCi/CdClaude CodeCloudsqlCursorDatadogDockerGCPGitlabGoogle BigqueryGoogle Cloud FunctionsGoogle Cloud RunGoogle Pub/SubGoogle SpannerGrafanaHelmIstioKafkaKubernetesMySQLPrometheusSQLTerraform
179K-225K Annually
Mid level
Fintech • Machine Learning • Payments • Software • Financial Services
Lead a portfolio of cloud-native full‑stack and SRE projects, mentor engineers, collaborate with product managers, and deliver resilient AWS/GCP/Azure services using Java, Python, containers, orchestration, and observability tooling.
Top Skills: Ai ToolingAWSDockerGCPHTML/CSSJavaKubernetesAzureNoSQLPythonRdbms
20 Days AgoSaved
In-Office
Seattle, WA, USA
167K-204K Annually
Senior level
167K-204K Annually
Senior level
Real Estate • PropTech
Embedded SRE focused on building self-service reliability infrastructure and observability across the full stack. Responsibilities include implementing end-to-end APM/RUM and tracing, operationalizing AI gateway monitoring and guardrails, integrating AI for incident diagnosis, advising teams on resilient architecture, and modernizing observable CI/CD pipelines. Role may include on-call rotation.
Top Skills: ApmAWSCi/CdCliDatadogDistributed TracingJavaJavaScriptKubernetesReactRumSpringTypescript
20 Days AgoSaved
Remote or Hybrid
New York, NY, USA
Senior level
Senior level
AdTech • Cloud • Digital Media • Information Technology • News + Entertainment • App development
Lead SRE for cloud-based live linear playout systems driving reliability, observability, incident response, SLIs/SLOs, automation, capacity planning, runbooks, and L1/L2 on-call support to ensure resilient distribution across NBCUniversal channels.
Top Skills: AmagiAWSCmafDockerEsamGrafanaH.264HarrisHevcHlsImagineIp NetworkingKubernetesLinuxMicrosoft TeamsRistScte-224Scte-35ServicenowSlackSnellSplunkSrtTs
Reposted 20 Days AgoSaved
Easy Apply
Hybrid
San Jose, CA, USA
Easy Apply
123K-175K Annually
Senior level
123K-175K Annually
Senior level
Cloud • Information Technology • Security • Software • Cybersecurity
The role involves creating scalable solutions using Linux and Kubernetes, troubleshooting performance issues, maintaining security, and writing automation tools.
Top Skills: AnsibleBashDockerFirewall TechnologiesGoKubernetesKvmLinuxMulti-Factor AuthenticationOpenstackPgpPkiPythonSshUnix
Reposted 20 Days AgoSaved
Hybrid
O'Fallon, MO, USA
122K-207K Annually
Senior level
122K-207K Annually
Senior level
Blockchain • Fintech • Payments • Consulting • Cryptocurrency • Cybersecurity • Quantum Computing
The Lead Site Reliability Engineer will ensure the reliability and performance of Mastercard's applications, mentor junior engineers, and improve service lifecycle through automation and DevOps practices.
Top Skills: GoJavaPythonSpring Framework
21 Days AgoSaved
In-Office or Remote
Basking Ridge, NJ, USA
113K-193K Annually
Senior level
113K-193K Annually
Senior level
Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
Lead enterprise SRE, DevOps, ITSM, and operational excellence for critical banking platforms. Drive reliability (SLI/SLO), automation, CI/CD, IaC, observability, incident/problem/change management, disaster recovery, and AI-enabled operational improvements while building and mentoring cross-functional teams.
Top Skills: AiopsAWSAzureChatopsCi/CdDatadogGitopsGrafanaInfrastructure-As-CodeKubernetesLlmObservabilityOn-PremOpenshiftOpentelemetryPrometheusSplunk
All Filters
JobType
New Jobs
Job Category
Experience
Industry
Company Name
Company Size

Sign up now Access later

Create Free Account