Maximum of 25 job preferences reached.
Top Site Reliability Engineer Jobs
Financial Services
Lead design and implementation of SRE practices, observability, reliability and AI-assisted operational workflows. Mentor engineers, define NFRs/SLOs, build automation, logging/metrics/tracing pipelines, containerized CI/CD/GitOps, and integrate AI agents for production-grade reliability.
Top Skills:
AutogenCassandraChaos MonkeyChromaClaudeCrewaiDatadogDockerDynamoDBDynatraceFlinkFluentdGithub CopilotGitopsGo (Golang)GrafanaGremlinHadoopInfluxdbJavaKafkaKubernetesLangchainLanggraphLitmuschaosLogstashModel Context Protocol (Mcp)MongoDBNeo4JPineconePrometheusPythonPyTorchRabbitMQScikit-LearnSparkSplunkSqsTensorFlowTerraformTigergraphTimescaledbVectorWeaviate
Financial Services
Operate and improve production reliability for critical services by building automation, monitoring, and runbooks. Triage incidents, reduce MTTR, improve observability, partner with engineering for root-cause fixes, and apply validated enterprise AI-assisted tools to support SRE workflows and reduce toil.
Top Skills:
.NetAWSCi/CdDatadogEnterprise-Authorized AiJavaKafkaMqPythonSpring Boot
Financial Services
Design, implement, and operate reliable, scalable application and platform infrastructure. Build IaC, CI/CD pipelines, observability, and automation; collaborate with engineers to improve availability and SRE best practices.
Top Skills:
.NetAzureDatadogDockerDynatraceEcsGitlabGrafanaJavaJenkinsKubernetesPrometheusPythonSplunkSpring BootTerraform
Reposted 15 Days AgoSaved
Financial Services
Lead design and delivery of reliability, observability, and SRE practices for large-scale data and AI/ML platforms. Define NFRs, SLI/SLOs, incident response, and implement reliable, secure, scalable infrastructure and automation while mentoring teams and driving safe, auditable AI-assisted operations.
Top Skills:
AWSAws GlueCi/CdDatabricksDatadogDockerDynatraceGrafanaKubernetesMapreducePrometheusPythonSparkSplunkTerraform
Financial Services
Lead SRE responsible for driving reliability, scalability, observability, and resilience of web hosting platforms. Lead design reviews, mentor engineers, define SLOs/error budgets, use data-driven analytics and enterprise AI to improve incident triage and SDLC/toolchain reliability, and champion SRE practices and guardrails across teams.
Top Skills:
Ai/MlAnsibleAWSAzureCi/CdCloudwatchDynatraceEnterprise AiGrafanaMonitoringObservabilityPrometheusPythonSdlcSplunkTelemetry
Financial Services
Designs, implements, and optimizes infrastructure and deployments to improve application availability, reliability, and scalability. Implements IaC and network-as-code, builds CI/CD pipelines, and enhances observability (monitoring, telemetry, SLOs). Collaborates cross-functionally, uses enterprise-authorized AI to accelerate incident triage and analysis, and drives iterative reliability improvements.
Top Skills:
.NetAi (Enterprise-Authorized)Ci/CdCloudContainer OrchestrationContainerizationInfrastructure As CodeJavaNetwork As CodeObservabilityPythonService Level ObjectivesSpring BootTelemetry
Financial Services
Design, implement, monitor, and optimize production infrastructure and applications. Build infrastructure-as-code, CI/CD pipelines, and observability. Use enterprise-authorized AI/ML to accelerate incident triage and automate remediation. Collaborate cross-functionally on reliability, SLI/SLO/SLA, incident management, and resilience improvements.
Top Skills:
.NetAWSCi/CdDatabricksDatadogDynatraceGrafanaJavaKubernetesPrometheusPysparkPythonSnowflakeSplunkSpring Boot
Financial Services
Lead SRE/DevOps engineer responsible for designing and delivering resilient, scalable Branch systems. Collaborates with product, architecture, security, and operations to embed reliability best practices across the SDLC, write and review production Java/Python code, implement IaC on AWS, use observability tooling, drive CI/CD and automation, and lead engineering communities of practice.
Top Skills:
AWSCi/CdDatadogDynatraceEnvironment As Code (Eac)GrafanaInfrastructure As Code (Iac)JavaObservabilityPythonSplunkTerraform
Financial Services
Lead design, develop, and troubleshoot production-grade software with emphasis on reliability, scalability, security, and automation. Serve as technical lead during incidents, implement SRE best practices, and adopt enterprise-authorized AI tools to improve code quality, testing, and operational decisioning. Mentor engineers, drive observability and CI/CD automation, and ensure responsible, auditable AI usage across the SDLC.
Top Skills:
.NetDatadogDynatraceEnterprise-Authorized Ai-Assisted Development ToolsGrafanaJava Spring BootPrometheusPythonSplunk
Reposted 16 Days AgoSaved
Artificial Intelligence • Cloud • Fintech • Information Technology • Insurance • Financial Services • Big Data Analytics
Lead application readiness and remediation coordination for AWS, EOL, and vulnerability patches. Validate impacts, define smoke and regression tests, drive automation, resolve dependencies, escalate blockers, and secure production sign-off to ensure audit-ready closure.
Top Skills:
AmiApi TestingAWSCertificatesCi/Cd PipelinesContainerizationDastDatabasesDockerEc2EksLibrariesMiddlewareNew RelicNew Relic MonitorsObservability ToolingRegression AutomationRuntimesSastScaService DashboardsSmoke TestingSyntheticsTerraform
Blockchain • Fintech • Payments • Consulting • Cryptocurrency • Cybersecurity • Quantum Computing
The BizOps Engineer II is responsible for enhancing system reliability, automating workflows, and ensuring operational excellence across technology services at Mastercard, including application monitoring and CI/CD processes.
Top Skills:
ArtifactoryBitbucketChefGitJavaJenkinsLinuxMainframeMaven
Artificial Intelligence • Cloud • Fintech • Information Technology • Insurance • Financial Services • Big Data Analytics
Own and operate endpoint patch deployment and remediation for workstations and user devices: manage pilot rings and rollback groups, monitor patch success and endpoint health, validate post-patch functionality, coordinate incident triage with cross-functional teams, and improve automation, reporting, and evidence capture for vulnerability remediation.
Top Skills:
DashboardingEdrEndpoint Management PlatformsItsmMecmMicrosoft IntunePowershellSccmTaniumVpn
New
Track Smarter, Apply Better.
Ditch the spreadsheets. Organize your job search with our freeApplication Tracker.
Use For Free
Healthtech • Software
Operate and maintain AWS-hosted MERN applications and large-scale data workflows. Manage serverless and Spark-based pipelines, perform incident response and on-call duties, engineer automation to eliminate operational toil, ensure HIPAA/SOC2/HITRUST compliance, build observability and lead blameless post-mortems.
Top Skills:
Amazon EcsAmazon EksAmazon EmrAthenaAws GlueAws LambdaAws SnsAws SqsCloudwatchEc2IamJavaScriptMernMySQLNode.jsOpentofuPysparkPythonRabbitMQTerraformTypescriptVpc
AdTech • Cloud • Digital Media • Information Technology • News + Entertainment • App development
The Staff Site Reliability Engineer is responsible for ensuring the reliability, performance, and security of workplace collaboration services, focusing on automation, incident management, and operational excellence while providing technical leadership and mentoring to engineers.
Top Skills:
Ai EngineeringAzure Virtual DesktopDefender For Office 365Exchange OnlineGraph ApiIntuneJamf ProMicrosoft 365Microsoft Entra IdMicrosoft PurviewOnedrivePowershellSharepoint OnlineTeams
Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity
Senior SRE owning availability, automation, and observability for CI/CD platform services. Build and operate infrastructure, run on-call, lead incident response, mentor engineers, drive design/capacity planning, integrate AI-assisted workflows, and improve cross-team reliability.
Top Skills:
Active DirectoryAnsibleApache AirflowSparkAWSAzureBashBazelBitbucketCassandraChefDatadogDnsFirewall RulesGCPGitGithub ActionsGitlabGitlab CiGoGrafanaHoneycombHumio/LogscaleJenkinsKafkaKubernetesLoad BalancersMongoDBMySQLNasNew RelicNfsObject StorageOpensearchOraclePostgresPowershellPrometheusPulsarPuppetPythonRabbitMQRedis/ValkeyRedpandaRoutingSaltSanSplunkTerraformVarnishVipsWindows Server
Fintech • Information Technology • Software • Financial Services
The role involves designing and automating infrastructure management, improving reliability, building internal tools, and contributing to architectural decisions. Responsibilities include working with Kubernetes and managing large-scale infrastructure, while participating in on-call rotations to prevent incidents.
Top Skills:
CloudwatchDatadogDockerEfkElkGoJavaScriptKubernetesPythonTerraform
Mobile • Software
Site Reliability Engineers will work on production infrastructure, focusing on AWS and Kubernetes while ensuring high availability and customer satisfaction.
Top Skills:
AirflowAWSCircleCICloudwatchEksGrafanaMongoDBPagerdutyPingdomRustScala SparkTerraformTypescript
Reposted 18 Days AgoSaved
Easy Apply
Easy Apply
Big Data • Cloud • Software • Database
The Senior Site Reliability Engineer will develop and support distributed storage services, ensuring reliability and operational safety, with a focus on automation and efficiency.
Top Skills:
AWSAzureDnsGoGoogle Cloud PlatformKubernetesLinuxPythonTcp/IpTls
Reposted 19 Days AgoSaved
Easy Apply
Easy Apply
AdTech
Design, build, and scale cloud-native infrastructure with automation-first approach. Develop Terraform modules, Helm charts, Istio routing, observability (Prometheus/Grafana/Datadog), maintain GCP databases, improve CI/CD, and use AI agents to automate and operationalize reliability and developer experience.
Top Skills:
Amazon KinesisAWSAws LambdaAws SnsCi/CdClaude CodeCloudsqlCursorDatadogDockerGCPGitlabGoogle BigqueryGoogle Cloud FunctionsGoogle Cloud RunGoogle Pub/SubGoogle SpannerGrafanaHelmIstioKafkaKubernetesMySQLPrometheusSQLTerraform
20 Days AgoSaved
Fintech • Machine Learning • Payments • Software • Financial Services
Lead a portfolio of cloud-native full‑stack and SRE projects, mentor engineers, collaborate with product managers, and deliver resilient AWS/GCP/Azure services using Java, Python, containers, orchestration, and observability tooling.
Top Skills:
Ai ToolingAWSDockerGCPHTML/CSSJavaKubernetesAzureNoSQLPythonRdbms
Real Estate • PropTech
Embedded SRE focused on building self-service reliability infrastructure and observability across the full stack. Responsibilities include implementing end-to-end APM/RUM and tracing, operationalizing AI gateway monitoring and guardrails, integrating AI for incident diagnosis, advising teams on resilient architecture, and modernizing observable CI/CD pipelines. Role may include on-call rotation.
Top Skills:
ApmAWSCi/CdCliDatadogDistributed TracingJavaJavaScriptKubernetesReactRumSpringTypescript
AdTech • Cloud • Digital Media • Information Technology • News + Entertainment • App development
Lead SRE for cloud-based live linear playout systems driving reliability, observability, incident response, SLIs/SLOs, automation, capacity planning, runbooks, and L1/L2 on-call support to ensure resilient distribution across NBCUniversal channels.
Top Skills:
AmagiAWSCmafDockerEsamGrafanaH.264HarrisHevcHlsImagineIp NetworkingKubernetesLinuxMicrosoft TeamsRistScte-224Scte-35ServicenowSlackSnellSplunkSrtTs
Cloud • Information Technology • Security • Software • Cybersecurity
The role involves creating scalable solutions using Linux and Kubernetes, troubleshooting performance issues, maintaining security, and writing automation tools.
Top Skills:
AnsibleBashDockerFirewall TechnologiesGoKubernetesKvmLinuxMulti-Factor AuthenticationOpenstackPgpPkiPythonSshUnix
Blockchain • Fintech • Payments • Consulting • Cryptocurrency • Cybersecurity • Quantum Computing
The Lead Site Reliability Engineer will ensure the reliability and performance of Mastercard's applications, mentor junior engineers, and improve service lifecycle through automation and DevOps practices.
Top Skills:
GoJavaPythonSpring Framework
Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
Lead enterprise SRE, DevOps, ITSM, and operational excellence for critical banking platforms. Drive reliability (SLI/SLO), automation, CI/CD, IaC, observability, incident/problem/change management, disaster recovery, and AI-enabled operational improvements while building and mentoring cross-functional teams.
Top Skills:
AiopsAWSAzureChatopsCi/CdDatadogGitopsGrafanaInfrastructure-As-CodeKubernetesLlmObservabilityOn-PremOpenshiftOpentelemetryPrometheusSplunk
Let Your Resume Do The Work
Upload your resume to be matched with jobs you're a great fit for.
Success! We'll use this to further personalize your experience.
Top Companies Hiring Site Reliability Engineers
See AllPopular Job Searches
All Software Engineer Jobs
.NET Developer Jobs
Aerospace Thermal Engineering Jobs
AI Engineer Jobs
Android Developer Jobs
Automation Engineer Jobs
Backend Developer Jobs
Blockchain Developer Jobs
C# Jobs
C++ Jobs
Cloud Architect Jobs
Cloud Engineer Jobs
Design Engineer Jobs
DevOps Engineer Jobs
Director Of Engineering Jobs
Electrical Engineering Jobs
Embedded Software Engineer Jobs
Engineering Jobs
Engineering Manager Jobs
Environmental Engineering Jobs
Field Engineer Jobs
Front End Developer Jobs
Full Stack Developer Jobs
Game Developer Jobs
Golang Jobs
Hardware Engineer Jobs
Industrial Engineering Jobs
iOS Developer Jobs
Java Developer Jobs
Javascript Developer Jobs
Linux Jobs
Manufacturing Engineer Jobs
Mechanical Engineering Jobs
Network Engineer Jobs
PHP Developer Jobs
Process Engineer Jobs
Project Engineer Jobs
Prompt Engineering Jobs
Python Jobs
QA Jobs
Robotics Engineer Jobs
Ruby on Rails Jobs
Salesforce Administrator Jobs
Salesforce Developer Jobs
Scala Jobs
Sharepoint Developer Jobs
Site Reliability Engineer Jobs
Software Engineering Manager Jobs
Solutions Architect Jobs
SQL Developer Jobs
Structural Engineer Jobs
System Engineer Jobs
Test Engineer Jobs
Web Developer Jobs
All Filters
Total selected ()
No Results
No Results




.png)














.png)





