Top Hybrid DevOps & Platform Engineering Jobs

8 Days AgoSaved
Easy Apply
Hybrid
3 Locations
Easy Apply
186K-232K Annually
Senior level
186K-232K Annually
Senior level
Artificial Intelligence • Big Data • Healthtech • Biotech • Pharmaceutical
Lead the infrastructure and SRE team, setting technical direction for reliable, secure, and scalable engineering systems. Hire, coach, and develop engineers; oversee multi-cloud infrastructure, Kubernetes, IaC, CI/CD, observability, incident response, support rotations, and regulated workloads. Partner across engineering, data science, security, and operations to improve developer experience and enable safe AI-assisted software delivery. Manage infrastructure roadmaps, budgets, resourcing, documentation, and organizational reliability practices.
Top Skills: Ai Model Training WorkloadsAWSAzureCi/CdContainerized ApplicationsCots SoftwareDatabasesDockerFoss SoftwareGCPGitInfrastructure As CodeKubernetesLoad BalancersMl PipelinesMulti-Cloud InfrastructureOpentofuPythonSecrets ManagementSnowflakeTerraformVercelVirtual Networking
8 Days AgoSaved
Hybrid
Salisbury, NC, USA
163K-245K Annually
Expert/Leader
163K-245K Annually
Expert/Leader
AdTech • eCommerce • Food • Marketing Tech • Retail
Owns enterprise hosting-platform strategy, governance, roadmaps, and cross-domain engineering outcomes. Designs reusable infrastructure automation, IaC pipelines, self-service capabilities, observability, lifecycle controls, and governed generative or agentic AI workflows. Leads evaluation, reliability, security, FinOps, disaster recovery, and operational readiness while influencing senior stakeholders and coaching technical teams. The role spans Windows, Linux, AIX, virtualization, storage, middleware, networking, ServiceNow, and production infrastructure across a large enterprise.
Top Skills: Active DirectoryAixAnsibleArmAzure Ai FoundryBackupBicepCi/CdCis BenchmarksConfiguration ManagementDatadogDnsEsxiFibre ChannelFirewallsGitGitopsGroup PolicyHmcHyper-VInfrastructure As CodeKubernetesLanggraphLinuxMicrosoft Agent FrameworkMiddlewareNagiosNasNetworkingNimNist Ai RmfNist Cybersecurity FrameworkNutanix AosOpentofuPci ControlsPowerPowershellPowervmPrismPythonRhelSanSemantic KernelServicenow CmdbServicenow DiscoveryServicenow ItsmShell ScriptingStorageTerraformTerraformVcenterVmware VsphereWindows ServerZero Trust
8 Days AgoSaved
Remote or Hybrid
United States
125K-210K Annually
Senior level
125K-210K Annually
Senior level
Artificial Intelligence • Cloud • Sales • Security • Software • Cybersecurity • Data Privacy
Leads the data infrastructure platform team responsible for production services supporting batch, streaming, analytics, and machine learning workloads. Oversees people leadership, platform strategy, reliability, security, scalability, cost efficiency, incident response, roadmap execution, and cross-functional collaboration. Builds self-service developer experiences, standardized CI/CD, observability, infrastructure automation, and service ownership practices across Airflow, Flink, Spark, Kafka, Snowflake, Iceberg, AWS, and ML platforms.
Top Skills: Amazon EksApache AirflowApache FlinkApache IcebergApache KafkaSparkAWSAws EmrAws SagemakerCi/CdDatabricksGitopsGoInfrastructure As CodeJavaKubernetesLlmsPythonSnowflakeTerraform
8 Days AgoSaved
Hybrid
Honolulu, HI, USA
135K-145K Annually
Entry level
135K-145K Annually
Entry level
Artificial Intelligence • Software
Deploy, operate, and maintain scalable Palantir infrastructure and services for government customers. Responsibilities include monitoring, configuration management, upgrades, migrations, troubleshooting production issues, optimizing reliability, supporting on-call operations, automating manual workflows, and developing infrastructure solutions for Foundry and Apollo platforms.
Top Skills: AIApolloBashConfiguration ManagementDistributed SystemsFoundryGoJavaJavaScriptLlmLoad BalancingMonitoring And AlertingPython
8 Days AgoSaved
Hybrid
Boston, MA, USA
150K-215K Annually
Senior level
150K-215K Annually
Senior level
Fitness • Hardware • Healthtech • Sports • Wearables
Designs and operates Kubernetes clusters on AWS, builds scalable and secure AI runtime infrastructure, and develops tooling for developer productivity and automated development workflows. The role improves platform reliability, supports long-running tasks on cost-efficient infrastructure, partners with application, security, and data teams, participates in incident response, and provides technical leadership and mentorship across the Platform organization.
Top Skills: Ai RuntimesAWSInfrastructure As CodeKubernetesTerraform
New

Cut your apply time in half.

Use ourAI Assistantto automatically fill your job applications.

Use For Free
Application Tracker Preview
11 Days AgoSaved
Hybrid
Seattle, WA, USA
185K-295K Annually
Senior level
185K-295K Annually
Senior level
AdTech • eCommerce • Information Technology • Software • Travel • Generative AI
Leads a team of 10–12 Workday engineers and platform specialists while owning Workday capabilities, platform operations, roadmap execution, implementations, integrations, security, governance, reliability, and continuous improvement. Partners with HR, Finance, Payroll, Product, and Technology leaders to translate business priorities into scalable solutions, develop engineering standards, manage team capacity and performance, and identify AI and automation opportunities.
Top Skills: AzureOracleSalesforceSAPServicenowWorkdayWorkday ExtendWorkday HcmWorkday IntegrationsWorkday Prism Analytics
11 Days AgoSaved
Remote or Hybrid
USA
190K-200K Annually
Senior level
190K-200K Annually
Senior level
eCommerce • Fintech • Food • Mobile • Social Impact
Own and evolve AWS cloud infrastructure supporting a financial and hospitality technology platform. Design highly available systems, deployment pipelines, observability, security controls, and infrastructure migrations. Lead incident response, reliability planning, capacity management, and root-cause analysis while participating in on-call rotations. Partner with application engineers, establish infrastructure standards, and mentor platform engineers as the organization scales.
Top Skills: AlbAWSCloudFormationCloudwatchEc2EcsEksElasticacheFargateIamNlbRuby on RailsRdsRedisTerraformValkeyVpc
11 Days AgoSaved
Easy Apply
Hybrid
San Francisco, CA, USA
Easy Apply
150K-220K Annually
Expert/Leader
150K-220K Annually
Expert/Leader
Fintech • Financial Services
Leads Forge’s Site Reliability Engineering team, improving system availability, observability, incident response, disaster recovery, automation, and production operations. Partners with Platform, Engineering, Security, Compliance, Risk, and Product teams to strengthen reliability and operational maturity in a regulated environment. Responsibilities include technical design, troubleshooting, team hiring, coaching, performance management, and career development while helping deliver secure, scalable, highly reliable products.
Top Skills: Amazon CloudwatchAnsibleAWSAzureCi/CdCloud InfrastructureDatadogDistributed SystemsInfrastructure-As-CodeKubernetesObservabilityTerraform
11 Days AgoSaved
Hybrid
New York, NY, USA
148K-211K Annually
Senior level
148K-211K Annually
Senior level
Artificial Intelligence • Cloud • Fintech • Information Technology • Insurance • Financial Services • Big Data Analytics
Leads cloud platform engineering and site reliability initiatives across AWS, GCP, and Azure. Designs reusable Terraform infrastructure, governed platform patterns, CI/CD integrations, and full-stack solutions for applications, data platforms, and AI services. Establishes SRE practices including SLOs, SLIs, error budgets, observability, incident response, and cost optimization. Partners with application, security, architecture, and business teams; mentors engineers, facilitates technical reviews, resolves operational issues, and drives standardization and process improvement.
Top Skills: .NetAgentic AiAmazon EksAPIsAWSAws CdkAzureC#Ci/CdCloudFormationFinopsGCPGitGithub ActionsHarnessJavaJavaScriptJfrog ArtifactoryKubernetesPowershellPythonRagSlo/SliSQLTerraformYaml
11 Days AgoSaved
Easy Apply
Hybrid
Chicago metropolitan area, Chicago, IL, USA
Easy Apply
64K-72K Annually
Junior
64K-72K Annually
Junior
Information Technology • Machine Learning • Natural Language Processing • Software • App development • Conversational AI • Big Data Analytics
Monitor applications, platforms, servers, hardware, and network connectivity in a 24/7 NOC. Investigate alerts, troubleshoot system and network errors, identify root causes, triage incidents to appropriate teams, document issues in tickets, and resolve problems within NOC scope using terminal commands. The role requires on-site day-shift work on a rotating schedule and strong networking, Linux, monitoring, communication, and troubleshooting skills.
Top Skills: AlertaJuniper Network EquipmentLinuxMS OfficeNagiosObserviumOsi ModelPythonTcp/Ip
All Filters
JobType
New Jobs
Job Category
Experience
Industry
Company Name
Company Size

Sign up now Access later

Create Free Account