Maximum of 25 job preferences reached.
Top Hybrid DevOps & Platform Engineering Jobs
Artificial Intelligence • Hardware • Machine Learning • Robotics • Software • Utilities
Lead design and implementation of reliability, observability, and CI/CD for Airflow/Astronomer pipelines and containerized workloads on AWS/Kubernetes. Build monitoring, alerting, SLOs, inference instrumentation, operational tooling, runbooks, and automate toil while partnering with data and engineering teams to scale production ML and data pipelines.
Top Skills:
Apache AirflowAstronomerAstronomer CliAWSContainerizationDockerDuplocloudEcrEcsGeospatialKubernetesLambdaLidarLinuxMl OpsPoint CloudS3SqsStep FunctionsTerraform
Artificial Intelligence • Cloud • Consumer Web • eCommerce • Information Technology • Software
Design, build, and optimize serverless, API-driven integrations on AWS. Develop Lambda-based functions, RESTful APIs, and event-driven data flows using services like API Gateway, SQS/SNS, Step Functions, and EventBridge. Implement security best practices, CI/CD and IaC, test and debug integrations, monitor with CloudWatch/X-Ray, and document designs while collaborating with architects and cross-functional teams.
Top Skills:
Amazon Api GatewayAmazon EventbridgeAmazon SnsAmazon SqsAws AppflowAws CloudwatchAws GlueAws LambdaAws SdksAws Step FunctionsAws X-RayCi/CdCloudFormationGitJavaJwtNode.jsOauthPythonRestful ApisSalesforceServicenowTerraformWorkday
Software
Own and operate OnRamp's AWS platform, IaC, CI/CD, and deployment tooling; build AI/agent execution infrastructure and observability; drive reliability, incident response, security/compliance, and automation using coding agents and LLM tools.
Top Skills:
AgentsAWSCdkCi/CdContainersInfrastructure-As-CodeLlm ToolingObservabilityTerraform
Fintech • Software • Financial Services
Design, build, and operate highly available, secure infrastructure for a multi-tenant financial SaaS platform. Lead automation (IaC), Kubernetes operations, CI/CD, observability, connectivity for market/custodian feeds, incident escalation and RCA, capacity planning, compliance (SOC 2/ISO27001), cost optimization, and mentor engineers.
Top Skills:
AnsibleAws Ec2BashCheckmkClaudeCyberarkDebianDnsEksElastic Stack (Elk/Efk)FixFtpsGithub ActionsGitlab CiGrafanaHarborHashicorp VaultHelmIstioJavaJenkinsJvmKubernetesLet'S EncryptLinkerdMySQLPostgresPrometheusProxmoxPythonRdsRhelS3SaltSftpSpinnakerSwiftTcp/IpTerraformTime-Series DatabasesTls/PkiUbuntuVmware VsphereVpn/Mpls
Healthtech • Software
The Software Engineer, Infrastructure will design, build, and operate core platform services while ensuring reliability, security, and efficiency through collaboration with various teams and improving developer experience.
Top Skills:
Ci/CdGoGoogle Cloud PlatformInfrastructure-As-CodeJavaScriptKubernetesPythonTerraform
New
Cut your apply time in half.
Use ourAI Assistantto automatically fill your job applications.
Use For Free
Reposted One Month AgoSaved
Easy Apply
Easy Apply
Artificial Intelligence • Big Data • Computer Vision • Information Technology • Machine Learning • Analytics • Defense
The Senior DevOps Engineer will manage AI deployments, enhance cloud environments, implement Infrastructure-as-Code, and ensure software reliability while guiding junior team members.
Top Skills:
AnsibleAWSAzureGCPGoHelmKubernetesLinuxPythonTerraform
eCommerce • Healthtech • Kids + Family • Retail • Social Media
Own and evolve Babylist's AWS infrastructure and developer platform using Terraform and Kubernetes. Improve CI/CD reliability, support engineers across environments, define monitoring and alerting standards, lead incident response and postmortems, and shape platform architecture to scale for millions of users.
Top Skills:
AWSCdnCircleCICronitorDatadogDnsEksGithub ActionsKubernetesLoad BalancersMySQLPagerdutyRdsRedisRuby On RailsSentrySidekiqTerraform
Cloud • Information Technology • Software
Leads Box’s engineering-first FinOps team responsible for forecasting and optimizing public cloud, SaaS, and AI spending. Builds scalable tooling, pipelines, and models using GCP, BigQuery, Python, and SQL; translates cost and usage data into executive insights; drives optimization and governance initiatives; establishes dashboards and operating metrics; develops senior engineers and data scientists; partners with Cloud Ops, SRE, Data, and Finance; and oversees on-call incident response and reliability improvements.
Top Skills:
BigQueryGoogle Cloud PlatformPythonSQL
Database • Analytics • Consulting
Serves as the senior technical authority for multi-cloud infrastructure across AWS, Google Cloud, and Azure. Defines architecture standards, Terraform practices, Kubernetes platforms, CI/CD automation, security controls, observability, reliability, and FinOps strategies. Leads complex architecture decisions, production escalations, client solutioning, pre-sales, design reviews, and technical mentorship. Builds reusable accelerators and reference architectures while remaining a hands-on individual contributor.
Top Skills:
AcrAksArgocdArtifact RegistryAWSAws Control TowerAzureAzure DevopsAzure Landing ZonesBashCi/CdCloud-InitDnsEcrEksFinopsFluxGithub ActionsGitlab CiGkeGCPIamJenkinsKubernetesOpaPowershellPrivate Service ConnectPrivatelinkPythonRbacSentinelService MeshShared VpcTerraformVMwareVnetVpc
Database • Analytics • Consulting
Principal-level cloud infrastructure engineer serving as the technical authority for AWS, Google Cloud, and Azure. Responsibilities include defining multi-cloud architectures, Terraform standards, production Kubernetes platforms, CI/CD automation, security controls, observability, reliability, FinOps, and incident response. The role leads technical solutioning and pre-sales, mentors senior engineers, establishes engineering standards, and builds reusable infrastructure accelerators for internal and client-facing engagements.
Top Skills:
Amazon EcrAmazon EksArgocdAWSAzureAzure AksAzure Container RegistryAzure DevopsBashCloud-InitFluxGithub ActionsGitlab CiGoogle Artifact RegistryGCPGoogle GkeJenkinsKubernetesOpaPowershellPrivate Service ConnectPrivatelinkPythonSentinelService MeshShared VpcTerraformVMware
Software
Lead global GPU capacity management across multi-cloud environments, including cluster acquisition, Kubernetes orchestration, workload migration, infrastructure automation, incident response, and GPU fleet maintenance. Build scalable capacity systems, develop Go-based operators, optimize reliability and cost through financial modeling, and coordinate infrastructure initiatives across SRE, Infra, and FDE teams.
Top Skills:
AWSGoGoogle Cloud PlatformKubernetesAzureMulti-Cloud InfrastructureNvidia B200Nvidia H100Python
Let Your Resume Do The Work
Upload your resume to be matched with jobs you're a great fit for.
Success! We'll use this to further personalize your experience.
All Filters
Total selected ()
No Results
No Results


.png)


















