Top DevOps & Platform Engineering Jobs in Lexington, KY

Reposted YesterdaySaved
Remote
United States
166K-331K Annually
Senior level
166K-331K Annually
Senior level
Software • Quantum Computing • Metaverse • Infrastructure as a Service (IaaS)
Lead and develop Site Reliability Engineering teams operating Microsoft Substrate services in regulated environments. Own reliability, availability, incident response, disaster recovery, SLOs, operational metrics, automation, and compliance. Serve in the on-call rotation, drive systemic post-incident improvements, conduct recovery exercises, and partner with engineering, product, security, infrastructure, and compliance teams to build resilient, auditable cloud services.
Top Skills: AutomationCloud ComputingDisaster RecoveryDistributed SystemsMicrosoft 365AzureMicrosoft SubstrateSlisSlos
Reposted YesterdaySaved
Remote
United States
102K-219K Annually
Junior
102K-219K Annually
Junior
Software • Quantum Computing • Metaverse • Infrastructure as a Service (IaaS)
Design, develop, and operate monitoring and observability solutions for flagship supercomputers. Build data pipelines to process large telemetry and logs, respond to incidents, implement mitigations, write postmortems, and improve reliability, performance, and operational runbooks at scale.
Top Skills: AzureCC#C++Data PipelinesJavaJavaScriptMonitoringObservabilityPythonSupercomputingTelemetry
Reposted YesterdaySaved
In-Office or Remote
4 Locations
143K-304K Annually
Senior level
143K-304K Annually
Senior level
Software • Quantum Computing • Metaverse • Infrastructure as a Service (IaaS)
Design, implement, and operate cloud-scale telemetry platforms and data pipelines for Azure Monitor. Lead cross-team architecture, mentor engineers, improve AI-enabled development practices, ensure security/compliance, and resolve complex production incidents.
Top Skills: AzureAzure Cosmos DbAzure Data FactoryAzure Event GridAzure MonitorAzure PostgresqlAzure Service BusAzure Sql DbAzure Synapse AnalyticsCC#C++ContainersGoKubernetesMicrosoft FabricMicrosoft SentinelOpentelemetryPower BIPythonRust
Reposted YesterdaySaved
Remote
US
131K-237K Annually
Expert/Leader
131K-237K Annually
Expert/Leader
Information Technology • Software
Lead architecture, design, and implementation of a cloud-native, event-driven TFDM platform using Kafka and streaming technologies. Drive AI-assisted development, containerization, IaC, CI/CD, and production reliability. Provide hands-on development, mentoring, troubleshooting, executive briefings, and oversight across integrations, security, testing, and deployments.
Top Skills: Ai Evals/GuardrailsAnsibleApache FlinkApache KafkaArgo CdArgo EventsArgo RolloutsArgo WorkflowAWSC++ChatgptClaudeClaude CodeCloudFormationContext EngineeringCucumberDockerGithub CopilotGitlabHelmIstioJavaJunitKafka ConnectKafka StreamsKsqldbKubernetesMcp-Based Tool IntegrationMockitoMulti-Agent OrchestrationOpenshiftPodmanPythonRagSeleniumTerraform
Reposted YesterdaySaved
In-Office or Remote
2 Locations
Senior level
Senior level
Retail
Design, configure, and administer the Splunk observability platform for enterprise-scale log and event ingestion. Optimize pipelines for performance, cost, and retention; integrate Splunk with cloud platforms (AWS, Azure, Kubernetes), CI/CD and deployment tools (Puppet, Ansible), and enterprise systems (ServiceNow, PagerDuty). Implement automation using Python/Bash, manage Splunk apps/roles, and support observability and APM initiatives while collaborating within SAFe using Jira and Confluence.
Top Skills: AnsibleAWSAzureBashCi/CdConfluenceJIRAKubernetesPagerdutyPuppetPythonServicenowSplunk
Reposted YesterdaySaved
Remote
United States
200K-250K Annually
Senior level
200K-250K Annually
Senior level
Artificial Intelligence • Cloud • Software • Infrastructure as a Service (IaaS)
Design, build, and operate Restate Cloud and BYOC deployments across multi-tenant SaaS and on-prem environments. Implement IaC and cloud orchestration for Kubernetes-based stateful workloads, ensure reliability and observability (SLOs, metrics, traces, logs, runbooks), automate fleet scaling, and participate in on-call rotations supporting production operations.
Top Skills: C++GoInfrastructure-As-CodeKubernetesRust
New

Cut your apply time in half.

Use ourAI Assistantto automatically fill your job applications.

Use For Free
Application Tracker Preview
Reposted YesterdaySaved
Remote
USA
1-1 Annually
Entry level
1-1 Annually
Entry level
Real Estate • Financial Services • PropTech
Contractor platform engineering role at SitusAMC supporting technology initiatives within the real estate industry. The posting provides general company information but does not specify platform engineering responsibilities, required qualifications, technologies, location, travel, or clearance requirements.
Reposted YesterdaySaved
Remote
United States
143K-200K Annually
Senior level
143K-200K Annually
Senior level
Software
Lead platform engineering focused on identity, authentication, and authorization across scalable, resilient backend services. Drive architecture, full‑stack TypeScript development, CI/CD automation, testing, on‑call service ownership, and mentor junior engineers while collaborating with cross‑functional teams to deliver cloud, hybrid, and on‑prem deployments.
Top Skills: AuthenticationAuthorizationAWSAws LambdaBitbucketCi/CdCircleCICypressDistributed SystemsJenkinsJestLinuxMochaNode.jsRestful ApisServerlessTypescript
YesterdaySaved
Remote
US
168K-194K Annually
Senior level
168K-194K Annually
Senior level
Big Data • eCommerce
Design, build, and operate shared cloud infrastructure across AWS, Kubernetes, Terraform, Databricks, and Cloudflare. Lead SRE and DevOps initiatives involving reliability, observability, CI/CD, incident response, disaster recovery, cost optimization, and developer self-service. Partner with application and data engineering teams to improve workload operations, deployment safety, and infrastructure scalability. Participate in on-call rotations and contribute to technical standards, architecture, documentation, and sustainable 24/7 operations.
Top Skills: AnthropicAWSAws CloudformationCi/CdCloudflareDatabricksDatadogInfrastructure As CodeKubernetesOpenaiTerraform
Reposted YesterdaySaved
Remote
United States
Expert/Leader
Expert/Leader
Software
Leads enterprise cloud infrastructure architecture and strategy across Azure, AWS, GCP, and hybrid environments. Establishes standards for security, reliability, observability, disaster recovery, identity governance, automation, and cost optimization. Drives cloud modernization, migrations, SaaS and AI integrations, and Infrastructure as Code practices. Acts as a principal technical authority, conducting design reviews, resolving complex cross-platform challenges, mentoring engineers, evaluating technologies, and aligning business and security requirements with scalable solutions.
Top Skills: ArmAWSBicepCi/CdCloud NetworkingDisaster RecoveryEnterprise Ai PlatformsGoogle Cloud PlatformHybrid CloudInfrastructure As CodeMicrosoft 365AzureMicrosoft Entra IdMulti-CloudObservabilityOktaPowershellPythonSaaSScimSsoTerraformZero Trust
YesterdaySaved
Remote
USA
140K-201K Annually
Senior level
140K-201K Annually
Senior level
Information Technology • Internet of Things
Designs and manages scalable GCP platforms for enterprise applications including Netcracker, Alepo, and Axiros. Responsibilities include developing Terraform infrastructure, CI/CD pipelines, upgrade automation, internal developer platforms, incident management, root cause analysis, performance optimization, rightsizing, and security guardrails. The role supports cross-functional engineering teams and participates in on-call rotations to maintain high service availability.
Top Skills: AdkAnsibleCi/CdCloud RunDatadogDynatraceElkGcp MonitoringGeminiGoogle Cloud Platform (Gcp)JavaKubernetesLinuxPrometheusPythonSaltstackTerraformVertex AiVpc
YesterdaySaved
Remote
United States
90-100 Hourly
Expert/Leader
90-100 Hourly
Expert/Leader
Mobile • Software
Leads DevOps and application operations for critical enterprise applications. Responsibilities include production reliability, SLO and SLA management, incident response, release and deployment operations, observability, automation, capacity planning, disaster recovery, security compliance, performance optimization, stakeholder service reviews, and mentoring AppOps engineers. The role supports 24x7 reliability through on-call participation and cross-functional collaboration with engineering, cloud, security, and compliance teams.
Top Skills: AWSAws CloudformationAzureAzure ArmBashBicepCi/CdDatadogDnsElkGrafanaInfrastructure As CodeOpensearchOpentelemetryPowershellPrometheusPythonSsl/TlsTerraform
All Filters
JobType
New Jobs
Job Category
Experience
Industry
Company Name
Company Size

Sign up now Access later

Create Free Account