Maximum of 25 job preferences reached.
Top DevOps & Platform Engineering Jobs in Lexington, KY
Software • Quantum Computing • Metaverse • Infrastructure as a Service (IaaS)
Lead and develop Site Reliability Engineering teams operating Microsoft Substrate services in regulated environments. Own reliability, availability, incident response, disaster recovery, SLOs, operational metrics, automation, and compliance. Serve in the on-call rotation, drive systemic post-incident improvements, conduct recovery exercises, and partner with engineering, product, security, infrastructure, and compliance teams to build resilient, auditable cloud services.
Top Skills:
AutomationCloud ComputingDisaster RecoveryDistributed SystemsMicrosoft 365AzureMicrosoft SubstrateSlisSlos
Software • Quantum Computing • Metaverse • Infrastructure as a Service (IaaS)
Design, develop, and operate monitoring and observability solutions for flagship supercomputers. Build data pipelines to process large telemetry and logs, respond to incidents, implement mitigations, write postmortems, and improve reliability, performance, and operational runbooks at scale.
Top Skills:
AzureCC#C++Data PipelinesJavaJavaScriptMonitoringObservabilityPythonSupercomputingTelemetry
Software • Quantum Computing • Metaverse • Infrastructure as a Service (IaaS)
Design, implement, and operate cloud-scale telemetry platforms and data pipelines for Azure Monitor. Lead cross-team architecture, mentor engineers, improve AI-enabled development practices, ensure security/compliance, and resolve complex production incidents.
Top Skills:
AzureAzure Cosmos DbAzure Data FactoryAzure Event GridAzure MonitorAzure PostgresqlAzure Service BusAzure Sql DbAzure Synapse AnalyticsCC#C++ContainersGoKubernetesMicrosoft FabricMicrosoft SentinelOpentelemetryPower BIPythonRust
Information Technology • Software
Lead architecture, design, and implementation of a cloud-native, event-driven TFDM platform using Kafka and streaming technologies. Drive AI-assisted development, containerization, IaC, CI/CD, and production reliability. Provide hands-on development, mentoring, troubleshooting, executive briefings, and oversight across integrations, security, testing, and deployments.
Top Skills:
Ai Evals/GuardrailsAnsibleApache FlinkApache KafkaArgo CdArgo EventsArgo RolloutsArgo WorkflowAWSC++ChatgptClaudeClaude CodeCloudFormationContext EngineeringCucumberDockerGithub CopilotGitlabHelmIstioJavaJunitKafka ConnectKafka StreamsKsqldbKubernetesMcp-Based Tool IntegrationMockitoMulti-Agent OrchestrationOpenshiftPodmanPythonRagSeleniumTerraform
Reposted YesterdaySaved
Retail
Design, configure, and administer the Splunk observability platform for enterprise-scale log and event ingestion. Optimize pipelines for performance, cost, and retention; integrate Splunk with cloud platforms (AWS, Azure, Kubernetes), CI/CD and deployment tools (Puppet, Ansible), and enterprise systems (ServiceNow, PagerDuty). Implement automation using Python/Bash, manage Splunk apps/roles, and support observability and APM initiatives while collaborating within SAFe using Jira and Confluence.
Top Skills:
AnsibleAWSAzureBashCi/CdConfluenceJIRAKubernetesPagerdutyPuppetPythonServicenowSplunk
Artificial Intelligence • Cloud • Software • Infrastructure as a Service (IaaS)
Design, build, and operate Restate Cloud and BYOC deployments across multi-tenant SaaS and on-prem environments. Implement IaC and cloud orchestration for Kubernetes-based stateful workloads, ensure reliability and observability (SLOs, metrics, traces, logs, runbooks), automate fleet scaling, and participate in on-call rotations supporting production operations.
Top Skills:
C++GoInfrastructure-As-CodeKubernetesRust
New
Cut your apply time in half.
Use ourAI Assistantto automatically fill your job applications.
Use For Free
Real Estate • Financial Services • PropTech
Contractor platform engineering role at SitusAMC supporting technology initiatives within the real estate industry. The posting provides general company information but does not specify platform engineering responsibilities, required qualifications, technologies, location, travel, or clearance requirements.
Software
Lead platform engineering focused on identity, authentication, and authorization across scalable, resilient backend services. Drive architecture, full‑stack TypeScript development, CI/CD automation, testing, on‑call service ownership, and mentor junior engineers while collaborating with cross‑functional teams to deliver cloud, hybrid, and on‑prem deployments.
Top Skills:
AuthenticationAuthorizationAWSAws LambdaBitbucketCi/CdCircleCICypressDistributed SystemsJenkinsJestLinuxMochaNode.jsRestful ApisServerlessTypescript
Big Data • eCommerce
Design, build, and operate shared cloud infrastructure across AWS, Kubernetes, Terraform, Databricks, and Cloudflare. Lead SRE and DevOps initiatives involving reliability, observability, CI/CD, incident response, disaster recovery, cost optimization, and developer self-service. Partner with application and data engineering teams to improve workload operations, deployment safety, and infrastructure scalability. Participate in on-call rotations and contribute to technical standards, architecture, documentation, and sustainable 24/7 operations.
Top Skills:
AnthropicAWSAws CloudformationCi/CdCloudflareDatabricksDatadogInfrastructure As CodeKubernetesOpenaiTerraform
Software
Leads enterprise cloud infrastructure architecture and strategy across Azure, AWS, GCP, and hybrid environments. Establishes standards for security, reliability, observability, disaster recovery, identity governance, automation, and cost optimization. Drives cloud modernization, migrations, SaaS and AI integrations, and Infrastructure as Code practices. Acts as a principal technical authority, conducting design reviews, resolving complex cross-platform challenges, mentoring engineers, evaluating technologies, and aligning business and security requirements with scalable solutions.
Top Skills:
ArmAWSBicepCi/CdCloud NetworkingDisaster RecoveryEnterprise Ai PlatformsGoogle Cloud PlatformHybrid CloudInfrastructure As CodeMicrosoft 365AzureMicrosoft Entra IdMulti-CloudObservabilityOktaPowershellPythonSaaSScimSsoTerraformZero Trust
Information Technology • Internet of Things
Designs and manages scalable GCP platforms for enterprise applications including Netcracker, Alepo, and Axiros. Responsibilities include developing Terraform infrastructure, CI/CD pipelines, upgrade automation, internal developer platforms, incident management, root cause analysis, performance optimization, rightsizing, and security guardrails. The role supports cross-functional engineering teams and participates in on-call rotations to maintain high service availability.
Top Skills:
AdkAnsibleCi/CdCloud RunDatadogDynatraceElkGcp MonitoringGeminiGoogle Cloud Platform (Gcp)JavaKubernetesLinuxPrometheusPythonSaltstackTerraformVertex AiVpc
Mobile • Software
Leads DevOps and application operations for critical enterprise applications. Responsibilities include production reliability, SLO and SLA management, incident response, release and deployment operations, observability, automation, capacity planning, disaster recovery, security compliance, performance optimization, stakeholder service reviews, and mentoring AppOps engineers. The role supports 24x7 reliability through on-call participation and cross-functional collaboration with engineering, cloud, security, and compliance teams.
Top Skills:
AWSAws CloudformationAzureAzure ArmBashBicepCi/CdDatadogDnsElkGrafanaInfrastructure As CodeOpensearchOpentelemetryPowershellPrometheusPythonSsl/TlsTerraform
Let Your Resume Do The Work
Upload your resume to be matched with jobs you're a great fit for.
Success! We'll use this to further personalize your experience.
All Filters
Total selected ()
No Results
No Results




















