Top Remote Infrastructure Engineer Jobs

Reposted 12 Days AgoSaved
Remote or Hybrid
4 Locations
165K-330K Annually
Junior
165K-330K Annually
Junior
Software
Develop and maintain components for machine learning inference platform. Focus on Kubernetes deployments, resource management, and performance monitoring.
Top Skills: GoKubernetesPython
Reposted 12 Days AgoSaved
Remote
United States
Entry level
Entry level
Artificial Intelligence • Information Technology
Loti AI seeks a DevOps engineer for decentralized infrastructure to deploy, scale, and maintain distributed systems, manage IPFS clusters, and oversee blockchain nodes.
Top Skills: ArweaveCi/CdDockerIpfsKubernetesThe Graph
Reposted 12 Days AgoSaved
Remote
USA
160K-210K Annually
Junior
160K-210K Annually
Junior
Automotive • Machine Learning • Robotics • Software • Transportation
Design, operate, and scale ML data and training pipelines for autonomous driving. Build distributed training, orchestration, CI/CD and infrastructure-as-code; maintain data/metadata stores and tooling to support large-scale model training and evaluation across cloud and cluster environments.
Top Skills: C++LinuxPythonPyTorch
Reposted 12 Days AgoSaved
Remote
USA
195K-270K Annually
Senior level
195K-270K Annually
Senior level
Healthtech • Software
The Engineering Director will lead and manage the cloud infrastructure and developer enablement, focusing on AI-native systems, team development, and security operations.
Top Skills: AWSAzureCi/CdGCPTerraform
Reposted 12 Days AgoSaved
Remote
United States
150K-200K Annually
Mid level
150K-200K Annually
Mid level
Artificial Intelligence • Information Technology • Consulting
As a System Engineer, you'll design, deploy, and maintain cloud systems for AI workloads, troubleshoot complex issues, and collaborate with teams to optimize performance.
Top Skills: BashGpuInfinibandLinuxNvlinkPciePython
Reposted 12 Days AgoSaved
In-Office or Remote
10 Locations
Senior level
Senior level
Energy
Design, deploy, and maintain highly available, secure cloud platforms; support CI, automation, monitoring, and migrations of legacy systems to AWS; resolve performance, vulnerability, and scalability issues; create operational documentation and collaborate with stakeholders.
Top Skills: AnsibleAWSAzureDockerGitIso-27001LinuxPostgresTerraform
Reposted 12 Days AgoSaved
Remote
TX, USA
120K-207K Annually
Senior level
120K-207K Annually
Senior level
Artificial Intelligence • Computer Vision • Hardware • Robotics • Metaverse
Primary point-of-contact for a customer, providing on-site and remote HPC, Ethernet, and AI infrastructure support. Troubleshoot Linux systems, networking protocols, and interoperability; reproduce and resolve complex issues; collaborate with engineering, marketing, and support; document support methodologies and improve processes.
Top Skills: ArpBgpChatgptCopilotCursorDockerEthernetGeminiGleanIgmpIpKubernetesLacpLfcsLinuxMlagNvidia Ethernet SwitchingNvidia Spectrum-XOspfPimRhcsaStpTcpTcpdumpUdpWireshark
Reposted 12 Days AgoSaved
In-Office or Remote
5 Locations
184K-357K Annually
Senior level
184K-357K Annually
Senior level
Artificial Intelligence • Computer Vision • Hardware • Robotics • Metaverse
Lead the optimization and performance analysis of distributed training and inference workloads on NVIDIA GPU platforms, with responsibilities including debugging, benchmarking, and ensuring reliability of large-scale AI systems.
Top Skills: C/C++Containerized EnvironmentsCudaInfinibandMegatronNcclNemoNsight SystemsNvlinkNvswitchPciePythonPyTorchRdmaRoceTensorrt-Llm
Reposted 12 Days AgoSaved
In-Office or Remote
San Francisco, CA, USA
270K-350K Annually
Senior level
270K-350K Annually
Senior level
Artificial Intelligence • Consumer Web • Wearables
Design, build, and own scalable platform infrastructure for low-latency voice/ML services. Anticipate and eliminate scaling bottlenecks, improve developer experience and deployment speed, integrate product with ML systems (credentials, proxy, routing, experimentation), and set technical direction as the platform team grows.
Reposted 12 Days AgoSaved
In-Office or Remote
3 Locations
110K-150K Annually
Senior level
110K-150K Annually
Senior level
Agency • Software • Consulting
The role involves managing full technical lifecycle of client cloud deployments, designing integrations, maintaining infrastructures, and engaging with clients directly.
Top Skills: Api IntegrationAWSAzureGCPHelmKubernetesLlm ApisTerraform
Reposted 13 Days AgoSaved
Remote
United States
106K-223K Annually
Senior level
106K-223K Annually
Senior level
Software • Quantum Computing • Metaverse • Infrastructure as a Service (IaaS)
Design and lead cloud and AI infrastructure solutions for commercial customers: modernize platform estates, migrate workloads (Windows, Linux, Oracle), design hybrid networking and secure architectures, run workshops, PoCs and pilots, and influence Azure-based deployment decisions to enable Frontier AI transformation and compliance.
Top Skills: Agentic AssessmentAgentic ToolingAzureAzure Container AppsAzure CopilotAzure Kubernetes Service (Aks)Defender For CloudGithub CopilotLinuxNetappOraclePostgresRed HatSAPSecure RoutingSQL ServerVirtual NetworksVMwareVpnWindows Server
13 Days AgoSaved
Remote
United States
180K-220K Annually
Senior level
180K-220K Annually
Senior level
Artificial Intelligence • Information Technology • Consulting
Serves as the primary technical partner for customers deploying and operating GPU clusters and AI infrastructure. Responsibilities include configuring and tuning GPU environments, troubleshooting hardware, networking, operating system, and cluster issues, coordinating with internal engineering teams, translating requirements into architectures and execution plans, documenting solutions, and identifying performance and reliability improvements.
Top Skills: Ai InfrastructureGpu ClustersHplLinuxNcclNvidia GpusNvidia Grace Blackwell
New

Track Smarter, Apply Better.

Ditch the spreadsheets. Organize your job search with our freeApplication Tracker.

Use For Free
Application Tracker Preview
Reposted 13 Days AgoSaved
Remote
US
190K-270K Annually
Senior level
190K-270K Annually
Senior level
Artificial Intelligence
Design and build scalable backend components and indexing pipelines for semantic and hybrid retrieval, build retrieval orchestration and knowledge-graph services, improve retrieval quality via evaluation and observability, design APIs, and optimize latency, throughput, cost, reliability, and security for large-scale AI inference and retrieval workloads.
Top Skills: C++ElasticEmbeddingsGoHybrid RetrievalJavaKnowledge GraphKubernetesLlmsObservability FrameworksOpensearchPineconePulumiPythonRagRustSemantic SearchTerraformVector Databases
14 Days AgoSaved
Remote
USA
Mid level
Mid level
Agency • Information Technology • Professional Services
Build and maintain cloud infrastructure, automation, configuration management, and CI/CD pipelines. Operate Kubernetes, Helm, and Docker workloads across multiple public clouds; configure observability tooling; troubleshoot production systems spanning networking, databases, serverless functions, and caching; and participate in a 24/7 on-call rotation. The role also involves owning infrastructure projects end to end and collaborating across engineering and product teams.
Top Skills: BashCachingCi/CdConfiguration ManagementDashboardsDnsDockerHelmInfrastructure As CodeKubernetesLoggingMetricsMonitoringNosql DatabasesPowershellPublic Cloud PlatformsPythonRelational DatabasesReverse ProxiesServerless FunctionsWeb Servers
Reposted 14 Days AgoSaved
Remote
USA
Senior level
Senior level
Healthtech • Insurance • Financial Services
Lead platform reliability: define SLOs/error budgets, own observability and deploy pipelines, harden integrations with dental systems, operate LLM-driven workflows safely, build incident practices, and raise engineering reliability across the company.
Top Skills: AnthropicAWSCi/CdCrewaiDatadogDockerEcsGoGoogle Vertex AiKubernetesLangchainLlamaindexMastraNode.jsOpenaiPostgresPythonReactTerraformTypescript
Reposted 14 Days AgoSaved
In-Office or Remote
Lehi, UT, USA
Senior level
Senior level
Healthtech • Software
Design, build, and maintain scalable, resilient data platform services for data integration, event processing, and orchestration. Improve data availability, discoverability, dependability, and security tooling while collaborating across cross-functional teams and delivering production-grade solutions at scale.
Top Skills: AWSBigtableC/C++GCPGoGoogle PubsubIcebergInfrastructure As CodeJavaKafkaKubernetesNoSQLPythonS3SpannerVerticaVitess
Reposted 14 Days AgoSaved
Remote or Hybrid
3 Locations
266K-395K Annually
Senior level
266K-395K Annually
Senior level
Artificial Intelligence • Cloud • Machine Learning • Infrastructure as a Service (IaaS)
Design, develop, and maintain high-performance distributed storage software and protocol APIs (file, block, object). Build scalable, resilient storage services, integrate with hardware (NVMe, GPU-direct), troubleshoot production data center issues, and participate in full SDLC for on-prem storage solutions.
Top Skills: CC++Ci/CdDockerDpuFibre ChannelGoGpuGpu-Direct StorageInfinibandIscsiKubernetesLinux Kernel InternalsLustreNfsNvmePythonRoceS3SmbSwift
Reposted 14 Days AgoSaved
Remote or Hybrid
3 Locations
314K-465K Annually
Expert/Leader
314K-465K Annually
Expert/Leader
Artificial Intelligence • Cloud • Machine Learning • Infrastructure as a Service (IaaS)
Lead design and implementation of high-performance distributed storage systems across object, block, and file paradigms. Drive architecture, mentor engineers, integrate storage with networking/compute/DPUs, optimize protocol performance, troubleshoot production data center issues, build benchmarking and observability tooling, and collaborate on cross-functional AI infrastructure deployments.
Top Skills: BlktraceBpftraceCC++CephDaosDpdkDpus (Nvidia Bluefield)EbpfFibre ChannelFioGoGpu-Direct StorageGrafanaInfiniband)IscsiKubernetesLustreMinioNfsNvmeNvme-OfPerfPrometheusRdma (RoceRustS3SmbSpdk
Reposted 14 Days AgoSaved
In-Office or Remote
3 Locations
Senior level
Senior level
Artificial Intelligence • Cloud • Information Technology • Software
As a Software Engineer in AI Infrastructure, you will design and develop core platform components, build APIs and services, enhance performance, and automate tooling while collaborating across teams and improving system reliability.
Top Skills: AnsibleGoHelmKubernetesPythonTerraform
Reposted 14 Days AgoSaved
Remote or Hybrid
United States
Senior level
Senior level
Artificial Intelligence • Software
The Senior Solutions Engineer will design and implement infrastructure for AI and HPC workloads, engage with customers, and lead technical discovery and architecture design.
Top Skills: Ai PlatformsCloud InfrastructureGpu InfrastructureHpc SystemsKubernetesLinuxMlopsNetworkingStorage Systems
Reposted 14 Days AgoSaved
Remote
United States
180K-224K Annually
Senior level
180K-224K Annually
Senior level
Artificial Intelligence • Information Technology • Consulting
Build and operate Nebius's network infrastructure: define SLIs/SLOs, improve site and inter-site reliability, lead incident response and postmortems, develop observability and alerting, automate change workflows, and collaborate with network and platform teams to embed operability.
Top Skills: Ci/CdContainer PlatformsGoInfrastructure As CodeLinuxPython
Reposted 14 Days AgoSaved
Remote
TN, USA
108K-207K Annually
Senior level
108K-207K Annually
Senior level
Artificial Intelligence • Computer Vision • Hardware • Robotics • Metaverse
Provide onsite and remote technical support for NVIDIA Ethernet and AI infrastructure, troubleshoot Linux-based systems, debug networking and interoperability issues, collaborate with engineering and marketing, document support methodologies, and act as primary customer contact, spending at least one week per month onsite.
Top Skills: AnsibleBashBgpChatgptCopilotCursorDockerEvpnGeminiGleanKubernetesLinuxNvidia Ethernet SwitchingNvidia Spectrum-XOspfPythonQosRoceTcpdumpVxlanWiresharkYaml
Reposted 14 Days AgoSaved
In-Office or Remote
Chicago, IL, USA
Expert/Leader
Expert/Leader
Database • Analytics • Consulting
Senior individual contributor who defines multi-cloud infrastructure standards and reference architectures, authors reusable Terraform IaC, drives Kubernetes production patterns, leads CI/CD and observability practices, advises on security and cost optimization, mentors senior engineers, and shapes pre-sales technical solutioning across AWS, GCP, and Azure.
Top Skills: AcrAksArgocdArtifact RegistryAWSAzureAzure DevopsBashCloud-InitDnsEcrEksFinopsFluxGithub ActionsGitlab CiGkeGCPJenkinsKubernetesOpaPowershellPrivate Service ConnectPrivatelinkPythonSentinelService MeshShared VpcTerraformVMwareVnetVpc
15 Days AgoSaved
In-Office or Remote
24 Locations
145K-275K Annually
Expert/Leader
145K-275K Annually
Expert/Leader
Artificial Intelligence • Consumer Web • Digital Media • Software
Design, build, and operate Rust services powering Scribd’s large-scale Content Library and core infrastructure. Own features end to end, evolve Python tooling, optimize DataFusion, Parquet, S3, caching, and PostgreSQL/Aurora systems, and strengthen governance through lineage, authorization, and deletion controls. Lead cross-functional technical projects, support platform users, mentor engineers, improve developer workflows, and maintain production reliability through observability, on-call, and incident response.
Top Skills: Amazon AuroraAmazon EcsAmazon S3Apache ArrowApache ParquetArrow-RsAWSAws LambdaClaude CodeDatafusionMySQLPostgresPythonRustTerraform
15 Days AgoSaved
Remote
WA, USA
75K-100K Annually
Mid level
75K-100K Annually
Mid level
Information Technology • Consulting • Cybersecurity
Deploy and implement enterprise infrastructure solutions across compute, storage, networking, and virtualization environments. Perform rack-and-stack installations, cabling, system integration, hardware deployment, and Layer 1/Layer 2 troubleshooting. Work directly with customers at data centers and colocation facilities, document implementations, provide knowledge transfer, and independently manage deployment priorities. The role requires approximately 60% regional travel and experience with enterprise networking, storage, SAN, virtualization, and command-line infrastructure platforms.
Top Skills: AristaArubaCiscoCliConverged InfrastructureDellDhcpDnsHyper-VJuniperNutanixPrivate CloudSanVlanVMware
All Filters
JobType
New Jobs
Job Category
Experience
Industry
Company Name
Company Size

Sign up now Access later

Create Free Account