Top Remote Infrastructure Engineer Jobs

Reposted 4 Days AgoSaved
Remote
TX, USA
75K-126K Annually
Senior level
75K-126K Annually
Senior level
Insurance
Lead NOC incident monitoring, troubleshooting, and resolution across SAN/NAS/object storage, Linux/Unix, Azure/AWS and hybrid platforms. Provide Level 2 operational support, drive MTTA/MTTR improvements, automate repetitive tasks, partner with engineering teams for production readiness, and contribute to service improvement and knowledge base initiatives.
Top Skills: AnsibleAWSAzureAzure Data Explorer (Adx)BashBrocadeCi/CdCisco MdsDatadogGitGithub CopilotHitachiJenkinsMicrosoft CopilotNasNetcoolNutanixObsPowershellPrism CentralPrism ElementPurePythonRedhat LinuxScalityServicenowStorage SanTivoliUnixWindows
4 Days AgoSaved
Remote or Hybrid
Santa Clara, CA, USA
248K-391K Annually
Expert/Leader
248K-391K Annually
Expert/Leader
Artificial Intelligence • Computer Vision • Hardware • Robotics • Metaverse
Leads architecture and operation of NVIDIA’s global compute platform across bare metal, virtualized, and cloud environments. Responsibilities include managing Kubernetes, OpenShift, KubeVirt, large-scale GPU infrastructure, automated remediation, telemetry, capacity planning, infrastructure-as-code, self-service platforms, GitOps, and legacy workload migrations. The role requires influencing autonomous engineering teams, handling hardware failures and large-scale infrastructure constraints, and establishing operational maturity for frontier AI inference systems.
Top Skills: ArgocdArmAWSBare MetalCloud InfrastructureGCPGitopsGoGpusHigh-Speed Backplane NetworkingKubernetesKubevirtMicroservicesNfsv4Nvme/TcpOpenshiftOpentofuPythonTerraformVdiVirtualization
Reposted 27 Days AgoSaved
Remote
USA
Mid level
Mid level
Blockchain • Payments • Financial Services
As an Infrastructure Engineer at Tempo, you'll build and manage infrastructure to enhance engineering efficiency and tackle challenges in the devops process.
Top Skills: ArgocdEthereumGoGrafanaHelmKubernetesLinuxPrometheusPythonRustTerraform
Reposted One Month AgoSaved
Easy Apply
Remote or Hybrid
Boston, MA, USA
Easy Apply
180K-210K Annually
Mid level
180K-210K Annually
Mid level
Healthtech • Software
The Software Engineer, Infrastructure will design, build, and operate core platform services while ensuring reliability, security, and efficiency through collaboration with various teams and improving developer experience.
Top Skills: Ci/CdGoGoogle Cloud PlatformInfrastructure-As-CodeJavaScriptKubernetesPythonTerraform
Reposted One Month AgoSaved
Easy Apply
Remote
USA
Easy Apply
218K-257K Annually
Expert/Leader
218K-257K Annually
Expert/Leader
Artificial Intelligence • Blockchain • Fintech • Financial Services • Cryptocurrency • NFT • Web3
Lead design and operation of test infrastructure at scale: test orchestration, smart selection, parallel sharding, flaky-test detection, observability, SLOs, and on-call. Define technical strategy, drive cross-team projects, mentor engineers, and reduce test feedback time to accelerate engineering velocity.
Top Skills: GeminiGleanGoLibrechat
Reposted One Month AgoSaved
Easy Apply
Remote or Hybrid
5 Locations
Easy Apply
127K-249K Annually
Senior level
127K-249K Annually
Senior level
Big Data • Cloud • Software • Database
The Security Software Engineer will design and implement security controls for MongoDB Atlas, collaborating across engineering teams and ensuring adherence to high security standards.
Top Skills: ApparmorC/C++CgroupsEbpfGoGrafanaJavaKubernetesPythonRustSeccompSelinuxSplunkTerraformVictoria Metrics
Reposted One Month AgoSaved
Easy Apply
Remote or Hybrid
5 Locations
Easy Apply
127K-249K Annually
Senior level
127K-249K Annually
Senior level
Big Data • Cloud • Software • Database
The Senior Site Reliability Engineer will lead security design and implementation for cloud infrastructures, mentor teams, and automate security solutions.
Top Skills: AnsibleAWSAzureCloud Security ToolsCloudFormationGCPGoTerraform
28 Days AgoSaved
Remote
CA, USA
224K-431K Annually
Senior level
224K-431K Annually
Senior level
Artificial Intelligence • Computer Vision • Hardware • Robotics • Metaverse
Build, scale, and harden deep learning training infrastructure for multi-thousand GPU clusters. Improve data loaders, distributed training, scheduling, and performance monitoring. Develop fault-resilient orchestration, training pipelines for massive video datasets, and collaborate with researchers and platform teams to maximize training efficiency, availability, and scalability.
Top Skills: DdpFsdpGpuInfinibandKubernetesLustreNcclPipeline ParallelismPythonPyTorchRoceSlurmTensor Parallelism
7 Days AgoSaved
In-Office or Remote
Portland, ME, USA
Senior level
Senior level
Fintech • Payments
Modernize SQL Server systems by refactoring stored procedures, optimizing queries, designing migrations, and implementing event-driven patterns. Build AI-ready data infrastructure, including embedding pipelines, vector search, RAG retrieval services, synchronization workflows, and evaluation monitoring. Develop ETL/ELT pipelines across relational, cloud, and NoSQL platforms, supported by infrastructure-as-code, data quality controls, testing, and observability. Collaborate with engineering teams, troubleshoot production issues, document solutions, mentor engineers, and use AI coding assistants to accelerate modernization and development.
Top Skills: Api IntegrationArmAWSAzureAzure Ai SearchBicepCdcCi/CdClaude CodeCosmos DbCursorEltEmbeddingsETLEvent StreamingGithub CopilotHybrid RetrievalKafkaMongoDBOpensearchPgvectorPineconePostgresRagSemantic SearchSnowflakeSQLSQL ServerT-SqlTerraformTransactional OutboxVector DatabasesVector Indexes
One Month AgoSaved
Remote or Hybrid
New York, NY, USA
180K-250K Annually
Senior level
180K-250K Annually
Senior level
Information Technology • Software • Design
Operate and maintain Monad's globally distributed validator, full node, and archive fleets; own infrastructure-as-code, observability, and release automation; build AI-driven agentic operations and guardrails; harden services, manage secrets, and codify runbooks and tooling to reduce toil and support safe, automated operation of production node and model workloads.
Top Skills: Ai AgentsAnsibleAtlantisBashFluxGitopsGrafanaKubernetesLinuxLlmsLokiPrometheusPythonSshSystemdTerraform
Entry level
Software
Design and implement infrastructure services for a GPU-as-a-Service platform. Build REST and gRPC APIs, bare-metal provisioning workflows, Kubernetes cluster lifecycle services, reconciliation loops, reliability mechanisms, and interfaces for hardware and cluster health. The role requires Go, Kubernetes controller patterns, bare-metal provisioning, asynchronous systems, and experience with infrastructure technologies such as Metal3, Cluster API, and Temporal.
Top Skills: ArgocdBmcCluster ApiFluxGitopsGoGrpcIpxeK0RdentKubernetesMetal3OpentofuPxeRedfishRestTemporalTerraform
One Month AgoSaved
In-Office or Remote
10 Locations
Senior level
Senior level
Financial Services
Lead and operate Kubernetes-based infrastructure (EKS/Porter), design CI/CD pipelines, ensure zero-downtime and high availability, and build observability and incident response practices. Collaborate with BizOps and data teams to support a modern lakehouse data platform and optimize SQL and data pipelines for performance and cost.
Top Skills: Apache IcebergAws GlueCdcCi/CdCoalesce.IoDbtDebeziumEksEstuaryKubernetesLakehousePci DssPorterSnowflakeSoc 2SQL
New

Cut your apply time in half.

Use ourAI Assistantto automatically fill your job applications.

Use For Free
Application Tracker Preview
7 Days AgoSaved
Remote
4 Locations
Entry level
Entry level
Artificial Intelligence • Information Technology
Build large-scale data pipelines and infrastructure for frontier AI model training. Develop data-processing models such as classifiers, quality filters, and labeling systems; design deduplication, quality scoring, labeling, and augmentation strategies; and create reliable tooling for researchers to explore and train on massive datasets. The role also evaluates how data quality and composition affect model outcomes, with web crawler experience as a bonus.
Top Skills: Kubernetes
Reposted One Month AgoSaved
Easy Apply
Remote
USA
Easy Apply
218K-257K Annually
Senior level
218K-257K Annually
Senior level
Artificial Intelligence • Blockchain • Fintech • Financial Services • Cryptocurrency • NFT • Web3
Own reliability, monitoring, and incident response for AI infrastructure; build automation and CI/CD tooling; manage Kubernetes/Docker production workloads; partner with infrastructure, security, and compliance; improve observability and documentation; develop internal full‑stack tooling in Go or Python.
Top Skills: AnsibleAWSBashChefCi/CdDockerEc2GitGoKubernetesLinuxLog AggregationNetwork SecurityPuppetPythonRubySaltTerraform
Reposted One Month AgoSaved
Easy Apply
Remote
USA
Easy Apply
186K-219K Annually
Senior level
186K-219K Annually
Senior level
Artificial Intelligence • Blockchain • Fintech • Financial Services • Cryptocurrency • NFT • Web3
Senior SRE on the IT Operations team owning reliability, monitoring, and incident response for AI infrastructure. Build automation, CI/CD and Kubernetes tooling, improve observability and documentation, and develop internal full-stack tools using Go or Python. Partner with Infrastructure, Security, and Compliance to scale secure, resilient AI deployment pipelines.
Top Skills: AnsibleAWSBashChefCi/CdDockerEc2GitGoKubernetesLinuxPuppetPythonRubySaltTerraform
7 Days AgoSaved
Remote
United States
150K-225K Annually
Senior level
150K-225K Annually
Senior level
Healthtech • Insurance
Build and maintain CI/CD pipelines, cloud infrastructure, Terraform automation, database systems, monitoring, and observability on Google Cloud. Support reliability, security, compliance, incident response, and shared on-call operations while applying backend engineering expertise to production services. The role manages PostgreSQL, Redis, BigQuery, Bigtable, and Firestore environments and contributes to disaster recovery and operational tooling.
Top Skills: BigQueryBigtableCloudsmithDockerFirestoreGitGoGoogle Cloud PlatformHcp TerraformIncident.IoKubernetesLinuxmacOSPostgresPythonRedisSentryTerraform
8 Days AgoSaved
Remote
USA
160K-190K Annually
Senior level
160K-190K Annually
Senior level
Energy
Own and evolve Voltus’s infrastructure platform, including Kubernetes and Nomad workloads, AWS architecture, identity and secrets, stateful systems, observability, infrastructure as code, CI/CD, and developer tooling. Lead business-critical migrations with minimal downtime, establish reliable deployment and rollback practices, mentor engineers, participate in on-call, and build secure infrastructure supporting AI-assisted development.
Top Skills: Amazon CognitoAmazon EksAmazon MskArgo CdAuth0AWSAws BedrockAws Control TowerAws OrganizationsBuildkiteC++DnsDockerElasticsearchFluxGitGitopsGoGrafanaHashicorp ConsulHashicorp NomadHashicorp VaultHelmIamJavaJenkinsKeycloakKubernetesMessage BrokersOidcOktaOpensearchOpentelemetryPrivatelinkPrometheusPythonRelational DatabasesService MeshSsoTerraformTime-Series DatabasesTransit GatewayVpc
8 Days AgoSaved
Remote or Hybrid
2 Locations
160K-322K Annually
Senior level
160K-322K Annually
Senior level
Artificial Intelligence • Computer Vision • Hardware • Robotics • Metaverse
Deploys and validates complete AI infrastructure software stacks on multi-node GPU systems, then converts implementations into technical documentation, automation, demos, training, and reference architectures. Tests prerelease software, evaluates interoperability and operational resilience, supports partners and field teams, collaborates with engineering and open-source communities, and recommends product improvements based on customer feedback. The role presents solutions through briefings, workshops, webinars, events, and internal training, with some travel required.
Top Skills: Ai InferenceAi TrainingAPIsBare-Metal ProvisioningBluefield DpuCertificate ManagementCi/CdCloud-NativeConfiguration ManagementContainersDgx CloudDocaEthernetGitopsGpu SystemsHelmHpcIdentity ManagementInfinibandInfrastructure As CodeKubernetesLinuxMulti-TenancyNvidia Ai EnterpriseNvidia DgxNvidia DsxObservabilityPythonSecrets ManagementShell ScriptingSlurmStorageTelemetry
8 Days AgoSaved
Remote
35 Locations
140K-220K Annually
Expert/Leader
140K-220K Annually
Expert/Leader
Artificial Intelligence • Software • Generative AI • Infrastructure as a Service (IaaS)
Own and scale Nango’s cloud platform, including compute, networking, databases, Kubernetes, IaC, and GitOps. Lead BYOC and private-cloud deployments, improve provisioning and observability, establish SLOs and incident processes, and manage Postgres, Redis, Elasticsearch, and ClickHouse at scale. Drive infrastructure security and compliance for SOC 2, GDPR, HIPAA, and enterprise reviews while shaping product strategy and supporting a remote-first engineering team.
Top Skills: AWSAzureClickhouseElasticsearchGCPGitopsKubernetesNode.jsPostgresRedisTerraformTypescript
Reposted One Month AgoSaved
In-Office or Remote
7 Locations
Entry level
Entry level
Cloud • Information Technology • Software • Infrastructure as a Service (IaaS)
The Infrastructure Engineer will build system-level software, focusing on distributed systems and OS level primitives while ensuring scalability and performance of services.
Top Skills: GoGrpcRust
9 Days AgoSaved
Remote
United States
209K-314K Annually
Senior level
209K-314K Annually
Senior level
Information Technology • Software
Provides global presales technical expertise for IT virtualization and infrastructure modernization solutions. Partners with account executives, engages customers, leads architecture discussions, presents solutions, advises on virtualization and migration strategies, develops service scopes, and supports pricing and delivery alignment. Requires deep virtualization architecture experience across VMware or related platforms, cloud, containers, security, AI, and infrastructure modernization.
Top Skills: Artificial IntelligenceContainersHpe Morpheus Vm EssentialsHybrid CloudHyperscalersIt Infrastructure SecurityNutanixOpen-Source ComponentsOperating SystemsRed HatVMware
10 Days AgoSaved
Remote
US
168K-194K Annually
Senior level
168K-194K Annually
Senior level
Big Data • eCommerce
Design, build, and operate shared cloud infrastructure across AWS, Kubernetes, Terraform, Databricks, and Cloudflare. Lead SRE and DevOps initiatives involving reliability, observability, CI/CD, incident response, disaster recovery, cost optimization, and developer self-service. Partner with application and data engineering teams to improve workload operations, deployment safety, and infrastructure scalability. Participate in on-call rotations and contribute to technical standards, architecture, documentation, and sustainable 24/7 operations.
Top Skills: AnthropicAWSAws CloudformationCi/CdCloudflareDatabricksDatadogInfrastructure As CodeKubernetesOpenaiTerraform
10 Days AgoSaved
Remote or Hybrid
8 Locations
110K-300K Annually
Mid level
110K-300K Annually
Mid level
Financial Services
Build and maintain scalable financial data infrastructure, including ingestion, validation, normalization, and integrations. Develop systems that calculate portfolio performance metrics efficiently across high-volume datasets, improve reliability and observability, and support distributed processing. Create internal tools, dashboards, reusable schemas, asset data feeds, and self-service analytics capabilities. Collaborate cross-functionally to expand market-data coverage, support B2B integrations, and strengthen operational resilience.
Top Skills: Business Intelligence DashboardsData InfrastructureData Ingestion PipelinesData ValidationDistributed Systems
Reposted 10 Days AgoSaved
Remote
2 Locations
85K-111K Annually
Senior level
85K-111K Annually
Senior level
Information Technology
Design, deploy, and manage enterprise digital fax and telephony infrastructure (RightFax/XM Fax, SBCs, SIP Trunks, G.711/T.38). Administer Windows Server, IIS, Active Directory/Azure AD, Azure, and VMware. Ensure security, compliance, high availability, and performance. Lead projects, mentor engineers, troubleshoot complex issues, and produce operational documentation while participating in Agile processes.
Top Skills: Active DirectoryConditional AccessG.711Group PolicyIisAzureMicrosoft Entra Id (Azure Ad)Privileged Account ManagementRightfaxRole-Based Access Control (Rbac)Security AuditingSession Border Controller (Sbc)Sip TrunkingT.38VMwareWindows ServerXm Fax
Reposted 10 Days AgoSaved
Remote or Hybrid
3 Locations
Senior level
Senior level
Artificial Intelligence • Big Data • Cloud • Software • Analytics • Infrastructure as a Service (IaaS) • Big Data Analytics
Design, build, and operate foundational, multi-cloud platform systems. Own platform strategy, make build vs. buy decisions, write design docs and code, ensure reliability at scale, document decisions and run architectural forums, and improve operational maturity for enterprise-grade SaaS platform services.
Top Skills: Apache AirflowAWSAzureGCPGoKubernetes
All Filters
JobType
New Jobs
Job Category
Experience
Industry
Company Name
Company Size

Sign up now Access later

Create Free Account