Maximum of 25 job preferences reached.
Top Engineering Jobs
Artificial Intelligence • Cloud • Software • Infrastructure as a Service (IaaS)
Own the technical lifecycle and operational health of Runpod’s global high-density GPU fleet. Responsibilities include hardware validation and benchmarking, fleet monitoring, SLA and downtime auditing, network and systems troubleshooting, incident coordination, partner technical support, AI-assisted operational automation, and infrastructure tooling. The role requires Linux, Docker, NVIDIA GPU, datacenter networking, and performance-tuning expertise, with potential future on-call participation.
Top Skills:
Ai AgentsBashDatadogDockerGoGrafanaInfinibandLinuxLlmsNvidia Software StackPrometheusPythonRdmaRoce
Artificial Intelligence • Cloud • Software • Infrastructure as a Service (IaaS)
Design, build, and maintain cloud infrastructure software (primarily in Go) to manage VMs across global data centers. Collaborate on product requirements, participate in code reviews, optimize performance and reliability, contribute to architecture decisions, and stay current with industry trends.
Top Skills:
CDockerGoHypervisorJavaScriptKernel/Driver DebuggingLinuxLxcPythonPython Machine Learning LibrariesRustTypescriptVirtual Machines
Artificial Intelligence • Cloud • Software • Infrastructure as a Service (IaaS)
Build developer tooling, demos, and integrations for Runpod; produce technical content and run scalable programs (content, events, community); attend and speak at developer/AI conferences; surface product feedback and ship solutions autonomously to improve developer onboarding and platform adoption.
Top Skills:
Agent FrameworksAi Libraries And FrameworksGoGpu ComputeJavaScriptMl InfrastructurePythonSdks
Artificial Intelligence • Cloud • Software • Infrastructure as a Service (IaaS)
Ensure stability and resilience of Runpod's distributed AI platform by defining SLIs/SLOs, leading incident response, building observability and reliability tooling, automating operational workflows, and partnering with engineering teams to reduce toil and improve production readiness.
Top Skills:
BashCi/CdContainerized Production SystemsGoGpu Observability ToolingGrafanaInfrastructure As CodeLinuxPrometheusPython
Artificial Intelligence • Cloud • Software • Infrastructure as a Service (IaaS)
Lead and scale core cloud and bare-metal infrastructure including SRE, global and HPC networking, and distributed storage. Define SLOs/SLAs, incident response, observability, and IaC. Architect InfiniBand/RoCE networks and high-performance storage to support large GPU workloads. Hire and mentor managers and senior ICs, partner with product and program teams to forecast capacity and drive reliability, throughput, and low-latency infrastructure at scale.
Top Skills:
AnsibleBare-MetalBgpCephContainer OrchestrationInfinibandInfrastructure As CodeKubernetesLustreNvlinkNvme-OfObservabilityRdmaRoceSpine-Leaf ArchitectureTerraformWeka
Artificial Intelligence • Cloud • Software • Infrastructure as a Service (IaaS)
Lead a product engineering team to own roadmaps, design scalable cloud-native systems, ship customer-facing features, ensure quality and reliability, hire and mentor engineers, and coordinate cross-functional delivery.
Top Skills:
APIsContainer RuntimesControl PlanesData StoresDockerEventingGoGpu/Accelerator WorkflowsKubernetesLinuxMicroservicesNetworkingOrchestrationPythonStorageTypescript
New
Track Smarter, Apply Better.
Ditch the spreadsheets. Organize your job search with our freeApplication Tracker.
Use For Free
Let Your Resume Do The Work
Upload your resume to be matched with jobs you're a great fit for.
Success! We'll use this to further personalize your experience.
All Filters
Total selected ()
No Results
No Results


