Top Tech Jobs & Startup Jobs

3 Hours AgoSaved
Remote or Hybrid
3 Locations
240K-356K Annually
Senior level
240K-356K Annually
Senior level
Artificial Intelligence • Cloud • Machine Learning • Infrastructure as a Service (IaaS)
Operate and scale bare-metal Kubernetes clusters at large scale, build control-plane services, operators, and automation for cluster lifecycle. Automate tooling in Go and Python, define SLOs/SLIs, run on-call rotations, troubleshoot incidents, and support customers integrating workloads, storage, and authentication.
Top Skills: ArgocdCi/CdCluster ApiCniCrdCsiEksFluentbitGitopsGkeGoGrafanaHelmKubeadmKubernetesKubernetes OperatorsPrometheusPython
3 Hours AgoSaved
Hybrid
2 Locations
169K-197K Annually
Senior level
169K-197K Annually
Senior level
Artificial Intelligence • Cloud • Machine Learning • Infrastructure as a Service (IaaS)
Lead FP&A for Public Cloud: own revenue forecasting, driver-based models, headcount and OpEx planning, GTM finance partnership, pricing analysis, customer segment and cohort analysis, internal allocation, and executive reporting to drive revenue, margin, and capital efficiency.
Top Skills: Ai/Ml WorkloadsCRMEnterprise Planning SystemsExcelFinancial Software SystemsGpu ComputePublic Cloud InfrastructureSQL
3 Hours AgoSaved
Remote or Hybrid
3 Locations
266K-395K Annually
Senior level
266K-395K Annually
Senior level
Artificial Intelligence • Cloud • Machine Learning • Infrastructure as a Service (IaaS)
Build and maintain scalable Kubernetes control-plane services, operators, and automation for GPU-accelerated AI workloads. Implement GPU-aware orchestration, integrate high-performance networking (InfiniBand/RDMA/GPUDirect), ensure observability (Prometheus, Grafana, tracing), develop inference platform services, create internal CLIs/tools, and support production systems via on-call rotations.
Top Skills: AksCiliumCloud InfrastructureCniContainersControllersCrdsCsiDcgmDevice PluginsDistributed TracingEksGkeGoGpudirectGrafanaInfinibandKaiKubernetesKueueLinuxMigMultusNcclNetwork OperatorNvidia Gpu OperatorOperatorsPrometheusPythonRdmaRoceSlurmVolcano
3 Hours AgoSaved
Remote or Hybrid
3 Locations
314K-465K Annually
Expert/Leader
314K-465K Annually
Expert/Leader
Artificial Intelligence • Cloud • Machine Learning • Infrastructure as a Service (IaaS)
Lead design and implementation of a highly available GPU/CPU host and instance lifecycle control plane. Drive hardware enablement, BIOS/firmware and DPU integrations, multi-tenant security, and durable execution models. Provide technical leadership across teams, translate customer and datacenter requirements into scalable infrastructure, and set engineering standards for mission-critical, large-scale AI cloud systems.
Top Skills: BiosCC++Compute Control PlaneDistributed SystemsDpu (Supernics)GoLinux Kernel InternalsPythonRustSoftware Defined NetworkingUefi
2 Days AgoSaved
Remote or Hybrid
2 Locations
266K-395K Annually
Senior level
266K-395K Annually
Senior level
Artificial Intelligence • Cloud • Machine Learning • Infrastructure as a Service (IaaS)
Design and implement a vendor-agnostic storage control plane for large-scale AI infrastructure. Build abstraction layers, reconciliation loops, CRD-based orchestration, capacity/placement engines, multi-tenant isolation, and observability to ensure performant, resilient storage across vendors and datacenters.
Top Skills: CC++CephCi/CdCustom SchedulersCxlDaosDdnDockerDpuEdsffGoGpuIbm Storage ScaleInfinibandKubernetesKubernetes Controllers/OperatorsKubernetes CrdsLinux KernelNetappNfsNumaObservability (Sli/Slo)Pcie TopologyPurePythonRoceS3Vast DataWekaZns Ssds
New

Track Smarter, Apply Better.

Ditch the spreadsheets. Organize your job search with our freeApplication Tracker.

Use For Free
Application Tracker Preview
3 Days AgoSaved
Hybrid
San Francisco, CA, USA
210K-305K Annually
Senior level
210K-305K Annually
Senior level
Artificial Intelligence • Cloud • Machine Learning • Infrastructure as a Service (IaaS)
Provide commercial legal support for Lambda's cloud business by drafting, reviewing, and negotiating customer, data center, partner, reseller, and marketing agreements; lead enterprise deal negotiations; advise cross-functionally on contract structure and risk; implement legal technology and process improvements; and develop templates and playbooks to scale contracting.
3 Days AgoSaved
Hybrid
San Jose, CA, USA
169K-225K Annually
Mid level
169K-225K Annually
Mid level
Artificial Intelligence • Cloud • Machine Learning • Infrastructure as a Service (IaaS)
Lead technical accounting research and memos for complex, non-routine U.S. GAAP transactions; analyze revenue, leases, equity, debt, business combos and stock comp; support IPO readiness, S-1/10-K prep, consolidated financials, footnotes, SOX readiness, external audits, and cross-functional accounting guidance.
Top Skills: ExcelNetSuiteWorkiva
Reposted 3 Days AgoSaved
Remote or Hybrid
2 Locations
291K-430K Annually
Senior level
291K-430K Annually
Senior level
Artificial Intelligence • Cloud • Machine Learning • Infrastructure as a Service (IaaS)
Own the billing and metering product for a usage-based GPU cloud: define metering accuracy, reconciliation, credits/invoicing lifecycle, pricing changes, public billing API, fraud and bad-debt tradeoffs, and instrumentation. Partner with finance, platform engineering, and go-to-market to ensure revenue correctness and a coherent billing experience across self-serve and sales-led motions.
Top Skills: APIsGpuSQL
Reposted 3 Days AgoSaved
Remote or Hybrid
2 Locations
291K-430K Annually
Senior level
291K-430K Annually
Senior level
Artificial Intelligence • Cloud • Machine Learning • Infrastructure as a Service (IaaS)
Own and drive the observability product across Lambda Cloud, converting telemetry into customer-facing diagnostics and internal telemetry for fleet operations. Partner with SRE, fleet engineering, console and API teams to define signals (GPU, cluster, job telemetry, InfiniBand), SLIs/SLOs, and deliver measurable features that help customers and operators debug distributed training at 64–1,024+ GPU scale.
Top Skills: APIsConsoleDatadogDcgmGpuGrafanaInfinibandMetrics PipelinesNcclPrometheusPyTorchSlasSlisSlos
Reposted 4 Days AgoSaved
Remote or Hybrid
2 Locations
338K-438K Annually
Senior level
338K-438K Annually
Senior level
Artificial Intelligence • Cloud • Machine Learning • Infrastructure as a Service (IaaS)
Own Lambda's GPU hardware product roadmap: select GPU platforms and node/cluster configurations, align partner silicon roadmaps, translate customer demand and benchmarks into fleet investment decisions, and partner across data center, supply chain, and engineering to launch, measure, and iterate hardware offerings.
Top Skills: B200Cloud InfrastructureDistributed Training InfrastructureGpu ArchitecturesH200HpcInfinibandMpiNcclNvidia GpusOdm/Oem
All Filters
JobType
New Jobs
Job Category
Experience
Industry
Company Name
Company Size

Sign up now Access later

Create Free Account