Top Tech Jobs & Startup Jobs

3 Days AgoSaved
Hybrid
Bellevue, WA, USA
Senior level
Senior level
Agency • HR Tech • Professional Services
Designs and operates virtualization and Kubernetes orchestration platforms for large-scale GPU and HPC workloads. Responsibilities include GPU cluster provisioning, workload scheduling, resource allocation, multi-tenant infrastructure, cluster lifecycle management, automation, reliability, security, and scalability. The role partners with hardware, networking, infrastructure, and AI platform teams while owning complex systems from architecture through production operations.
Top Skills: High-Performance NetworkingInfrastructure As CodeKubernetesKubernetes Device PluginsLinuxNvidia Gpu OperatorSlurm
3 Days AgoSaved
Hybrid
Bellevue, WA, USA
Senior level
Senior level
Agency • HR Tech • Professional Services
Build and scale distributed infrastructure for large-scale AI model training across GPU clusters. Responsibilities include improving reliability, fault tolerance, checkpointing, recovery, resource utilization, training pipelines, developer tooling, and operational processes. The role partners with platform, orchestration, performance, and machine learning teams to diagnose training issues and support production-ready AI workloads.
Top Skills: Containerized Ai WorkloadsDeepspeedGpu ClustersKubernetesMegatron-LmPytorch DistributedRay
4 Days AgoSaved
Hybrid
Bellevue, WA, USA
Senior level
Senior level
Agency • HR Tech • Professional Services
Optimize GPU kernels and data-plane performance for large-scale AI training and inference. Profile workloads, identify bottlenecks, improve latency, throughput, and utilization, develop benchmarking practices, and evaluate GPU technologies across distributed infrastructure. Collaborate with AI infrastructure, machine learning, and platform engineering teams to improve GPU efficiency, scalability, and reliability.
Top Skills: Amd GpusCudaGpu Compiler TechnologiesNsight ComputeNsight SystemsNvidia GpusRocm
Reposted 4 Days AgoSaved
Hybrid
Bellevue, WA, USA
Senior level
Senior level
Agency • HR Tech • Professional Services
Build and scale distributed infrastructure for large-scale AI model training across GPU clusters. Responsibilities include improving reliability, fault tolerance, checkpointing, recovery, throughput, resource utilization, and cost efficiency. The role integrates models into production training pipelines, develops automation for AI researchers, diagnoses distributed training issues, and establishes platform reliability practices. Candidates should have experience with distributed training systems, foundation models, multi-node GPU workloads, complex distributed systems, and ML infrastructure at scale.
Top Skills: Containerized Ai WorkloadsDeepspeedDistributed SystemsGpu ClustersKubernetesMachine Learning InfrastructureMegatron-LmPytorch DistributedRay
4 Days AgoSaved
Hybrid
Bellevue, WA, USA
Expert/Leader
Expert/Leader
Agency • HR Tech • Professional Services
Lead the architecture, mechanical design, configuration, validation, and production deployment of GPU servers and rack-scale AI infrastructure. Partner with NVIDIA, AMD, ODMs, OEMs, data center engineering, networking, and operations teams. Own rack layouts, power distribution, cooling, airflow, cable management, serviceability, qualification, and hardware standards while balancing performance, reliability, manufacturability, scalability, and cost.
Top Skills: Amd GpusCable ManagementGpu ServersLiquid CoolingNvidia GpusOdm/Oem ManufacturingPower DistributionRack-Scale InfrastructureThermal Management
New

Cut your apply time in half.

Use ourAI Assistantto automatically fill your job applications.

Use For Free
Application Tracker Preview
4 Days AgoSaved
Hybrid
Bellevue, WA, USA
Entry level
Entry level
Agency • HR Tech • Professional Services
Optimize GPU kernels and data-plane performance for large-scale AI training and inference workloads. Profile bottlenecks, improve utilization, latency, throughput, and scalability, develop benchmarking practices, evaluate GPU technologies, and collaborate with infrastructure, machine learning, and platform engineering teams.
Top Skills: Amd GpusCudaGpu Compiler TechnologiesGpu Kernel DevelopmentGpu ProfilingNsight ComputeNsight SystemsNvidia GpusRocmRocm Profiling Tools
4 Days AgoSaved
Hybrid
Bellevue, WA, USA
Senior level
Senior level
Agency • HR Tech • Professional Services
Build and operate production-grade AI model-serving and inference systems. Optimize large-model workloads for throughput, latency, GPU utilization, scalability, reliability, and cost. Collaborate with training, GPU performance, orchestration, and infrastructure teams, while developing monitoring, alerting, and operational practices. Investigate performance and capacity issues and contribute to platform architecture and engineering standards.
Top Skills: BatchingCachingCloud InfrastructureDistributed SystemsGpu ComputingKubernetesQuantizationSglangTensorrt-LlmTriton Inference ServerVllm
Reposted 4 Days AgoSaved
Hybrid
Bellevue, WA, USA
Senior level
Senior level
Agency • HR Tech • Professional Services
Design, build, and operate virtualization and Kubernetes orchestration platforms for large-scale GPU and HPC workloads. Develop automated provisioning, scheduling, resource allocation, multi-tenant capacity management, and cluster lifecycle systems. Improve platform reliability, security, scalability, and operational maturity while partnering with hardware, networking, infrastructure, and AI platform teams. Own complex systems from architecture through production operation and contribute to engineering standards and platform strategy.
Top Skills: Infrastructure As CodeKubernetesKubernetes Device PluginsLinuxNvidia Gpu OperatorSlurm
17 Days AgoSaved
Hybrid
Bellevue, WA, USA
Senior level
Senior level
Agency • HR Tech • Professional Services
Conduct applied research on AI models, inference systems, accelerator ecosystems, and data center infrastructure. Evaluate emerging technologies and translate findings into engineering, product, infrastructure, financial, and commercial strategies. Own technical views on GPU hall capacity, power, cooling, density, cost, and performance. Support customer and partner discussions, infrastructure diligence, benchmarking, and technology investment decisions. The role requires hands-on inference, compiler, kernel, runtime, and multi-accelerator experience, with up to 25% international travel.
Top Skills: Ai Model ServingBatchingCudaDirect-To-Chip CoolingGpu AcceleratorsImmersion CoolingKv-Cache OptimizationLiquid CoolingQuantizationRocm/HipSglangTensorrt-LlmTritonVllmXla
17 Days AgoSaved
Hybrid
Bellevue, WA, USA
Entry level
Entry level
Agency • HR Tech • Professional Services
Evaluate and shape networking strategy for large-scale AI infrastructure. Track networking technologies, vendors, standards, and research; define scale-up and scale-out fabric positions; validate performance, cost, isolation, and observability requirements; and guide build-versus-buy decisions. Diagnose production collective communication issues, assess GPU networking and tenant-facing capabilities, support finance, sales, delivery, and diligence activities, and translate applied research into engineering and commercial decisions.
Top Skills: Co-Packaged OpticsEthernetGpu ClustersHpc FabricsInfinibandNcclNdrNetwork ObservabilityNetwork TelemetryNvlinkNvswitchOpticsRcclRocev2Spectrum-XUltra EthernetXdr
All Filters
JobType
New Jobs
Job Category
Experience
Industry
Company Name
Company Size

Sign up now Access later

Create Free Account