Top Tech Jobs & Startup Jobs

Reposted YesterdaySaved
Hybrid
Bellevue, WA, USA
Senior level
Senior level
Agency • HR Tech • Professional Services
Design, build, and operate GPU-focused virtualization and Kubernetes orchestration systems for multi-tenant AI/HPC clusters. Implement automated provisioning, workload scheduling, resource management, and tooling to scale, secure, and run GPU compute reliably across production environments while partnering with hardware, networking, and platform teams.
Top Skills: Container OrchestrationGpu Cluster ProvisioningInfrastructure As CodeKubernetesKubernetes Device PluginsLinuxNvidia Gpu OperatorSlurmVirtualization
Reposted YesterdaySaved
Hybrid
Bellevue, WA, USA
Senior level
Senior level
Agency • HR Tech • Professional Services
Build and operate production-grade model-serving and inference systems focused on high throughput, low latency, GPU efficiency, scalability, and reliability. Collaborate with GPU performance, training, orchestration, and infrastructure teams to optimize inference performance, design scalable systems, implement monitoring/alerting, and resolve production capacity and reliability issues.
Top Skills: BatchingCachingCloud InfrastructureGpu ComputingGpu SchedulingKubernetesQuantizationTensorrt-LlmTriton Inference ServerVllm
Reposted YesterdaySaved
Hybrid
Bellevue, WA, USA
Senior level
Senior level
Agency • HR Tech • Professional Services
Develop and maintain the company’s technical perspective on AI data center infrastructure, including power, cooling, density, rack architecture, siting, and economics. Evaluate emerging technologies, vendors, sites, and infrastructure partners; build defensible product and engineering recommendations; own technical reference views for GPU halls; develop cost models; and influence finance, sales, product, engineering, and investment decisions. The role is a senior individual contributor requiring hands-on experience with large-scale data centers, GPU environments, power procurement, cooling, and infrastructure economics.
Top Skills: 800 VdcAi Inference TechnologiesBackup GenerationData Center InfrastructureDirect-To-Chip CoolingGpu-Based ComputeImmersion CoolingLiquid Cooling
Reposted YesterdaySaved
Hybrid
Bellevue, WA, USA
Entry level
Entry level
Agency • HR Tech • Professional Services
Evaluate and shape networking strategy for large-scale AI infrastructure. Track networking technologies, vendors, standards, and research; define scale-up and scale-out fabric positions; validate performance, cost, isolation, and observability requirements; and guide build-versus-buy decisions. Diagnose production collective communication issues, assess GPU networking and tenant-facing capabilities, support finance, sales, delivery, and diligence activities, and translate applied research into engineering and commercial decisions.
Top Skills: Co-Packaged OpticsEthernetGpu ClustersHpc FabricsInfinibandNcclNdrNetwork ObservabilityNetwork TelemetryNvlinkNvswitchOpticsRcclRocev2Spectrum-XUltra EthernetXdr
Reposted YesterdaySaved
Hybrid
Bellevue, WA, USA
Senior level
Senior level
Agency • HR Tech • Professional Services
Build and scale distributed infrastructure for large-scale AI model training across GPU clusters. Responsibilities include improving reliability, fault tolerance, checkpointing, recovery, throughput, resource utilization, and cost efficiency. The role integrates models into production training pipelines, develops automation for AI researchers, diagnoses distributed training issues, and establishes platform reliability practices. Candidates should have experience with distributed training systems, foundation models, multi-node GPU workloads, complex distributed systems, and ML infrastructure at scale.
Top Skills: Containerized Ai WorkloadsDeepspeedDistributed SystemsGpu ClustersKubernetesMachine Learning InfrastructureMegatron-LmPytorch DistributedRay
New

Cut your apply time in half.

Use ourAI Assistantto automatically fill your job applications.

Use For Free
Application Tracker Preview
Reposted YesterdaySaved
Hybrid
Bellevue, WA, USA
Expert/Leader
Expert/Leader
Agency • HR Tech • Professional Services
Lead design of GPU servers and rack-scale infrastructure for AI workloads, partnering with GPU vendors, ODMs/OEMs, and data center teams. Own mechanical layout, power distribution, thermal design, cable management, validation, and production readiness while balancing performance, reliability, serviceability, and cost. Influence hardware standards and strategy across large-scale deployments.
Top Skills: AirflowAmdCable ManagementData Center EngineeringGpu Server ArchitectureGpusHigh-Density RacksHpcLiquid CoolingNvidiaOdmOemPower DistributionRack-Scale InfrastructureThermal Management
Reposted YesterdaySaved
Hybrid
Bellevue, WA, USA
Senior level
Senior level
Agency • HR Tech • Professional Services
Design, build, and operate virtualization and Kubernetes orchestration platforms for large-scale GPU and HPC workloads. Develop automated provisioning, scheduling, resource allocation, multi-tenant capacity management, and cluster lifecycle systems. Improve platform reliability, security, scalability, and operational maturity while partnering with hardware, networking, infrastructure, and AI platform teams. Own complex systems from architecture through production operation and contribute to engineering standards and platform strategy.
Top Skills: Infrastructure As CodeKubernetesKubernetes Device PluginsLinuxNvidia Gpu OperatorSlurm
24 Days AgoSaved
Hybrid
Bellevue, WA, USA
Expert/Leader
Expert/Leader
Agency • HR Tech • Professional Services
Lead the engineering organization responsible for bringing data center hardware into production-ready AI and HPC infrastructure. Own automated provisioning, Linux deployment, configuration, validation, GPU clusters, Kubernetes environments, monitoring, and workload readiness. Establish standards across servers, GPUs, networking, storage, firmware, and software automation while partnering with hardware, network, SRE, data center operations, and software teams. Build and develop a high-performing engineering organization capable of scaling infrastructure across thousands of servers or GPUs.
Top Skills: AnsibleBashBmcContainersCudaDcgmDistributed SystemsEthernetForemanGpu OperatorInfinibandIpmiIronicKubernetesLinuxMaasNvidia DriversNvmlPxeRdmaRedfishRoceSlurmTerraformXcat
One Month AgoSaved
Hybrid
Bellevue, WA, USA
Expert/Leader
Expert/Leader
Agency • HR Tech • Professional Services
The CIO will build and lead the enterprise technology function for a rapidly scaling AI company. Responsibilities include enterprise technology strategy, architecture, ERP and business systems, AI-enabled automation, IT operations, governance, controls, vendor management, budgeting, and cost optimization. The leader will recruit and develop the technology organization, oversee implementation partners, establish scalable systems and standards, and partner with executive leadership to align technology investments with business growth and public-company readiness.
Top Skills: Agentic WorkflowsAIApi-First ArchitectureAutomationCloud InfrastructureData CentersErpHrisIdentity And Access ManagementIt General ControlsMaster Data ManagementProcurement SystemsService ManagementSox
Reposted One Month AgoSaved
In-Office
Indianapolis, IN, USA
Senior level
Senior level
Agency • HR Tech • Professional Services
The IT Systems Engineer will manage Windows Server environments, VMware, Citrix infrastructure, and Ivanti Endpoint Management, while providing technical support and implementing security practices in an enterprise environment.
Top Skills: Active DirectoryCitrixCitrix NetscalerCitrix Virtual AppsDhcpDnsGroup PolicyIvanti Endpoint ManagementMicrosoft IntunePowershellPythonSolarwindsVmware EsxiVsphereWindows ServerXenappXendesktop
All Filters
JobType
New Jobs
Job Category
Experience
Industry
Company Name
Company Size

Sign up now Access later

Create Free Account