Top Tech Jobs & Startup Jobs

14 Days AgoSaved
In-Office or Remote
28 Locations
Senior level
Senior level
Artificial Intelligence • Information Technology • Consulting
Develop and optimize GPU clusters and InfiniBand networks for hyperscale HPC environments. Responsibilities include troubleshooting hardware and infrastructure issues, integrating new GPUs through Kubernetes, QEMU, and KVM, managing GPU devices and InfiniBand fabrics, tuning performance, and improving automated monitoring and fault resolution.
Top Skills: CC++GoGpuInfinibandKubernetesKvmLinuxLinux KernelMpiNcclNicsPciePythonPyTorchQemuRdmaRoceSoftware-Defined NetworkingTensorFlow
15 Days AgoSaved
Remote
28 Locations
Senior level
Senior level
Artificial Intelligence • Information Technology • Consulting
Conduct frontier applied research on efficient AI model architectures, including sparse attention, long-context memory, selective computation, dynamic inference, reasoning, and continual adaptation. Formulate research questions, design rigorous experiments at meaningful scales, develop efficient implementations, publish findings, contribute to open-source tools and models, mentor researchers, and help direct the research stream.
Top Skills: Attention MechanismsDistributed TrainingEfficient InferencePythonTransformer Architectures
15 Days AgoSaved
In-Office or Remote
9 Locations
141K-176K Annually
Entry level
141K-176K Annually
Entry level
Artificial Intelligence • Information Technology • Consulting
Develop and manage AI and ISV partner business development initiatives for Nebius’s AI cloud platform. The role involves collaborating with partners, managing strategic relationships, and supporting growth across AI infrastructure and cloud services. The posting provides limited specific responsibilities or qualifications, but emphasizes English proficiency and authorization to work in the country of application.
Top Skills: Ai/MlCloud ComputingCloud InfrastructureGpu Orchestration
15 Days AgoSaved
In-Office
Zürich, CHE
Senior level
Senior level
Artificial Intelligence • Information Technology • Consulting
Own optimization of LLM and VLM inference systems from model artifacts through production deployment. Improve latency, throughput, memory efficiency, GPU utilization, quality, reliability, and cost per token. Deploy and extend inference engines, build model-compression workflows, implement decoding and KV-cache optimizations, create benchmark harnesses, and diagnose bottlenecks across models, kernels, runtimes, schedulers, gateways, and clusters. Collaborate with research, infrastructure, product, kernel, and customer teams.
Top Skills: AwqCudaEagleFlashinferFp8GptqInt4Int8KserveLmcacheMedusaMxfp4Nvfp4Nvidia DynamoPythonPyTorchRayRay ServeSglangSmoothquantTensorrt-LlmTritonTriton Inference ServerVllm
15 Days AgoSaved
Remote
26 Locations
Senior level
Senior level
Artificial Intelligence • Information Technology • Consulting
Own production optimization of LLM and VLM inference systems, improving latency, throughput, memory efficiency, GPU utilization, quality, reliability, and cost. Deploy and benchmark inference engines, build compression and quantization workflows, implement acceleration techniques, diagnose serving bottlenecks, and collaborate with kernel, platform, research, product, and customer teams.
Top Skills: CudaFlashinferKserveKubernetesLmcacheNvidia DynamoPythonPyTorchRayRay ServeSglangTensorrt-LlmTritonTriton Inference ServerVllm
New

Track Smarter, Apply Better.

Ditch the spreadsheets. Organize your job search with our freeApplication Tracker.

Use For Free
Application Tracker Preview
16 Days AgoSaved
In-Office
London, Greater London, England, GBR
Senior level
Senior level
Artificial Intelligence • Information Technology • Consulting
Own optimization of LLM and VLM inference systems from model artifacts through production deployment. Improve latency, throughput, memory efficiency, GPU utilization, quality, reliability, and cost per token. Deploy and extend inference engines, build compression and quantization workflows, implement decoding and KV-cache optimizations, create reproducible benchmark harnesses, diagnose production regressions, and collaborate with kernel, platform, research, infrastructure, product, and customer teams.
Top Skills: AwqCudaEagleFlashinferFp8GptqInt4Int8KserveLmcacheMedusaMxfp4Nvfp4Nvidia DynamoPythonPyTorchRayRay ServeSglangSmoothquantTensorrt-LlmTritonTriton Inference ServerVllm
16 Days AgoSaved
In-Office
Tel Aviv, ISR
Senior level
Senior level
Artificial Intelligence • Information Technology • Consulting
Design, build, and own reliable production data pipelines for analytics and machine learning using Python and SQL. Implement idempotent data processing, transformations, validation, quality checks, and performance optimization. Orchestrate workloads with Airflow or equivalent frameworks, and run them using Docker and Kubernetes. Collaborate with Analytics, Data Science, and ML teams while contributing technical leadership, CI/CD automation, cloud operations, and cost-efficient infrastructure management.
Top Skills: Apache AirflowSparkCi/CdCloud ComputingDockerInfrastructure As CodeKubernetesLinuxNon-Relational DatabasesPreemptible ComputingPythonRelational DatabasesSpot ComputingSQL
18 Days AgoSaved
In-Office
Hyderabad, Telangana, IND
Entry level
Entry level
Artificial Intelligence • Information Technology • Consulting
Investigate complex datacenter incidents involving servers, GPUs, firmware, BIOS/BMC, Linux, and hardware-software interactions. Lead root cause analysis, identify recurring patterns, drive escalations with R&D and ODM vendors, validate firmware rollouts, and create durable fixes. Develop runbooks and troubleshooting guides to improve L1/L2 capabilities. Provide hands-on support during critical incidents and travel to datacenters as needed.
Top Skills: BashBiosBmcDcgmiGpu ServersIpmitoolLinuxNvidia-SmiOcpPythonRedfish
18 Days AgoSaved
In-Office
Béthune, Pas-de-Calais, Hauts-de-France, FRA
Entry level
Entry level
Artificial Intelligence • Information Technology • Consulting
Investigate complex datacenter incidents involving servers, GPUs, firmware, hardware, and Linux. Lead root-cause analysis, identify recurring issues, coordinate evidence-based escalations with R&D and vendors, validate BIOS/BMC firmware rollouts, and create runbooks that improve L1/L2 support. The role includes hands-on datacenter troubleshooting, platform readiness, incident leadership, and technical enablement across Europe and the US.
Top Skills: BashBiosBmcDcgmiGpu ServersIpmitoolLinuxNvidia-SmiOcpOdm PlatformsPythonRedfish
18 Days AgoSaved
In-Office
Amsterdam, NLD
Entry level
Entry level
Artificial Intelligence • Information Technology • Consulting
Own the L3 knowledge system for datacenter servers, GPU platforms, firmware, out-of-band management, and Linux diagnostics. Develop and maintain SOPs, runbooks, troubleshooting guides, and error catalogs; capture findings from L3 investigations; validate procedures with support teams; and prepare documentation for new hardware platforms. The role requires hands-on technical expertise, strong Confluence skills, precise technical writing, and occasional datacenter travel.
Top Skills: Atlassian ConfluenceBashBiosBmcGitIpmiIpmitoolLinuxNvidia DcgmiNvidia Nvidia-SmiOpen Compute Project (Ocp)PythonRedfish
All Filters
JobType
New Jobs
Job Category
Experience
Industry
Company Name
Company Size

Sign up now Access later

Create Free Account