Top Tech Jobs & Startup Jobs

Reposted 4 Days AgoSaved
In-Office
San Francisco, CA, USA
125K-153K Annually
Senior level
125K-153K Annually
Senior level
Artificial Intelligence • Cloud • Software • Infrastructure as a Service (IaaS)
The Manager, Customer Success Engineering at DigitalOcean leads a team supporting strategic customers in cloud and AI/ML workloads, focusing on operational excellence and customer satisfaction.
Top Skills: AIAWSAzureDatabasesGCPGpu InfrastructureKubernetesMlPythonPyTorchRestful ApisScikit-LearnTensorFlow
Reposted 4 Days AgoSaved
In-Office
Bengaluru, Bengaluru Urban, Karnataka, IND
Senior level
Senior level
Artificial Intelligence • Cloud • Software • Infrastructure as a Service (IaaS)
Lead the Use & Accounts engineering team, defining technical strategy, enhancing customer console experience, ensuring system stability, and managing team performance.
Top Skills: AWSAzureChefGCPGoHelmJavaScriptNode.jsTerraform
Reposted 4 Days AgoSaved
In-Office
Bengaluru, Bengaluru Urban, Karnataka, IND
Senior level
Senior level
Artificial Intelligence • Cloud • Software • Infrastructure as a Service (IaaS)
Lead and manage the Security Products engineering team, driving technical roadmap, ensuring quality, and fostering team growth. Oversee product delivery and security practices.
Top Skills: Apache FlinkAWSAzureChefGCPGoHelmJavaScriptKafkaNode.jsTerraform
Reposted 5 Days AgoSaved
In-Office
Boston, MA, USA
191K-239K Annually
Senior level
191K-239K Annually
Senior level
Artificial Intelligence • Cloud • Software • Infrastructure as a Service (IaaS)
Lead architecture and low-level optimizations for LLM inference to maximize throughput and minimize latency. Profile and optimize GPU kernels, implement quantization and parallelization strategies (FP8/INT8/FP4), tune libraries (AITER, Triton, CUDA, ROCm), and mentor engineers while translating hardware limits into shippable platform features.
Top Skills: Ai/Ml InferenceAmd AiterAmd Mi355XBf16CudaDeepseekFlashattentionFp4Fp8Glm-5Gpu Kernel DevelopmentInt8MoeMulti-Node Gpu ClusteringOpenai TritonQwen3-235BRmsnormRocmTensorrtTflopsTriton Compiler
Reposted 5 Days AgoSaved
In-Office
Seattle, WA, USA
191K-239K Annually
Senior level
191K-239K Annually
Senior level
Artificial Intelligence • Cloud • Software • Infrastructure as a Service (IaaS)
Lead design and implementation of scalable, multi-tenant serverless inference services focused on throughput, GPU utilization, resiliency, observability, and operational tooling. Partner with platform and GPU teams, debug production performance issues, mentor engineers, drive incident response, and create reusable patterns, standards, and automation for inference workloads.
Top Skills: Api GatewayCloud-NativeGoGpuKubernetesMicroservicesService MeshTensorrt-LlmTritonVllm
New

Track Smarter, Apply Better.

Ditch the spreadsheets. Organize your job search with our freeApplication Tracker.

Use For Free
Application Tracker Preview
Reposted 5 Days AgoSaved
In-Office
San Francisco, CA, USA
191K-239K Annually
Senior level
191K-239K Annually
Senior level
Artificial Intelligence • Cloud • Software • Infrastructure as a Service (IaaS)
Lead benchmarking and performance optimization for AI inference at GPU/kernel levels. Solve memory bandwidth and compute bottlenecks, implement quantization and parallelization (multi-node), advise on hardware/software stacks, mentor engineers, and collaborate to ship high-performance inference features.
Top Skills: AiterAmd GpusAskBf16CkCudaCustom Cuda KernelsFlashattentionFp4Fp8Int8MoeNvidia GpusOpenai TritonRmsnormRocmTensorrtTritonTriton Compiler
Reposted 5 Days AgoSaved
In-Office
Denver, CO, USA
191K-239K Annually
Senior level
191K-239K Annually
Senior level
Artificial Intelligence • Cloud • Software • Infrastructure as a Service (IaaS)
Lead architecture and low-level optimizations for high-performance AI inference. Benchmark and tune GPU kernels, memory and precision management, parallelize across multi-node GPU clusters, implement quantization (FP8/INT8/FP4), advise on hardware/software stacks (CUDA/ROCm/Triton/TensorRT), mentor engineers, and collaborate to turn hardware limits into shippable inference products.
Top Skills: Amd AiterAsk (Assembly)Bf16Ck (Composable Kernel)CudaFlashattentionFp4Fp8Int8Openai TritonRocmTensorrtTriton Compiler
Reposted 5 Days AgoSaved
In-Office
Austin, TX, USA
191K-239K Annually
Senior level
191K-239K Annually
Senior level
Artificial Intelligence • Cloud • Software • Infrastructure as a Service (IaaS)
Lead benchmarking and low-level performance optimization for GPU-based inference. Solve memory, precision, and parallelization bottlenecks, tune kernels (CUDA/Triton/ROCm/AMD AITER), implement quantization (FP8/INT8/FP4), advise on hardware/software integration, mentor engineers, and translate hardware limits into shippable features for high-throughput, low-latency inference fleets.
Top Skills: Amd Aiter (Ck/Ask)Amd GpusAmd Mi355XBf16CudaCuda KernelsFlashattentionFp4Fp8Int8Moe (Mixture Of Experts)Nvidia GpusOpenai TritonRms NormRocmTensorrtTriton Compiler
Reposted 6 Days AgoSaved
In-Office
Austin, TX, USA
139K-174K Annually
Senior level
139K-174K Annually
Senior level
Artificial Intelligence • Cloud • Software • Infrastructure as a Service (IaaS)
Lead design and implementation of high-scale, resilient AI inference data plane services. Architect distributed LLM serving systems, optimize performance (tensor/data parallelism, KV-cache routing), build on Kubernetes-native inference frameworks, contribute to open-source projects, mentor engineers, and operate production services with strong observability and SLOs.
Top Skills: GoGrpcKubernetesLlm-DModular MaxNixlNvidia DynamoPythonRay ServeSglangTensorrtTensorrt-LlmTgiVllm
Reposted 6 Days AgoSaved
In-Office
San Francisco, CA, USA
139K-174K Annually
Senior level
139K-174K Annually
Senior level
Artificial Intelligence • Cloud • Software • Infrastructure as a Service (IaaS)
Lead design and delivery of high-scale, resilient distributed inference data-plane services for LLM hosting. Optimize performance (tensor/data parallelism, KV-cache routing, prefill/decode disaggregation), collaborate cross-functionally, contribute upstream to open-source inference projects, mentor engineers, and operate production services with strong SLOs and observability.
Top Skills: GoGrpcKserveKubernetesLlm-DModular MaxNixlNvidia DynamoPythonRay ServeSglangTensorrtTensorrt-LlmTgiVllm
All Filters
JobType
New Jobs
Job Category
Experience
Industry
Company Name
Company Size

Sign up now Access later

Create Free Account