Maximum of 25 job preferences reached.
Top AI & Machine Learning Jobs
Artificial Intelligence • Cloud • Software • Infrastructure as a Service (IaaS)
Leads DigitalOcean’s Applied Research function for agentic AI, setting the research agenda across model routing, memory, observability, and reinforcement learning. Builds and manages a research team, connects research investments to product outcomes, and partners with engineering to ship production capabilities that improve reliability, cost, latency, and task success. Represents the company externally through publications, talks, open-source work, and research collaborations.
Top Skills:
Adaptive RoutingAgent ObservabilityAgentic AiArtificial IntelligenceDpoGrpoInference InfrastructureLarge Language Models (Llms)Machine LearningMemory SystemsModel SelectionModel ServingOpen-Source ModelsOpen-Weight ModelsPpoReinforcement LearningRetrieval SystemsRlaifRlhf
Artificial Intelligence • Cloud • Software • Infrastructure as a Service (IaaS)
Design and optimize scalable, multi-tenant serverless inference infrastructure and APIs for large-scale AI workloads. Build highly available distributed services, improve GPU utilization, observability, capacity management, fault tolerance, and operational automation. Partner with platform and GPU infrastructure teams, contribute to architecture decisions, participate in on-call and incident response, and provide technical leadership and mentorship to engineers.
Top Skills:
Api GatewaysCloud-Native ArchitecturesDistributed SystemsGoGpu InfrastructureKubernetesLlm ServingMicroservicesObservabilityService MeshSreTensorrt-LlmTritonVllm
Artificial Intelligence • Cloud • Software • Infrastructure as a Service (IaaS)
Serve as a technical partner to sales and customer success teams, advising high-value customers on cloud infrastructure and AI/ML solutions. Responsibilities include architecture reviews, proofs of concept, demos, onboarding, technical workshops, infrastructure optimization, LLM deployment and fine-tuning, GPU workload optimization, AI application design, data pipeline and database integration, escalations, documentation, and development of tools for premium customers. Collaborate globally with Engineering, Product, Support, and account teams.
Top Skills:
Ai/MlAnsibleAWSAzureBashChefCi/CdCiscoCloud-Native DatabasesCoreosCudaData PipelinesDigitaloceanDockerFp8GCPGenerative AiGitGoHugging FaceInt4Int8JuniperKubernetesKvmLinuxLlmsManaged DbaasMongoDBMySQLNfsNvidia GpuObject StorageOracle CloudPostgresPuppetPythonPyTorchRedisRubySaltstackTensorFlowTensorrtTerraformVagrantVllmXen
Artificial Intelligence • Cloud • Software • Infrastructure as a Service (IaaS)
Own DigitalOcean’s enterprise AI-native architecture, including agent orchestration, shared AI platform services, tool and capability layers, enterprise-system integrations, identity, authorization, governance, observability, evaluation, and cost controls. Build production prototypes and reference implementations, write code, establish technical standards, mentor distributed engineers, and re-engineer business workflows with functional teams. Partner across platform engineering, security, identity, data, and business technology to deliver governed, scalable agentic systems.
Top Skills:
APIsArizeAutogenBoomiBraintrustCrewaiEu Ai ActEvent-Driven ArchitectureGoGraphragGreenhouseIso/Iec 42001JavaKubernetesLangfuseLanggraphLangsmithModel Context Protocol (Mcp)MulesoftN8NNetSuiteNist Ai RmfOpenai Agents SdkOpentelemetryOwasp Agentic AiPythonSalesforceSemantic KernelServerlessServicenowSpiffe/SpireTypescriptVector StoresWorkatoWorkday
Artificial Intelligence • Cloud • Software • Infrastructure as a Service (IaaS)
Lead and scale a Customer Success Engineering team supporting strategic cloud and AI/ML customers. Own 24x7 support operations, escalations, KPIs, and technical enablement across Kubernetes, databases, compute, GPUs and MLOps. Act as technical escalation for Sev1/Sev2 incidents, partner with TAMs and Product/Engineering, improve escalation runbooks and knowledge base, and drive automation and tooling to improve CSAT, response/resolution times, and support quality.
Top Skills:
AWSAzureBare Metal Gpu ProvisioningDatabasesGenerative AiGCPGpu Infrastructure (Nvidia H100H200)Hugging FaceJIRAKubernetes (Doks)LangchainLarge Language Models (Llms)MlopsNlpPythonPyTorchRestful ApisScikit-LearnTensorFlow
New
Track Smarter, Apply Better.
Ditch the spreadsheets. Organize your job search with our freeApplication Tracker.
Use For Free
Artificial Intelligence • Cloud • Software • Infrastructure as a Service (IaaS)
The Staff Forward Deployed Engineer drives AI adoption for strategic customers by solving complex cloud infrastructure challenges, developing scalable assets, and influencing product roadmaps through collaboration and technical expertise.
Top Skills:
CrewaiCudaGoGpuKubernetesLanggraphLlamaindexOpenai TritonPulumiPythonRocmTensorrtTerraform
Artificial Intelligence • Cloud • Software • Infrastructure as a Service (IaaS)
The Staff Forward Deployed Engineer will work with AI-Native customers to solve cloud infrastructure challenges, build scalable tools, optimize AI workloads, and influence product development by embedding with clients and establishing best practices.
Top Skills:
CudaGoKubernetesOpenai TritonPulumiPythonRocmTensorrtTerraform
Artificial Intelligence • Cloud • Software • Infrastructure as a Service (IaaS)
Lead architecture and low-level optimizations for LLM inference to maximize throughput and minimize latency. Profile and optimize GPU kernels, implement quantization and parallelization strategies (FP8/INT8/FP4), tune libraries (AITER, Triton, CUDA, ROCm), and mentor engineers while translating hardware limits into shippable platform features.
Top Skills:
Ai/Ml InferenceAmd AiterAmd Mi355XBf16CudaDeepseekFlashattentionFp4Fp8Glm-5Gpu Kernel DevelopmentInt8MoeMulti-Node Gpu ClusteringOpenai TritonQwen3-235BRmsnormRocmTensorrtTflopsTriton Compiler
Artificial Intelligence • Cloud • Software • Infrastructure as a Service (IaaS)
Lead design and implementation of scalable, multi-tenant serverless inference services focused on throughput, GPU utilization, resiliency, observability, and operational tooling. Partner with platform and GPU teams, debug production performance issues, mentor engineers, drive incident response, and create reusable patterns, standards, and automation for inference workloads.
Top Skills:
Api GatewayCloud-NativeGoGpuKubernetesMicroservicesService MeshTensorrt-LlmTritonVllm
Artificial Intelligence • Cloud • Software • Infrastructure as a Service (IaaS)
Lead benchmarking and performance optimization for AI inference at GPU/kernel levels. Solve memory bandwidth and compute bottlenecks, implement quantization and parallelization (multi-node), advise on hardware/software stacks, mentor engineers, and collaborate to ship high-performance inference features.
Top Skills:
AiterAmd GpusAskBf16CkCudaCustom Cuda KernelsFlashattentionFp4Fp8Int8MoeNvidia GpusOpenai TritonRmsnormRocmTensorrtTritonTriton Compiler
Let Your Resume Do The Work
Upload your resume to be matched with jobs you're a great fit for.
Success! We'll use this to further personalize your experience.
Top Companies Hiring AI & Machine Learning Roles
See AllAll Filters
Total selected ()
No Results
No Results

