Maximum of 25 job preferences reached.
Top AI & Machine Learning Jobs
Artificial Intelligence • Cloud • Software • Infrastructure as a Service (IaaS)
Build and maintain scalable backend services and concurrent asynchronous workflows in Python/Go, and develop data-dense React/TypeScript UIs for configuring, running, and visualizing evaluations of AI agents. Integrate secure execution environments, design resilient APIs, manage large client-side state and streaming updates, and collaborate with cross-functional teams to ensure reliability, performance, and excellent developer experience.
Top Skills:
C++DockerFigmaGoGraphQLGrpcJavaJavaScriptKubernetesLlm ApisMicrovmsPythonReactRestTypescriptZod
Artificial Intelligence • Cloud • Software • Infrastructure as a Service (IaaS)
Lead and grow an engineering team to design, build, and scale a production LLM inference platform. Drive architecture, scheduling, GPU utilization, orchestration, observability, security, and cross-functional delivery for large Kubernetes clusters and multi-tenant AI workloads.
Top Skills:
Amd)CriuCudaGpu (NvidiaGvisorHamiKai-SchedulerKata ContainersKubernetesMicrovmsNumaNvidia Cuda-CheckpointNvidia GroveNvlinkOci Image VolumesPcieRocmSglangVllm
Artificial Intelligence • Cloud • Software • Infrastructure as a Service (IaaS)
Embed with customers to design, build, and deploy production agentic AI systems. Architect multi-agent workflows, prototype and harden LLM-based agent frameworks, implement stateful memory and evaluation/safety guardrails, optimize latency and cost, and translate field feedback into platform improvements while collaborating with product and engineering teams.
Top Skills:
Agentic Ai FrameworksAsyncioAutogenDistributed Training ToolsEvals FrameworksFastapiHugging Face TransformersLanggraphLlmMulti-Agent FrameworksPydanticPythonPyTorchState Space ModelsTensorFlowVlm
Artificial Intelligence • Cloud • Software • Infrastructure as a Service (IaaS)
Lead design, deployment, and optimization of production-scale LLM inference systems. Embed with customers to debug latency, profile GPU usage, implement cluster-scale strategies (prefill/decode disaggregation, KV-cache routing, tiered caching, MoE parallelism), drive hardware efficiency (quantization, parallelism, batching), build internal tooling, and contribute improvements upstream to open-source inference frameworks.
Top Skills:
Continuous BatchingData ParallelismDocker/ContainersFp4Fp8GoGpuGrpcKubernetesKv-CacheLlm-DModular MaxMoeNvidia DynamoPaged AttentionPythonRay ServeSglangTensor ParallelismTensorrt-LlmVllm
Artificial Intelligence • Cloud • Software • Infrastructure as a Service (IaaS)
Technical leader designing and implementing massive-scale, stateful multi-agent multi-turn simulation and evaluation systems. Build persona synthesis pipelines, what-if benchmarking frameworks, durable workflow orchestration, and high-performance APIs while integrating LLMs and agentic architectures. Drive architecture, mentor engineers, lead cross-functional strategy, and ensure scalability, reliability, and observability for AI feedback and evaluation infrastructure.
Top Skills:
AutogenCrewaiGoGrpcLangchainLlm OrchestrationLlmsMessage/Event-Driven ArchitecturesPythonStreaming Llm Token HandlingWorkflow Orchestration Engines
New
Cut your apply time in half.
Use ourAI Assistantto automatically fill your job applications.
Use For Free
Artificial Intelligence • Cloud • Software • Infrastructure as a Service (IaaS)
Lead and grow the Inference Orchestration engineering team to design, build, and operate Kubernetes-based AI infrastructure at scale. Drive scheduling, GPU utilization, topology-aware placement, checkpoint/restore for long jobs, fault tolerance, model distribution, security isolation, and cross-functional delivery to meet performance, cost, and reliability goals.
Top Skills:
Amd GpusCriuGvisorHamiKai-SchedulerKata ContainersKubernetesMicrovmsNumaNvidia Cuda-CheckpointNvidia GpusNvidia GroveNvlinkOci Image VolumesPcieSglangTritonVllm
Artificial Intelligence • Cloud • Software • Infrastructure as a Service (IaaS)
Lead and grow an AI engineering team to build internal copilots, agents, and an AI platform. Act as a player-coach contributing to architecture, prototypes, and code. Define technical roadmap, partner with business leaders to ship AI-native workflows, and ensure governance, observability, evaluation, and cost controls for LLM-driven systems.
Top Skills:
AgentsAutogenClaudeClaude CodeCrewaiCursorEvaluation HarnessesGithub CopilotGoGreenhouseGrpcJavaKubernetesLanggraphLlmopsLlmsMcp GatewayModel Context Protocol (Mcp)NetSuiteObservability StacksPythonRetrieval-Augmented Generation (Rag)SalesforceServerlessTypescriptVector StoresWorkday
Artificial Intelligence • Cloud • Software • Infrastructure as a Service (IaaS)
Lead and scale a Customer Success Engineering team supporting strategic cloud and AI/ML customers. Own 24x7 support operations, escalations, KPIs, and technical enablement across Kubernetes, databases, compute, GPUs and MLOps. Act as technical escalation for Sev1/Sev2 incidents, partner with TAMs and Product/Engineering, improve escalation runbooks and knowledge base, and drive automation and tooling to improve CSAT, response/resolution times, and support quality.
Top Skills:
AWSAzureBare Metal Gpu ProvisioningDatabasesGenerative AiGCPGpu Infrastructure (Nvidia H100H200)Hugging FaceJIRAKubernetes (Doks)LangchainLarge Language Models (Llms)MlopsNlpPythonPyTorchRestful ApisScikit-LearnTensorFlow
Artificial Intelligence • Cloud • Software • Infrastructure as a Service (IaaS)
The Staff Forward Deployed Engineer drives AI adoption for strategic customers by solving complex cloud infrastructure challenges, developing scalable assets, and influencing product roadmaps through collaboration and technical expertise.
Top Skills:
CrewaiCudaGoGpuKubernetesLanggraphLlamaindexOpenai TritonPulumiPythonRocmTensorrtTerraform
Artificial Intelligence • Cloud • Software • Infrastructure as a Service (IaaS)
The Staff Forward Deployed Engineer will work with AI-Native customers to solve cloud infrastructure challenges, build scalable tools, optimize AI workloads, and influence product development by embedding with clients and establishing best practices.
Top Skills:
CudaGoKubernetesOpenai TritonPulumiPythonRocmTensorrtTerraform
Let Your Resume Do The Work
Upload your resume to be matched with jobs you're a great fit for.
Success! We'll use this to further personalize your experience.
Top Companies Hiring AI & Machine Learning Roles
See AllAll Filters
Total selected ()
No Results
No Results

