Top Tech Jobs & Startup Jobs

3 Days AgoSaved
Hybrid
San Francisco, CA, USA
Expert/Leader
Expert/Leader
Artificial Intelligence • Cloud • Generative AI • Infrastructure as a Service (IaaS)
Own the vision, strategy, and roadmap for FriendliAI’s AI inference platform, including model serving, deployment, orchestration, APIs, and developer capabilities. Partner with engineering, research, enterprise customers, and go-to-market teams to improve inference performance, scalability, reliability, usability, and cost. Lead product initiatives from discovery through launch, establish product planning and measurement practices, and mentor Product Managers and Designers.
Top Skills: Ai Inference PlatformsAPIsCloud-Native InfrastructureDistributed SystemsGpu InfrastructureLlm ServingModel ServingMulti-Tenant SaasSdksSglangTensorrt-LlmVllm
21 Days AgoSaved
Hybrid
San Francisco, CA, USA
Mid level
Mid level
Artificial Intelligence • Cloud • Generative AI • Infrastructure as a Service (IaaS)
Customer-facing engineer who deploys, scales, and operates LLM and multimodal inference workloads. Build and manage containerized deployments on Kubernetes, create Terraform/Helm automation, diagnose production performance issues, integrate with customers' CI/CD pipelines, and provide reliability and observability guidance through workshops and hands-on support.
Top Skills: AWSCi/CdDeepspeed-InferenceDockerEksElkGCPGpuGrafanaHelmKubernetesLokiOciOpentelemetryPrometheusTensorrtTerraformTritonVllm
26 Days AgoSaved
Hybrid
San Francisco, CA, USA
Senior level
Senior level
Artificial Intelligence • Cloud • Generative AI • Infrastructure as a Service (IaaS)
Design, build, and operate a multi-cluster, multi-tenant Kubernetes fleet for GPU inference. Extend Kubernetes with controllers/CRDs, implement GPU scheduling and autoscaling, own the network data plane, design cross-cluster connectivity and service mesh, drive reliability/SLOs, deliver IaC with Terraform/Helm/GitOps, and collaborate with platform, SRE, and security teams.
Top Skills: AnsibleApi ServerAWSCiliumCniCrdsCustom ControllersDevice PluginsDnsDynamic Resource Allocation (Dra)EbpfEfaEtcdGitopsGoHelmInfinibandIngressIpamKube-ProxyKubeletKubernetesKubesprayLoad Balancing (L4/L7)MtlsNcclNvidia Gpu OperatorOperatorsPythonRdmaSchedulerService MeshSr-IovTerraform
Reposted One Month AgoSaved
Hybrid
San Francisco, CA, USA
Senior level
Senior level
Artificial Intelligence • Cloud • Generative AI • Infrastructure as a Service (IaaS)
Own and evolve core backend microservices for an AI inference platform: build production-grade APIs, multi-tenant SaaS features (auth, RBAC, billing), design OLTP/OLAP data models, collaborate on multi-cloud orchestration, ensure reliability/performance, and drive engineering quality through testing and CI/CD.
Top Skills: Ci/CdClickhouseFastapiGraphQLGrpcHugging FaceKubernetesLlm ServingNext.JsOpentelemetryPostgresPythonReactRestSQL
Reposted One Month AgoSaved
Hybrid
San Francisco, CA, USA
Mid level
Mid level
Artificial Intelligence • Cloud • Generative AI • Infrastructure as a Service (IaaS)
Design, build, and maintain agent APIs and production agent applications for document understanding, advanced RAG, and customer support automation. Integrate open-source models, collaborate with backend and infra for deployment and monitoring, and ensure APIs are robust, scalable, and developer-friendly.
Top Skills: Agent FrameworksAPIsCloud-Native DevelopmentContainer OrchestrationHugging FaceKubernetesLangchainLlamaindexMultimodal ModelsOcrPythonRagSummarization
New

Track Smarter, Apply Better.

Ditch the spreadsheets. Organize your job search with our freeApplication Tracker.

Use For Free
Application Tracker Preview
Reposted One Month AgoSaved
Hybrid
San Francisco, CA, USA
Mid level
Mid level
Artificial Intelligence • Cloud • Generative AI • Infrastructure as a Service (IaaS)
Design and optimize high-performance GPU kernels (GEMM, attention, routing) for AI inference across NVIDIA and AMD GPUs. Implement CUDA/C++ and low-level assembly code, build reduced-precision/quantized (FP8/FP4) kernels, benchmark cross-vendor performance, contribute to internal GPU libraries, accelerate multi-modal pipelines, and integrate next-generation GPU features into production.
Top Skills: AmdC++CudaCutlassFp4Fp8GemmGpu AssemblyHipNvidiaRocmTriton
Reposted One Month AgoSaved
Hybrid
San Francisco, CA, USA
Senior level
Senior level
Artificial Intelligence • Cloud • Generative AI • Infrastructure as a Service (IaaS)
Design, build, and maintain a scalable web platform and APIs for deploying and monitoring multimodal AI models and agent workflows. Collaborate with product, infrastructure, and design teams to optimize performance, ensure reliability, drive CI/CD and testing, and contribute to long-term architecture decisions for a cloud-native, multi-tenant SaaS system.
Top Skills: Ci/CdCloud-NativeFastapiGraphQLGrpcHugging FaceKubernetesLlmNext.JsOpentelemetryPostgresPythonRbacReactRestSQLTypescript
Reposted One Month AgoSaved
Hybrid
San Francisco, CA, USA
Senior level
Senior level
Artificial Intelligence • Cloud • Generative AI • Infrastructure as a Service (IaaS)
Lead end-to-end enterprise sales for FriendliAI's AI inference platform: generate pipeline, close high-value deals, run technical POCs, engage AI/ML communities, collaborate with engineering, and inform product roadmap.
Top Skills: Ai/MlCloud PlatformsDeveloper ToolsHugging FaceInference ServingLlmMlops
Reposted One Month AgoSaved
Hybrid
San Francisco, CA, USA
Senior level
Senior level
Artificial Intelligence • Cloud • Generative AI • Infrastructure as a Service (IaaS)
Design, implement, and optimize GPU kernels, kernel compiler, memory planner, and runtime for low-latency generative AI inference. Analyze performance bottlenecks across hardware and software, collaborate with infrastructure teams, and maintain production profiling, benchmarking, and validation tooling while supporting new model architectures and multi-GPU strategies.
Top Skills: BenchmarkingC++Compiler InfrastructureDiffusion ModelsDistributed InferenceGpu KernelsKernel CompilerMulti-GpuProfilingPythonRuntime SystemsTransformer Models
Reposted One Month AgoSaved
Hybrid
San Francisco, CA, USA
Mid level
Mid level
Artificial Intelligence • Cloud • Generative AI • Infrastructure as a Service (IaaS)
Design, deploy, and operate large-scale LLM and multimodal inference architectures. Work hands-on with customer engineering teams to containerize, scale, monitor, and troubleshoot GPU-based inference workloads across Kubernetes, CI/CD, and hybrid/on-prem environments. Create Helm charts, Terraform modules, and observability tooling while delivering workshops and platform reliability insights.
Top Skills: AWSCi/CdDeepspeed-InferenceDockerDocker ImagesEksElkGCPGpu ComputingGrafanaHelmHugging FaceKubernetesLokiOciOtelPrometheusTensorrtTerraformTritonVllm
All Filters
JobType
New Jobs
Job Category
Experience
Industry
Company Name
Company Size

Sign up now Access later

Create Free Account