Top Tech Jobs & Startup Jobs

2 Days AgoSaved
Hybrid
London, Greater London, England, GBR
75K-95K Annually
Entry level
75K-95K Annually
Entry level
Artificial Intelligence • Machine Learning • Software
Support ML engineering teams running production training and inference workloads across Kubernetes, cloud, Linux, and GPU infrastructure. Diagnose distributed PyTorch, CUDA, NCCL, networking, storage, scheduling, and performance issues; analyze observability data; advise customers during incidents; and improve reliability through tooling, automation, documentation, runbooks, and operational processes. The role is hybrid in London, requires at least two office days weekly, and supports EMEA shifts.
Top Skills: Bare Metal InfrastructureCloud InfrastructureContainerizationCudaDistributed SystemsGpu InfrastructureGrafanaInfinibandKubeflowKubernetesLinuxNcclOpentelemetryPrometheusPythonPyTorchRayRdmaSlurm
2 Days AgoSaved
Hybrid
4 Locations
180K-250K Annually
Mid level
180K-250K Annually
Mid level
Artificial Intelligence • Machine Learning • Software
Build and deploy production AI systems for customers, translating business objectives into scalable technical solutions. Responsibilities include architecture, proof-of-concepts, software development, deployment, monitoring, debugging, inference optimization, and distributed systems operation. The engineer partners with customer engineering teams, collaborates with product and engineering, improves reusable platform capabilities, and owns technical engagements from discovery through production scaling.
Top Skills: APIsDistributed SystemsDockerGoGpu-Accelerated WorkloadsKubernetesLanggraphModel Serving SystemsPythonRayReactTensorrtTypescriptVector DatabasesVllmWorkflow Orchestration Systems
2 Days AgoSaved
Remote or Hybrid
4 Locations
180K-220K Annually
Senior level
180K-220K Annually
Senior level
Artificial Intelligence • Machine Learning • Software
Build and operate production software, APIs, tooling, and automation for large-scale GPU, bare-metal, and HPC infrastructure. Responsibilities include provisioning, configuration, monitoring, lifecycle management, observability, hardware integration, reliability improvements, and infrastructure capacity deployment. The role partners with networking, data center, platform, and infrastructure teams to design scalable systems and define technical direction.
Top Skills: APIsBare-Metal InfrastructureBmcContainerizationDell HardwareGpu ServersHpcIpmiJuniper NetworksLinuxOrchestrationPalo Alto FirewallsPxe/IpxePythonRedfishSonicVast
2 Days AgoSaved
Hybrid
New York, NY, USA
155K-220K Annually
Expert/Leader
155K-220K Annually
Expert/Leader
Artificial Intelligence • Machine Learning • Software
Leads Lightning AI’s global customer experience organization, including Technical Account Managers and Support Engineers. Owns strategy, onboarding, adoption, technical success, retention, expansion, account health, escalations, and customer experience metrics. Partners closely with Product, Engineering, Infrastructure, Sales, and Finance to resolve customer issues, influence product priorities, and scale high-touch support through automation and self-service. The role requires strong technical fluency, executive communication, operational rigor, and experience supporting complex SaaS, cloud, developer platform, or AI infrastructure products.
Top Skills: Distributed TrainingGpu ClustersInferenceKubernetesModel ServingPyTorchSlurm
2 Days AgoSaved
Hybrid
2 Locations
180K-220K Annually
Entry level
180K-220K Annually
Entry level
Artificial Intelligence • Machine Learning • Software
Designs, implements, and maintains secure network infrastructure for Lightning AI’s production and AI environments. Responsibilities include firewall, VPN, IDS/IPS, segmentation, vulnerability assessment, penetration testing, SIEM monitoring, incident response, compliance documentation, access controls, and security automation. The role partners with infrastructure and engineering teams to mitigate threats and optimize security controls.
Top Skills: Cisco AsaCloud SecurityEncryptionEndpoint SecurityFirewallsFortinetIds/IpsNessusNetwork SegmentationPalo AltoQradarQualysSd-WanSecurity AutomationSIEMSplunkVpn
New

Cut your apply time in half.

Use ourAI Assistantto automatically fill your job applications.

Use For Free
Application Tracker Preview
2 Days AgoSaved
Hybrid
New York, NY, USA
160K-245K Annually
Senior level
160K-245K Annually
Senior level
Artificial Intelligence • Machine Learning • Software
Drive data-informed decisions across Product, Sales, and Engineering by analyzing usage patterns, business performance, and trends. Build data pipelines, write SQL and Python analyses, create Looker and Data Studio dashboards, monitor opportunities and risks, and enable teams to use analytics tools effectively.
Top Skills: AWSBigQueryData StudioEfsGCPLookerLooker StudioPythonS3SQL
2 Days AgoSaved
Remote or Hybrid
4 Locations
Entry level
Entry level
Artificial Intelligence • Machine Learning • Software
This is a general talent community opportunity rather than a specific open position. Lightning AI invites candidates to submit their information for consideration as future roles become available across U.S. and London hubs, with occasional remote opportunities. The company develops tools and infrastructure for building, training, and deploying AI systems and values ownership, urgency, communication, teamwork, continuous improvement, and long-term thinking.
2 Days AgoSaved
Hybrid
New York, NY, USA
175-200 Annually
Senior level
175-200 Annually
Senior level
Artificial Intelligence • Machine Learning • Software
Own strategic channel and technology partner relationships, develop joint business plans, create co-selling motions, identify partner-sourced opportunities, and drive pipeline, revenue, customer adoption, and field engagement. The role coordinates with sales, solutions engineering, marketing, customer success, and partner teams; manages enablement, executive alignment, quarterly reviews, partner workflows, and performance metrics while helping build an early-stage partner ecosystem.
2 Days AgoSaved
Remote or Hybrid
3 Locations
165K-310K Annually
Senior level
165K-310K Annually
Senior level
Artificial Intelligence • Machine Learning • Software
Design, optimize, and deploy large language model training and post-training pipelines. Improve model quality through fine-tuning, reinforcement learning, preference optimization, evaluation, and experimentation. Build PyTorch-based infrastructure, optimize distributed multi-GPU training, diagnose performance and convergence issues, and develop production-ready AI systems. Collaborate with researchers, infrastructure engineers, platform teams, and customers while contributing to open-source projects and reusable training capabilities.
Top Skills: CudaDeepspeedDistributed TrainingDpoFsdpGpu Performance OptimizationGrpoHugging Face TransformersLightning FabricMegatron-LmMixed PrecisionMulti-Gpu SystemsNvidia MoltPeftPpoPythonPyTorchReinforcement LearningReward ModelingRlhfSglangSupervised Fine-TuningTensorrtTransformer-Based Language ModelsTritonTrlVllm
2 Days AgoSaved
Remote or Hybrid
4 Locations
180K-220K Annually
Senior level
180K-220K Annually
Senior level
Artificial Intelligence • Machine Learning • Software
Operate, scale, and optimize distributed storage infrastructure supporting large-scale AI/ML and HPC workloads. Build Python automation, manage Linux bare-metal systems, troubleshoot storage, hardware, networking, and operating system issues, and improve performance, reliability, monitoring, capacity planning, and lifecycle management. Collaborate with infrastructure, networking, platform, and data center teams on storage deployments and scaling strategies.
Top Skills: CephGpu Direct StorageLinuxNfsPythonRdmaS3Vast
All Filters
JobType
New Jobs
Job Category
Experience
Industry
Company Name
Company Size

Sign up now Access later

Create Free Account