Lambda
Jobs at Lambda
Let Your Resume Do The Work
Upload your resume to be matched with jobs you're a great fit for.
Success! We'll use this to further personalize your experience.
Recently posted jobs
Artificial Intelligence • Cloud • Machine Learning • Infrastructure as a Service (IaaS)
Lead the technical vision and development of Lambda’s managed Kubernetes platform for AI workloads on bare metal. Design GPU-aware orchestration, multi-tenant control planes, networking, storage, inference services, Slurm integration, self-healing automation, and resilient managed services. Provide cross-functional infrastructure leadership, mentor engineers, drive technical standards, lead chaos engineering, and collaborate with customers, NVIDIA, and open-source communities.
Artificial Intelligence • Cloud • Machine Learning • Infrastructure as a Service (IaaS)
Build and operate distributed systems powering Lambda’s GPU cloud, including APIs, control planes, schedulers, workflows, and operational tooling. Own systems through architecture, implementation, deployment, observability, on-call, incident response, and continuous improvement. Improve reliability, performance, security, and scalability while collaborating across infrastructure, networking, storage, security, and SRE teams. Senior engineers lead complex domain work; Staff engineers shape cross-team architecture and mentor others.
Artificial Intelligence • Cloud • Machine Learning • Infrastructure as a Service (IaaS)
Design and build Lambda’s managed Kubernetes platform for large-scale AI workloads. Responsibilities include developing control-plane services, operators, controllers, cluster lifecycle automation, GPU-aware scheduling, inference infrastructure, autoscaling, internal deployment tools, observability, and production support. The role requires deep Kubernetes and distributed systems expertise, strong Go and Python skills, and collaboration across networking, storage, compute, and AI infrastructure teams.
Artificial Intelligence • Cloud • Machine Learning • Infrastructure as a Service (IaaS)
Designs and scales enterprise security across IAM, corporate SaaS, endpoints, and infrastructure. Responsibilities include automating joiner-mover-leaver workflows, access reviews, zero-trust controls, SaaS security posture management, endpoint protections, vendor risk assessments, and AI-powered security automation. The role partners with IT, Legal, HR, and Compliance on security frameworks, documentation, audits, and strategic roadmap initiatives.
Artificial Intelligence • Cloud • Machine Learning • Infrastructure as a Service (IaaS)
Leads the organization responsible for software services supporting secure, reliable, high-performance AI cloud networking. Owns distributed network systems, control planes, observability, security, traffic engineering, technical strategy, operational excellence, and organizational growth. Hires and develops engineering leaders, establishes SLOs and incident practices, partners with network, product, and operations teams, and represents technical direction to executives and customers.
Artificial Intelligence • Cloud • Machine Learning • Infrastructure as a Service (IaaS)
Leads Lambda’s global cloud and AI network engineering organization, owning architecture, security, performance, availability, capacity, operations, budgets, and data center readiness. Develops network roadmaps, internet peering and backbone strategy, automation, telemetry, and secure-by-design infrastructure. Builds and manages multidisciplinary engineering teams, develops leaders, partners with product and security executives, oversees incident response, and delivers reliable, cost-efficient connectivity for large-scale AI workloads.
Artificial Intelligence • Cloud • Machine Learning • Infrastructure as a Service (IaaS)
Developers Relations technical practitioner who helps enterprise teams deploy production AI workloads on Lambda. Creates technical guides, demonstrations, open-source examples, talks, workshops, and videos; grows reach through professional networks, communities, customers, and partners; represents Lambda at conferences and podcasts; gathers field feedback; and collaborates with engineering, MLE, marketing, and go-to-market teams.
Artificial Intelligence • Cloud • Machine Learning • Infrastructure as a Service (IaaS)
Own the multi-year strategy and technical direction for Lambda’s large-scale networking software across GPU fabrics, cloud networks, backbone, and edge infrastructure. Lead architecture, design, implementation, operations, customer engagements, and cross-functional delivery of distributed network services. Mentor engineers, resolve complex technical issues, participate in on-call operations, and guide mission-critical enterprise projects while improving system performance, availability, scalability, and cost.
Artificial Intelligence • Cloud • Machine Learning • Infrastructure as a Service (IaaS)
Manage debt and covenant compliance, debt service, hedge settlements and documentation, letters of credit, structured finance administration, intercompany funding, KYC/AML support, and debt forecasting. Maintain banking and counterparty relationships while preparing treasury analyses, forecasts, and leadership materials. The role requires strong financial modeling, Excel, communication, and coordination skills, with on-site presence in San Francisco four days per week.
Artificial Intelligence • Cloud • Machine Learning • Infrastructure as a Service (IaaS)
Manage core HR operations across the employee lifecycle, including onboarding, offboarding, employee support, payroll backup, compliance, leave administration, workplace accommodations, grievances, HRIS data integrity, and employee experience initiatives. Partner with HR, IT, managers, and employees while maintaining accurate records, handling confidential information, and improving HR processes and programs.
Artificial Intelligence • Cloud • Machine Learning • Infrastructure as a Service (IaaS)
Architect scalable compute platforms for AI/ML, simulation, and high-throughput workloads. Define system standards, reference designs, roadmaps, and technical requirements across hardware and software. Evaluate CPU, GPU, accelerator, networking, power, cooling, and density tradeoffs. Lead platform validation and performance characterization, collaborate with product and engineering teams, and mentor systems engineers on sizing and optimization.
Artificial Intelligence • Cloud • Machine Learning • Infrastructure as a Service (IaaS)
Own mechanical design standards and cooling architectures for large-scale AI data centers. Design and size central plants, hydronic distribution, air-cooling, and liquid-cooling systems; perform thermal, hydraulic, and equipment calculations; review designs and vendor documents; coordinate with engineering, construction, and operations teams; and support procurement, testing, commissioning, failure investigations, and site activities.
Artificial Intelligence • Cloud • Machine Learning • Infrastructure as a Service (IaaS)
Architect high-performance, low-latency networking for AI cloud platforms, GPU clusters, storage systems, and multi-tenant environments. Define network topologies, standards, reference designs, and scalability roadmaps; evaluate advanced InfiniBand and Ethernet technologies; guide automation and telemetry strategies; and build resilient, fault-tolerant architectures. Partner with compute and storage architects, mentor engineers, and influence technical decisions across teams.
Artificial Intelligence • Cloud • Machine Learning • Infrastructure as a Service (IaaS)
Own controls standards and reference architectures for mechanical systems in large-scale AI data centers. Design control narratives, sequences, automation, instrumentation, alarms, PLC/DDC/BMS requirements, and system responses to failures and changing loads. Review vendor designs, coordinate cross-functional controls requirements, support testing and commissioning, troubleshoot control issues, and incorporate operational lessons into future standards. Travel to project sites and partner facilities as needed.
Artificial Intelligence • Cloud • Machine Learning • Infrastructure as a Service (IaaS)
Lead architecture, design, scaling, automation, and operation of Lambda’s large-scale AI cloud network. Own major network domains, lead cross-team projects, resolve complex production issues, qualify hardware and software, improve observability and self-healing, mentor engineers, shape technical strategy, and participate in on-call operations. The role requires deep expertise in data center, backbone, internet, cloud, Linux, network automation, and multi-vendor networking technologies.
Artificial Intelligence • Cloud • Machine Learning • Infrastructure as a Service (IaaS)
Owns electrical design standards and reference architectures for low-voltage power systems supporting high-density AI data centers. Leads design across transformers, switchgear, UPS systems, distribution, busway, and rack-level delivery. Performs or oversees electrical studies, develops technical documentation, evaluates designs, coordinates cross-functional requirements, and supports equipment selection, testing, commissioning, construction, and failure investigations. The role requires occasional travel to project sites, manufacturers, and partner facilities.
Artificial Intelligence • Cloud • Machine Learning • Infrastructure as a Service (IaaS)
Owns electrical design standards and reference architectures for medium-voltage systems supporting large-scale AI data centers. Leads utility interconnection, campus distribution, studies, equipment specifications, protection schemes, procurement, testing, construction, commissioning, and failure investigations. Collaborates with utilities, engineering partners, manufacturers, construction teams, and operators. The role requires on-site presence in San Francisco four days weekly and travel to project sites and partner facilities as needed.
Artificial Intelligence • Cloud • Machine Learning • Infrastructure as a Service (IaaS)
Own product strategy, definition, roadmap, pricing, packaging, and hardware productization for Lambda’s compute, storage, and networking infrastructure. Lead the unified customer experience across AI cloud primitives, partner with NVIDIA on technology integration, establish cross-functional operating mechanisms, and hire and develop product managers. The role requires deep cloud infrastructure expertise and strong judgment across technical, commercial, customer, and hardware considerations.
Artificial Intelligence • Cloud • Machine Learning • Infrastructure as a Service (IaaS)
Lead Lambda’s Detection and Response function by building and managing a high-performing team, defining threat management and incident response frameworks, expanding automation and threat hunting, and ensuring 24/7 operational coverage. The role partners with engineering and executives, drives security tooling and AI-powered detection initiatives, establishes roadmaps and metrics, and leads blameless post-incident improvements across cloud and bare-metal infrastructure.
Artificial Intelligence • Cloud • Machine Learning • Infrastructure as a Service (IaaS)
Develop and execute executive communications strategies for Lambda’s leadership team. Responsibilities include shaping executive voices and public profiles, writing talking points, keynotes, long-form content and social posts, securing media and speaking opportunities, building relationships with AI infrastructure reporters and analysts, creating communications frameworks, and managing agency relationships.



