Technical Success Engineer

Posted Yesterday
Be an Early Applicant
2 Locations
Hybrid
251K-335K Annually
Mid level
Artificial Intelligence • Cloud • Machine Learning • Infrastructure as a Service (IaaS)
The Role
Owns technical customer deployments from contract through production readiness. Validates GPU, cloud, compute, storage, networking, and connectivity configurations; troubleshoots issues; coordinates Infrastructure, Engineering, Product, and Data Center teams; tracks risks and dependencies; provides status updates; guides onboarding; documents reusable runbooks; and serves as the customer’s technical contact until the environment is stable.
Summary Generated by Built In

Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers. Our customers range from AI researchers to enterprises and hyperscalers. Lambda's mission is to make compute as ubiquitous as electricity and give everyone the power of superintelligence. One person, one GPU.

If you'd like to build the world's best AI cloud, join us.

*Note: This position requires presence in our San Francisco or San Jose office location 4 days per week; Lambda’s designated work from home day is currently Tuesday.


About this role

The Superintelligence Technical Success Engineer is part of Lambda's Superintelligence business unit, dedicated to our largest, most strategic customers operating in the most complex environments. This role owns taking signed deployment from contract to a live, fully operational production environment. The role requires strong technical acumen to validate the build against what was promised, spot gaps, and work effectively with engineering and infrastructure teams to get issues resolved and requirements clearly understood.

You'll work alongside the account team, to make sure the customer's requirements and expectations are clearly understood and addressed throughout deployment. This role calls for someone who's ready to dive in wherever the engagement needs them.

What You'll Do

  • Take signed deployments, from contract, to live, to working production environments, validating configuration, connectivity, storage, and compute against what was promised

  • Bring technical depth to validation and troubleshooting conversations — asking the right questions, spotting gaps, and working closely with engineering and infrastructure teams to drive issues to resolution

  • Coordinate with Infrastructure, Engineering, Product, and Data Center teams to close technical dependencies and resolve blockers

  • Own a current, accurate technical picture of the deployment — what's built, what's open, what's at risk

  • Keep stakeholders informed with clear, regular status updates: RAG status, top risks, and what's being done about them

  • Guide the customer through onboarding to their first successful production workload

  • Be the customer's go-to technical contact through deployment and early production

  • Feed recurring technical patterns back into reusable runbooks, checklists, or automation

  • Transition out once the customer is stable and self-sufficient, keeping the broader account team informed along the way

You

  • 4+ years of hands-on technical experience with GPU/HPC infrastructure, cloud platforms, Kubernetes, or large-scale Linux systems

  • Comfortable being the technical voice in the room able to validate builds, spot gaps, and hold engineering teams accountable for resolving them

  • Strong troubleshooting instincts and a willingness to get into the weeds of networking, storage, or compute issues

  • Track record of coordinating across engineering and infrastructure teams to close out technical dependencies

  • Clear written and verbal communication for status updates and technical documentation

  • A bias toward diving in and taking ownership, rather than waiting for a fully defined process

Nice to Have

  • Exposure to large-scale GPU cluster deployments

  • Familiarity with project/program tracking tools and structured status reporting

  • Experience building runbooks or checklists that outlived the engagement they were built for

Salary Range Information

The annual salary range for this position has been set based on market data and other factors. However, a salary higher or lower than this range may be appropriate for a candidate whose qualifications differ meaningfully from those listed in the job description.

About Lambda

  • Founded in 2012, with 500+ employees, and growing fast

  • Our investors notably include TWG Global, US Innovative Technology Fund (USIT), Andra Capital, SGW, Andrej Karpathy, ARK Invest, Fincadia Advisors, G Squared, In-Q-Tel (IQT), KHK & Partners, NVIDIA, Pegatron, Supermicro, Wistron, Wiwynn, Gradient Ventures, Mercato Partners, SVB, 1517, and Crescent Cove

  • We have research papers accepted at top machine learning and graphics conferences, including NeurIPS, ICCV, SIGGRAPH, and TOG

  • Our values are publicly available: https://lambda.ai/careers

  • We offer generous cash & equity compensation

  • Health, dental, and vision coverage for you and your dependents

  • Wellness and commuter stipends for select roles

  • 401k Plan with 2% company match (USA employees)

  • Flexible paid time off plan that we all actually use

Equal Opportunity Employer

Lambda is an Equal Opportunity employer. Applicants are considered without regard to race, color, religion, creed, national origin, age, sex, gender, marital status, sexual orientation and identity, genetic information, veteran status, citizenship, or any other factors prohibited by local, state, or federal law.

Skills Required

  • 4+ years of hands-on technical experience with GPU/HPC infrastructure, cloud platforms, Kubernetes, or large-scale Linux systems
  • Ability to validate technical builds, identify gaps, and hold engineering teams accountable for resolution
  • Strong troubleshooting skills involving networking, storage, or compute issues
  • Experience coordinating across engineering and infrastructure teams to resolve technical dependencies
  • Clear written and verbal communication for status updates and technical documentation
  • Demonstrated ownership and willingness to work in undefined or evolving processes
  • Exposure to large-scale GPU cluster deployments
  • Familiarity with project or program tracking tools and structured status reporting
  • Experience building reusable runbooks or checklists
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: San Francisco, CA
750 Employees
Year Founded: 2012

What We Do

Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure and the #1 GPU Cloud for ML/AI teams. Their mission is to make compute as ubiquitous as electricity and give everyone the power of superintelligence.

Similar Jobs

Hybrid
2 Locations
106 Employees
251K-335K Annually

OpenAI Logo OpenAI

Demo Experience Engineer, Technical Success

Artificial Intelligence • Machine Learning • Generative AI
Hybrid
San Francisco, CA, USA
4500 Employees
234K-260K Annually

Macroscope Logo Macroscope

Technical Customer Success Engineer

Artificial Intelligence • Machine Learning • Software • Analytics
In-Office
San Francisco, CA, USA
26 Employees

Similar Companies Hiring

Legora Thumbnail
Artificial Intelligence • Legal Tech • Software
New York, New York
700 Employees
Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account