Customer Support Engineer

Posted 2 Hours Ago
Be an Early Applicant
San Francisco, CA, USA
In-Office
Mid level
Artificial Intelligence • Cloud • Information Technology • Infrastructure as a Service (IaaS)
The Role
Own enterprise customer support for GPU and AI infrastructure. Triage compute, networking, scheduling, orchestration, performance, hardware, and facility issues using monitoring tools. Respond to incidents, coordinate escalations with engineering and vendors, communicate clearly under pressure, and improve runbooks, SOPs, knowledge bases, and support processes. Participate in 24/7 shift coverage, including occasional after-hours, weekend, and holiday incident support.
Summary Generated by Built In
About San Francisco Compute

San Francisco Compute exists to give ambitious teams large scale supercomputers without taking on a catastrophic balance sheet risk. Most GPU clouds force you to sign long term contracts with no way out. If you underbuy, you miss the window. If you overbuy, you’ll blow up with no way out. It’s a desperate gamble that San Francisco Compute does not force you to play.

Originally, SFC was an AI lab that bought too big of a GPU cluster and was forced to sublease it or go under. Today, we build data centers, supercomputers, and then lease those machines on contracts that let our customers sublease. More than anyone else in the industry, we deeply felt the pain and fear of being forced to bet the farm with your neocloud vendor. To make it possible to sublease, we invented the first compute market and became the first one to offer physical settlement. That means we operate GPU clusters like a neocloud and derisk our customers like a market.

Customers buying GPU clusters don’t want to work through a network of brokers. They want to transact with the folks that have the keys to the cluster and can own the SLA.

To keep the industry safe from unnecessary risk, SFC needs to scale fast. To do that, we serve financially motivated parties as a technical partner to build, operate, & lease supercomputers on their behalf. Basically, we help folks who want an economic return build GPU clusters on their balance sheet, but operated in full by us. That allows us to retain the technical control needed to operate a cloud, but scale fast like a market. SFC’s ability to offer subleasing can significantly increase the levered returns for cluster owners. Our hybrid model out-performs brokered contracts or other “compute markets”, who have to go through a network of counterparties to solve problems. It also lets us design custom solutions for our customers, like clusters deployed in regions next to your dataset or physical eval set or a large CPU fleet deployed colocated with your GPU cluster.

Our team includes senior & technical leadership from places like Lambda, Crusoe, Digital Ocean, AWS, and Hut8. Our CTO is the cofounder and former CEO of Voltage Park. In prior roles, our team has deployed 8GW of datacenter capacity & hundreds of thousands of GPUs. We’ve been described as having “the highest talent density in the space.” If you are an honorable, gritty person with eyes wide open & good epistemics, who wants to see AI go right, we’d love to work with you. There has never been a more critical moment in history and it is up to us to shape it on behalf of those who come after us.

What You'll Do

Customer-Facing Support

  • Own inbound support tickets and live escalations for enterprise customers running workloads on our GPU/AI infrastructure

  • Triage and resolve technical issues across compute, networking, and platform layers, spanning both scheduling/orchestration problems and performance issues

  • Communicate clearly with technical customers under pressure - set expectations, give real status updates, close the loop

  • Hand off cleanly across timezones so customers never feel the seams of follow-the-sun coverage

Technical Troubleshooting & Escalation

  • Use monitoring and alerting tools to diagnose issues before or as customers report them

  • Escalate hardware, data-center, or facility-level issues to the right internal or external party (engineering, colo partners, hardware OEMs) with a clear, well-documented handoff

  • Serve as a first responder on incidents, working alongside engineering through to resolution

Process & Tooling

  • Work from and help improve runbooks, SOPs, and the knowledge base - flag gaps, don't just work around them

  • Use AI-assisted tooling to work faster without losing quality or judgment

  • Track and care about your own CSAT, first-response, and time-to-resolve numbers - these aren't just manager metrics, they're your feedback loop

About You
  • 3-5 years in technical support, customer support engineering, or a similar customer-facing technical role

  • Comfortable troubleshooting infrastructure-level issues - Linux administration, basic shell or Python scripting, and hands-on use of monitoring/observability tools (e.g. Prometheus, Grafana, Datadog); GPU/AI/HPC experience is a strong plus, not a requirement

  • Can explain technical problems clearly to both technical customers and internal engineering teams

  • Calm under pressure - you don't rattle when a customer is frustrated or a system is down

  • Comfortable working shift-based hours as part of a 24/7 global coverage model, including occasional after-hours, weekend, or holiday coverage during incidents

  • Genuinely care about getting the customer to a good outcome, not just closing the ticket

  • Excited to help build process and coverage from scratch, not just operate inside an existing one

  • Excellent written communication skills, to both customers and internal teams

Nice to Haves
  • Experience with GPU/AI infrastructure, specialized hardware, or managed services

  • Familiarity with data-center or colocation operations

  • Experience with modern support tooling (ticketing, monitoring/alerting platforms) and AI-assisted workflows

  • Scripting or basic programming ability for diagnostics and automation

  • Experience with incident management

BenefitsGenerous equity grant

Team members are offered a competitive salary along with equity in the company

Visa Sponsorships

Yes, we sponsor visas and work permits

Retirement matching

We match 401(k) plans up to 4%

Medical, dental & vision

We offer competitive medical, dental, vision insurance for employees and dependents and cover 100% of premiums

Time off

We offer unlimited paid time off as well as 10+ observed holidays

Parental leave

We offer biological, adoptive, and foster parents paid time off to spend quality time with family

Daily lunch

We cover lunch daily for employees

Unlimited office book budget

You can buy as many books for the office as you want

The San Francisco Compute Company is committed to maintaining a workplace free from discrimination and harassment.

We make employment decisions based on business needs, job requirements, and individual qualifications, without regard to race, color, religion, belief, national origin, social or ethical origin, age, physical, mental, or sensory disability, sexual orientation, gender identity or expression, marital status, civil union or domestic partnership status, past or present military service, HIV status, family medical history or genetic information, family or parental status including pregnancy, or any other status protected by law.

We welcome the opportunity to consider qualified applicants with prior arrest or conviction records. Our commitment to diversity includes hiring talented individuals regardless of their criminal history, in accordance with local, state, and federal laws, including San Francisco’s Fair Chance Ordinance and California’s ban-the-box laws.

Skills Required

  • 3-5 years of experience in technical support, customer support engineering, or a similar customer-facing technical role
  • Experience troubleshooting infrastructure-level issues
  • Linux administration experience
  • Basic shell or Python scripting ability
  • Hands-on experience with monitoring and observability tools
  • Ability to explain technical problems clearly to customers and engineering teams
  • Ability to work calmly under pressure during customer escalations or system outages
  • Availability for shift-based hours in a 24/7 global coverage model, including occasional after-hours, weekend, or holiday coverage
  • Strong customer-outcome orientation
  • Ability to build and improve support processes from scratch
  • Excellent written communication skills
  • Experience with GPU or AI infrastructure, specialized hardware, or managed services
  • Familiarity with data-center or colocation operations
  • Experience with modern support tooling and AI-assisted workflows
  • Experience with incident management
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
30 Employees
Year Founded: 2023

What We Do

San Francisco Compute Company operates a marketplace for large-scale GPU clusters, enabling users to buy and sell compute contracts with flexible terms. They aim to make AI compute more accessible and affordable by creating a liquid market for GPU offtake.

Similar Jobs

Vapi Logo Vapi

Support Engineer

Artificial Intelligence • Software
Hybrid
San Francisco, CA, USA
48 Employees
154K-169K Annually
In-Office
San Diego, CA, USA
249 Employees
115K-200K Annually

Nexthop AI Logo Nexthop AI

Support Engineer

Artificial Intelligence • Cloud • Information Technology • Software
In-Office
Santa Clara, CA, USA
142 Employees

Rillet Logo Rillet

Support Engineer

Artificial Intelligence • Software
Hybrid
2 Locations
200 Employees
85K-120K Annually

Similar Companies Hiring

Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software
New York, New York
30 Employees
Vega Thumbnail
Artificial Intelligence • Automotive • Insurance • Transportation
US
43 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account