Principal Cloud Platform Engineer

Reposted 18 Days Ago
2 Locations
In-Office
210K-280K Annually
Mid level
Artificial Intelligence • Hardware • Machine Learning • Natural Language Processing • Software • Generative AI
SambaNova is the #1 platform for business AI.
The Role
Own and operate a global AI inferencing platform: ensure uptime, low-latency performance, scalability and cost efficiency. Build monitoring, alerting, CI/CD, IaC, autoscaling, capacity planning, incident response, and participate in a shared on-call rotation to maintain 24/7 service reliability.
Summary Generated by Built In

The era of pervasive AI has arrived. In this era, organizations will use generative AI to unlock hidden value in their data, accelerate processes, reduce costs, drive efficiency and innovation to fundamentally transform their businesses and operations at scale.

SambaNova Suite™ is the first full-stack, generative AI platform, from chip to model, optimized for enterprise and government organizations. Powered by the intelligent SN40L chip, the SambaNova Suite is a fully integrated platform, delivered on-premises or in the cloud, combined with state-of-the-art open-source models that can be easily and securely fine-tuned using customer data for greater accuracy. Once adapted with customer data, customers retain model ownership in perpetuity, so they can turn generative AI into one of their most valuable assets.

About the team

The Cloud Platform team owns the production inferencing service that serves SambaNova's models to customers on RDU accelerators, including capacity planning, deployment, monitoring, and incident response across regions in the United States, Asia, Europe, and Latin America.

About the role

As a Principal Cloud Platform Engineer, you will be specializing in our AI Inferencing Service and will be the guardian of its reliability, performance, and scalability. You will bridge the gap between software development and operations, applying an engineering mindset to solve operational challenges. Your primary focus will be ensuring our inference endpoints have exceptional uptime, low-latency response times, and efficient resource utilization, directly impacting the experience of our customers and the success of our AI products. This role includes participating in a shared on-call rotation to maintain 24/7 service reliability. 

Responsibilities 

Some of your responsibilities will include:

  • Shared ownership of the production inferencing service across regions, covering availability, latency, performance, change management, and capacity planning
  • Standing-up and automating AI infrastructure in new regions
  • Participating in a shared primary/secondary on-call rotation, and leading incident response
  • Building monitoring, alerting, and dashboards in Prometheus, Grafana, and Datadog for service health, model latency and throughput, and accelerator utilization
  • Finding and eliminating performance bottlenecks
  • Designing auto-scaling policies that handle variable inference loads
  • Managing cloud and on-prem infrastructure as code in Terraform and Ansible
  • Building CI/CD pipelines that safely deploy new model versions and service updates
  • Forecasting infrastructure needs against the product roadmap and usage trends, and working with finance to manage cloud spend
  • Defining and reporting on SLOs and SLIs for the inferencing platform, using that data to prioritize reliability work
Required Qualifications
  • B.S. in Computer Science, Computer Engineering, or related field
  • 5+ years of experience in a Site Reliability Engineering, DevOps
  • Experience supporting a large-scale, customer-facing service in a public cloud environment (AWS, GCP, Azure)
  • Strong programming and scripting skills in languages like Python, Go, Rust, or Java
  • Proven experience with containerization and orchestration technologies (Docker and Kubernetes)
  • Deep understanding of monitoring and observability principles and tools (e.g., Prometheus, Grafana, ELK Stack, Datadog)
  • Experience with Infrastructure as Code (e.g., Terraform, CloudFormation)
  • Experience with CI/CD principles and tools (e.g., Jenkins, GitHub Actions, ArgoCD)
  • Strong Linux/Unix system administration fundamentals
Preferred Qualifications
  • Experience in a hybrid environment bridging cloud and on-premise/data center infrastructure.
  • Direct experience supporting ML/AI inferencing services in production.
  • Familiarity with GPU-accelerated computing and optimizing workloads for NVIDIA GPUs for purposes of mapping to RDUs.
  • Knowledge of model serving frameworks like vLLM, SGLang or Ray.
  • Understanding of MLOps principles and practices.
  • Experience with managing and tuning databases (SQL or NoSQL) and caching systems (Redis, Memcached)

Base Salary Range:

Base Pay Range
$210,000$280,000 USD

Submission Guidelines
Please note that in order to be considered an applicant for any position at SambaNova Systems, you must submit an application form for each position for which you believe you are qualified. 

EEO Policy
SambaNova Systems is an Equal Opportunity/Affirmative Action Employer. All qualified applicants will receive consideration for employment without regard basis of age (40 and over), color, disability, gender identity, genetic information, marital status, military or veteran status, national origin/ancestry, race, religion, creed, sex (including pregnancy, childbirth, breastfeeding), sexual orientation, and any other applicable status protected by federal, state, or local laws.

Benefits Summary for US-Based, Full-Time Employment Positions
SambaNova offers a competitive total rewards package, including the base salary, plus equity and benefits. We cover 95% premium coverage for employee medical insurance, and 77% premium coverage for dependents and offer a Health Savings Account (HSA) with employer contribution. We also offer Dental, Vision, Short/Long term Disability, Basic Life, Voluntary Life, and AD&D insurance plans in addition to Flexible Spending Account (FSA) options like Health Care, Limited Purpose, and Dependent Care. Our library of well-being benefits available to you and your dependents includes a full subscription to Headspace, Gympass+ membership with access to physical gyms, One Medical membership, counseling services with an Employee Assistance Program, and much more.

Skills Required

  • Bachelor's degree in Computer Science, Engineering, or related field (or equivalent practical experience).
  • 3-5+ years experience in Site Reliability Engineering, DevOps, or similar role supporting large-scale, customer-facing services in public cloud (AWS, GCP, or Azure).
  • Strong programming/scripting skills in Python, Go, or Java.
  • Proven experience with containerization and orchestration (Docker, Kubernetes).
  • Deep understanding of monitoring and observability tools and principles (Prometheus, Grafana, ELK Stack, Datadog).
  • Solid experience with Infrastructure as Code (e.g., Terraform, CloudFormation).
  • Familiarity with CI/CD principles and tools (e.g., Jenkins, GitHub Actions, ArgoCD).
  • Excellent problem-solving skills and systematic approach to troubleshooting distributed systems.
  • Participate in a shared on-call rotation to provide 24/7 support for the inference service.
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Palo Alto, CA
500 Employees
Year Founded: 2017

What We Do

AI is changing the world and at SambaNova, we believe that you don’t need unlimited resources to take advantage of the most advanced, valuable AI capabilities - capabilities that are helping organizations explore the universe, find cures for cancer, and giving companies access to insights that provide a competitive edge. We deliver the world’s fastest and only complete AI solution for enterprises and governments with world-record inference performance and accuracy. Powered by the SambaNova SN40L Reconfigurable Dataflow Unit (RDU), organizations can build a technology backbone for the next decade of AI innovation with SambaNova Suite. Our fully integrated hardware-software system, DataScale®, enables organizations to train, fine-tune, and deploy the most demanding AI workloads using the largest and most challenging models. Most recently, with the launch of our newest offering, SambaNova Cloud, developers can supercharge AI-powered applications on Llama 3.2 models. SambaNova was founded in 2017 in Palo Alto, California, by a group of industry luminaries, business leaders, and world-class innovators who understand AI. Today, we’ve built an incredibly smart and motivated team dedicated to making a lasting impact on the industry and equipping our customers to thrive in the new era of AI.

Why Work With Us

As a talent first company, we aim to hire the greatest and most innovative minds in the industry- driving the next generation of AI computing where no barrier is too high and the possibilities are truly limitless. We encourage our peers to take risks and take the initiative to make a lasting impact on the AI and ML industries.

Gallery

Gallery

Similar Jobs

Hybrid
4 Locations
289097 Employees

Coursera + Udemy  Logo Coursera + Udemy

Fp&a Manager

Artificial Intelligence • Consumer Web • Edtech • Enterprise Web • HR Tech • Social Impact • Generative AI
Remote or Hybrid
United States
1500 Employees
111K-162K Annually

Citizens Logo Citizens

Wealth Advisor - Lansdale, PA

Digital Media • Fintech • Information Technology • Machine Learning • Financial Services • Cybersecurity • Automation
In-Office or Remote
2 Locations
17000 Employees
105K-250K Annually

Sprout Social Logo Sprout Social

Lead GTM Financial Planning & Analysis Analyst

Marketing Tech • Social Media • Software • Analytics • Business Intelligence
Easy Apply
Remote or Hybrid
US
1400 Employees
114K-188K Annually

Similar Companies Hiring

LTX Thumbnail
Robotics • Conversational AI • Generative AI
Jerusalem, Israel
200 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Revel.io Thumbnail
Aerospace • Hardware • Robotics • Software
US
50 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account