Cloud Platform Architect

Reposted 25 Days Ago
2 Locations
In-Office
245K-325K Annually
Senior level
Artificial Intelligence • Hardware • Machine Learning • Natural Language Processing • Software • Generative AI
SambaNova is the #1 platform for business AI.
The Role
As a Senior Cloud SRE, you'll ensure the reliability and performance of our AI inferencing service, handle incident management, and optimize cloud infrastructure and resource utilization.
Summary Generated by Built In

SambaNova is a leader in next-generation AI infrastructure, delivering a full-stack inference platform for customers worldwide. At the core of SambaNova's technology is the RDU (Reconfigurable Dataflow Unit) — a chip built on a dataflow architecture rather than the traditional GPU model. Its decode performance is especially strong for agentic workloads like multi-turn agents, code generation, and long-running applications. RDUs are packaged into SambaRack, rack-scale hardware that lets customers deploy state-of-the-art models with better performance, greater energy efficiency, and faster time to value.

About the team

The Cloud Platform team owns the production inferencing service that serves SambaNova's models to customers on RDU accelerators, including capacity planning, deployment, monitoring, and incident response across regions in the United States, Asia, Europe, and Latin America.

About the role

The Cloud Operations team is seeking an experienced engineering leader to scale the platform our internal and external customers use to access SambaNova RDUs.

Responsibilities

In this role, you'll architecting our next-generation system from the ground up, running the Kubernetes infrastructure that powers some of the most advanced AI workloads in the industry, and bridging multi-cloud and on-prem environments in ways no generic SaaS company can offer. Your work will directly impact the productivity of every engineer at SambaNova and by extension, the speed at which we ship the future of AI computing.

  • Architect, build, and maintain our next-generation internal developer platform, automating and streamlining our cloud and on-prem infrastructure
  • Design, write, and manage Terraform modules to provision and manage resources across AWS, GCP, and Azure, ensuring consistency and reproducibility
  • Build and manage highly available, secure, and performant Kubernetes clusters that serve as the primary runtime for our diverse AI workloads
  • Design and implement robust networking solutions (VPCs, load balancers, firewalls, service meshes) that seamlessly connect our multi-cloud and hybrid environments
  • Collaborate with AI and software engineering teams to understand their needs, provide golden paths to production, and build internal tools that accelerate their development cycles
  • Implement best practices for observability (monitoring, logging, tracing) to ensure system reliability and performance, and participate in on-call rotation
Required Qualifications
  • 7+ years of experience in DevOps, Site Reliability Engineering (SRE), or Cloud Infrastructure roles
  • Proficiency in at least one programming language (e.g., Python, Go, Rust)
  • Expertise with Kubernetes (EKS, GKE, or self-managed) in production environments - pods, operators, CRDs, CNIs, etc.
  • Expertise with Infrastructure as Code with the ability to manage complex, multi-cloud environments
  • Strong proficiency with at least one major cloud provider (AWS, GCP, or Azure), with a solid understanding of the core services (compute, storage, networking, IAM)
  • Networking fundamentals (TCP/IP, DNS, HTTP, load balancing) and security best practices in the cloud
Preferred Qualifications
  • Experience in a hybrid environment bridging cloud and on-premise/data center infrastructure
  • Experience managing infrastructure for data-intensive or ML/AI workloads
  • Knowledge of building and maintaining CI/CD pipelines (e.g., GitLab CI, Jenkins, ArgoCD)
  • Experience with service mesh technologies (e.g., Istio, Linkerd)
  • Contributions to open-source projects or a public portfolio of code (GitHub)

Base Salary Range:

Base Pay Range
$245,000$325,000 USD

Submission Guidelines
Please note that in order to be considered an applicant for any position at SambaNova Systems, you must submit an application form for each position for which you believe you are qualified. 

EEO Policy
SambaNova Systems is an Equal Opportunity/Affirmative Action Employer. All qualified applicants will receive consideration for employment without regard basis of age (40 and over), color, disability, gender identity, genetic information, marital status, military or veteran status, national origin/ancestry, race, religion, creed, sex (including pregnancy, childbirth, breastfeeding), sexual orientation, and any other applicable status protected by federal, state, or local laws.

Benefits Summary for US-Based, Full-Time Employment Positions
SambaNova offers a competitive total rewards package, including the base salary, plus equity and benefits. We cover 95% premium coverage for employee medical insurance, and 77% premium coverage for dependents and offer a Health Savings Account (HSA) with employer contribution. We also offer Dental, Vision, Short/Long term Disability, Basic Life, Voluntary Life, and AD&D insurance plans in addition to Flexible Spending Account (FSA) options like Health Care, Limited Purpose, and Dependent Care. Our library of well-being benefits available to you and your dependents includes a full subscription to Headspace, Gympass+ membership with access to physical gyms, One Medical membership, counseling services with an Employee Assistance Program, and much more.

Skills Required

  • Bachelor's degree in Computer Science, Engineering, or a related field
  • 5-8+ years of experience in Site Reliability Engineer, DevOps, or related role
  • Strong programming/scripting skills in languages like Python, Go, or Java
  • Proven experience with Docker and Kubernetes
  • Deep understanding of monitoring tools like Prometheus or Grafana
  • Solid experience with Infrastructure as Code like Terraform
  • Familiarity with CI/CD principles and tools
  • Excellent problem-solving skills
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Palo Alto, CA
500 Employees
Year Founded: 2017

What We Do

AI is changing the world and at SambaNova, we believe that you don’t need unlimited resources to take advantage of the most advanced, valuable AI capabilities - capabilities that are helping organizations explore the universe, find cures for cancer, and giving companies access to insights that provide a competitive edge. We deliver the world’s fastest and only complete AI solution for enterprises and governments with world-record inference performance and accuracy. Powered by the SambaNova SN40L Reconfigurable Dataflow Unit (RDU), organizations can build a technology backbone for the next decade of AI innovation with SambaNova Suite. Our fully integrated hardware-software system, DataScale®, enables organizations to train, fine-tune, and deploy the most demanding AI workloads using the largest and most challenging models. Most recently, with the launch of our newest offering, SambaNova Cloud, developers can supercharge AI-powered applications on Llama 3.2 models. SambaNova was founded in 2017 in Palo Alto, California, by a group of industry luminaries, business leaders, and world-class innovators who understand AI. Today, we’ve built an incredibly smart and motivated team dedicated to making a lasting impact on the industry and equipping our customers to thrive in the new era of AI.

Why Work With Us

As a talent first company, we aim to hire the greatest and most innovative minds in the industry- driving the next generation of AI computing where no barrier is too high and the possibilities are truly limitless. We encourage our peers to take risks and take the initiative to make a lasting impact on the AI and ML industries.

Gallery

Gallery

Similar Jobs

FleetPride Logo FleetPride

Architect

Other • Transportation
In-Office
Dallas, TX, USA
3000 Employees

Entrust Logo Entrust

Architect

Information Technology • Security • Software
In-Office
5 Locations
2800 Employees
121K-219K Annually

Liberty Mutual Insurance Logo Liberty Mutual Insurance

Associate Claims Adjuster, Workers Compensation

Artificial Intelligence • Fintech • Insurance • Marketing Tech • Software • Analytics
Hybrid
5 Locations
40000 Employees
50K-78K Annually

Atlassian Logo Atlassian

Account Manager

Cloud • Information Technology • Productivity • Security • Software • App development • Automation
In-Office or Remote
Austin, TX, USA
11000 Employees
112K-175K Annually

Similar Companies Hiring

Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Revel.io Thumbnail
Aerospace • Hardware • Robotics • Software
US
50 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software • Productivity
US
15 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account