Senior Software Engineer - Capacity

Posted One Month Ago
Be an Early Applicant
Bellevue, WA, USA
Hybrid
200K-288K Annually
Senior level
Artificial Intelligence • Big Data • Cloud • Machine Learning • Software • Database • Analytics
Let's build a world where data and AI turn possibilities into reality.
The Role
Design and build Snowflake’s multi-cloud capacity platform for CPU and GPU resources across AWS, Azure, and GCP. Own canonical capacity data, demand forecasting, procurement, reservations, allocation, utilization, supply-risk visibility, and cost optimization systems. Collaborate with cloud providers and core services, AI/ML, warehouse, and finance teams. Ensure high availability, reliability, and performance through production support, troubleshooting, on-call participation, and incident management.
Summary Generated by Built In

At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done.

Senior Software Engineer, Capacity Engineering

At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic, fast-moving environments and approach challenges with an experimental mindset, rapidly testing emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done.

Snowflake’s infrastructure is expanding rapidly across AWS, Azure, and GCP. The Capacity team plays a pivotal role in provisioning the cloud resources essential for Snowflake's operations and ongoing growth. Capacity Engineering accurately models demand, forecasts requirements, and delivers optimal CPU and GPU capacity on schedule. We drive hardware cost-efficiency and price/performance while continually maximizing fleet utilization. To achieve this across all major cloud providers, the team is building a centralized, self-serve internal capacity platform. This software-driven system provides early visibility into supply risks, ensures sufficient lead time for capacity deployment, and maintains high utilization across committed cloud resources.

The technical problem spans the full lifecycle. We model demand and supply as first-class data, reconcile heterogeneous provider telemetry and commitments into a single canonical capacity layer, and make the live state of the fleet legible and actionable in real time. That includes CPU and GPU procurement and reservation lifecycle for AI/ML workloads (training, fine-tuning, and model serving), demand forecasting, cloud resource and cost optimization, hardware evolution analysis as new generations become available (price/performance, cross-family flexibility, migration paths), and the allocation and efficiency systems that close utilization gaps with the teams that own those workloads.

We are actively looking for a senior software engineer. If you love solving problems at scale, prefer to write scalable, reliable, and testable software, are an ace troubleshooter, and are deeply technical, then this is the role for you! Snowflake’s growth and multi-cloud footprint in a constrained capacity environment demand real engineering maturity in the systems that plan and land compute. While the domains below describe the shape of our current goals, the engineer will drive the strategy and deliverables for clear company impact.

AS A SENIOR SOFTWARE ENGINEER AT SNOWFLAKE YOU WILL:
  • Design and build the capacity platform that unifies CPU and GPU allocation, procurement, reservation lifecycle, and utilization across all three clouds.

  • Own the canonical capacity data layer: ingest and reconcile demand forecasts, provider supply signals, commitments, and fleet utilization into a single, trustworthy model consumed across the company.

  • Serve as the liaison to Cloud Service Providers managing and integrating vendor relationships into the capacity planning and procurement workflow.

  • Build planning and allocation systems that translate demand into hardware requirements (shape, quantity, region, timing) and surface supply risk early, with real-time visibility into fleet and reservation health.

  • Drive efficiency: instrument utilization across CPU and GPU accelerator workloads, establish price/performance baselines, and build the tooling that recovers stranded capacity and right-sizes commitments.

  • Integrate hardware evolution into the platform: evaluate new CPU and GPU generations and their price/performance, and build the flexibility (backup and cross-family fallbacks) that keeps plans aligned to the hardware roadmap.

  • Partner with core services, warehouse, AI/ML, and finance teams to forecast and procure capacity ahead of launches, support AI/ML workloads reliably, and turn insights into procurement and allocation decisions.

  • Ensure high availability, reliability, and performance of capacity systems by participating in on-call rotations and incident management.

WHAT WE LOOK FOR:
  • 7+ years of industry experience designing, building, and supporting large-scale systems in production.

  • Hands-on experience working with cloud providers on compute cluster and cloud services provisioning (CPU and/or GPU fleets).

  • Experience with capacity planning, procurement, resource management, or efficiency work on systems built on large private clouds or public cloud providers.

  • Deep system and architectural analysis experience to identify actionable performance, availability, and efficiency insights across CPU and GPU accelerator fleets.

  • Proficiency in programming languages such as Go, Python, or Java.

  • Excellent problem-solving skills and ability to troubleshoot complex issues in a production environment.

  • Strong communication skills and the ability to collaborate effectively in a team environment.

  • BS / MS in Computer Science, Engineering or related fields.

  • Experience developing or using observability infrastructure such as OpenTelemetry or Prometheus is a plus.

  • Familiarity with accelerator/GPU fleets, hardware price/performance analysis, or Kubernetes-based compute at scale is a plus.

  • Prior background working with Modeling, Forecasting, Cloud Spend Optimization, and LLMs is a plus.

Snowflake is growing fast, and we’re scaling our team to help enable and accelerate our growth. We are looking for people who share our values, challenge ordinary thinking, and push the pace of innovation while building a future for themselves and Snowflake.

Snowflake is growing fast, and we’re scaling our team to help enable and accelerate our growth. We are looking for people who share our values, challenge ordinary thinking, and push the pace of innovation while building a future for themselves and Snowflake.

How do you want to make your impact?

For jobs located in the United States, please visit the job posting on the Snowflake Careers Site for salary and benefits information: careers.snowflake.com

Skills Required

  • 7+ years of industry experience designing, building, and supporting large-scale production systems
  • Hands-on experience working with cloud providers on compute cluster and cloud services provisioning for CPU and/or GPU fleets
  • Experience with capacity planning, procurement, resource management, or efficiency work on large private or public clouds
  • Deep systems and architectural analysis experience identifying performance, availability, and efficiency insights across CPU and GPU fleets
  • Proficiency in Go, Python, or Java
  • Excellent problem-solving and production troubleshooting skills
  • Strong communication and team collaboration skills
  • BS or MS in Computer Science, Engineering, or a related field
  • Experience developing or using observability infrastructure such as OpenTelemetry or Prometheus
  • Familiarity with accelerator or GPU fleets, hardware price/performance analysis, or Kubernetes-based compute at scale
  • Background in modeling, forecasting, cloud spend optimization, or LLMs

Snowflake Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Snowflake and has not been reviewed or approved by Snowflake.

  • Fair & Transparent Compensation — Pay is often characterized as top‑of‑market across multiple roles. The company also points to a Fair Pay Workplace certification, signaling externally reviewed pay‑equity practices.
  • Equity Value & Accessibility — Equity is a meaningful part of total compensation, with new‑hire grants, refresh potential, and a discounted ESPP with a favorable lookback. Feedback suggests this ownership component materially boosts perceived total rewards.
  • Leave & Time Off Breadth — Parental leave is described as up to 26 weeks paid in the U.S., paired with flexible or generous PTO and multiple leave types. Family‑building benefits and a dedicated parental‑leave hub further expand support.

Snowflake Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Bozeman, MT
9,023 Employees
Year Founded: 2012

What We Do

Snowflake powers the end-to-end data lifecycle – from ingesting and processing data to analyzing and modeling it, to building and sharing data and AI applications – helping engineers, analysts, and leaders innovate faster and achieve more with their data. We're on a mission to empower every enterprise to achieve its full potential through data and AI.

Why Work With Us

Snowflake is where data does more, and so do you. More innovating, more growing, and more collaborating. Here, you’ll find the sweet spot between building big and moving fast, in technology and your career.

Gallery

Gallery

Similar Jobs

In-Office
4 Locations
26259 Employees
75K-215K Annually

Atlassian Logo Atlassian

Research Intern, 2027 Summer U.S.

Cloud • Information Technology • Productivity • Security • Software • App development • Automation
In-Office
Seattle, WA, USA
11000 Employees
40K-55K Hourly

Shield AI Logo Shield AI

Chief Engineer

Aerospace • Artificial Intelligence • Machine Learning • Robotics • Software
In-Office
Seattle, WA, USA
250K-350K Annually

Metropolis Technologies Logo Metropolis Technologies

Head of Growth, Retail

Artificial Intelligence • Computer Vision • Machine Learning • Payments • Real Estate • PropTech
Easy Apply
In-Office
Seattle, WA, USA
23100 Employees
200K-230K Annually

Similar Companies Hiring

Revel Thumbnail
Aerospace • Hardware • Robotics • Software
Marina Del Rey, California
60 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software
New York, New York
30 Employees
Vega Thumbnail
Artificial Intelligence • Automotive • Insurance • Transportation
US
43 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account