Senior Platform Engineer

Posted Yesterday
Be an Early Applicant
2 Locations
Hybrid
130K-190K Annually
Senior level
Cloud • Software
The Role
Own and scale CloudZero’s Kafka and Kubernetes platforms, including Amazon EKS operations, reliability, monitoring, governance, and performance tuning. Build Infrastructure as Code, GitOps, and CI/CD workflows using tools such as Terraform and GitHub Actions. Partner with software engineers to improve application reliability and developer experience, automate infrastructure, and support production systems through on-call responsibilities.
Summary Generated by Built In
About the Role

CloudZero is growing fast. Our customer base is expanding, the data challenges we're solving are getting more complex, and the platform is scaling to match. As we look to the next evolution of CloudZero's infrastructure, we're looking for a Senior Platform Engineer who can help shape what comes next.

You'll own our core messaging and container orchestration platforms, bringing deep technical expertise to decisions that will shape how we build, scale, and operate our infrastructure. You'll have the autonomy to identify opportunities, determine the right solutions, and drive improvements in reliability, performance, and developer experience. Rather than simply maintaining what exists today, you'll help define the foundations our engineering organization will rely on as we continue to grow.

In this role, you'll lead the operations, scaling, and tuning of our Kafka clusters and Kubernetes environments. Because we leverage Amazon EKS, you won't be bogged down with from-scratch control plane management. Instead, you'll focus on optimizing performance, implementing GitOps best practices, and empowering our engineering teams to ship code and process data seamlessly.

What You'll Do
  • Lead the operations and reliability of our Kafka platform, including deployment, scaling, and topic governance, and define the SLOs for availability and latency

  • Administer and optimize our Amazon EKS environments, handling cluster scaling, node group management, namespace provisioning, and overall container resilience

  • Implement comprehensive monitoring, logging, and alerting for both Kafka and Kubernetes using Sumo Logic to catch bottlenecks before they impact the business

  • Partner closely with software engineering teams to help them design reliable, scalable applications that effectively leverage Kubernetes and Kafka

  • Manage our cloud environments using Infrastructure as Code (e.g., Terraform, CloudFormation, Pulumi) integrated seamlessly with our GitHub Actions workflows

  • Build and maintain our CI/CD pipelines and infrastructure deployments using GitHub Actions, ensuring all configuration changes are version-controlled, automated, and predictable

  • On call experience

What You Bring
  • Production Kafka experience, with a deep understanding of Kafka architecture, broker configuration, topic management, replication, and troubleshooting in high-throughput environments

  • A proven track record running containerized workloads in production. Strong experience with any Kubernetes environment is great, whether that's EKS, GKE, AKS, or self-hosted

  • Familiarity with core cloud infrastructure concepts (compute, networking, IAM). While we operate on AWS, strong experience with any major cloud provider (GCP, Azure, etc.) is perfectly fine

  • Solid experience with Terraform, CloudFormation, or similar Infrastructure as Code tools to manage infrastructure predictably

  • Ability to write automation and tooling in Python, Go, or Node, and Bash

  • Some experience designing and operating GitOps workflows using GitHub Actions or something similar

Nice to Have
  • Experience with Kafka ecosystem tools (Kafka Connect, Schema Registry, ksqlDB)

  • Experience writing custom GitHub Actions or composite actions

  • Experience running stateful workloads and data infrastructure on Kubernetes

  • Passion for Developer Experience (DevEx) and optimizing CI/CD workflows to improve engineering velocity

About CloudZero
CloudZero is the AI ROI Company. We built the financial control plane for AI: the system finance, IT, and engineering use to connect every AI dollar to the outcome it produced. Across every provider. In real time.
AI spend is the fastest-growing line on enterprise P&Ls and the least understood. Only 14% of CFOs can prove AI ROI today. CloudZero answers the question no one else can: what did it cost to produce this outcome, for this customer, on this model.
The largest cloud spenders on the planet already run on CloudZero, including Coinbase, Duolingo, DoorDash, and Shutterstock. We processed 14 trillion billing events in the last twelve months. We're the first listed partner on Anthropic's cost and usage API. We've raised over $119 million, including a $56 million Series C backed by leading venture capital firms.

Why Join Our Team?
At CloudZero, you’ll find a collaborative, fast-moving environment where your work makes a direct impact. We’re a team that values ownership, creativity, and curiosity — and we’re tackling some of the most complex challenges in the cloud space. If you’re excited by working with cutting-edge technology, driving meaningful outcomes, and growing with a company that’s scaling fast, we’d love to hear from you!

Skills Required

  • Production experience with Kafka, including architecture, broker configuration, topic management, replication, and troubleshooting high-throughput environments
  • Production experience running containerized workloads in Kubernetes or a comparable Kubernetes environment
  • Familiarity with cloud infrastructure concepts including compute, networking, and IAM
  • Experience with Terraform, CloudFormation, or similar Infrastructure as Code tools
  • Ability to write automation and tooling in Python, Go, Node.js, and Bash
  • Experience designing and operating GitOps workflows using GitHub Actions or similar tools
  • On-call experience
  • Experience with Kafka ecosystem tools such as Kafka Connect, Schema Registry, or ksqlDB
  • Experience writing custom or composite GitHub Actions
  • Experience running stateful workloads and data infrastructure on Kubernetes
  • Interest in developer experience and optimizing CI/CD workflows

CloudZero Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about CloudZero and has not been reviewed or approved by CloudZero.

  • Healthcare Strength — Healthcare coverage is described as comprehensive, spanning medical, dental, and vision. This breadth is consistently presented as a core part of the total rewards package.
  • Leave & Time Off Breadth — Paid time off is presented as flexible and generous, with practices like Focus Fridays supporting balance. Remote-first policies and periodic meetups complement the time-off approach.
  • Equity Value & Accessibility — Equity grants are included broadly, giving employees a stake in the company’s success. This equity component is positioned as a meaningful part of total compensation.

CloudZero Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Boston, MA
180 Employees
Year Founded: 2016

What We Do

CloudZero is the only cloud cost intelligence platform that puts engineering in control by connecting technical decisions to business results. CloudZero ingests cost data from AWS and Snowflake, organizes it for analysis, and delivers the insights to engineering teams who can understand how their work is impacting the business. You can answer question like: * Who are my most expensive customers? * Which product, feature, and team is spending the most? * Has the profitability of my product changed quarter over quarter? The outcome is real-time intelligence that helps companies control their cost of goods sold (COGS) and gross margins — aligning engineering and finance teams once and for all.

Similar Jobs

DraftKings Logo DraftKings

Senior Platform Engineer

Digital Media • Gaming • Information Technology • Software • Sports • Esports • Big Data Analytics
Hybrid
Boston, MA, USA
6400 Employees
120K-149K Annually

Capital One Logo Capital One

Artificial Intelligence Engineer

Fintech • Machine Learning • Payments • Software • Financial Services
Remote or Hybrid
5 Locations
55000 Employees
286K-392K Annually

CrowdStrike Logo CrowdStrike

Senior Engineer

Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity
Remote or Hybrid
USA
11000 Employees
140K-215K Annually

Jackpocket Logo Jackpocket

Senior Platform Engineer

Consumer Web • Gaming • Mobile • News + Entertainment • Software
Hybrid
Boston, MA, USA
330 Employees
120K-149K Annually

Similar Companies Hiring

Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Revel Thumbnail
Aerospace • Hardware • Robotics • Software
Marina Del Rey, California
60 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software
New York, New York
30 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account