Staff Software Engineer, Infrastructure

Posted Yesterday
Be an Early Applicant
2 Locations
Remote or Hybrid
Expert/Leader
Cloud • Machine Learning • Other • Software
The ultimate developer toolkit for in-app chat & activity feeds.
The Role
Staff Software Engineer responsible for designing and building Stream’s high-scale platform and infrastructure. The role includes setting SLOs, leading distributed-systems decisions, developing deployment and debugging tools, improving incident response, optimizing infrastructure costs, and guiding technical direction through design reviews and RFCs. The engineer will work with Go, Python, cloud platforms, Kubernetes, databases, observability systems, and real-time workloads in a small, autonomous team.
Summary Generated by Built In
Staff Software Engineer, Platform & Infrastructure

The role

We are hiring a Staff Software Engineer for our Infrastructure team.

Stream runs behind apps that more than a billion people use: billions of API requests a month, a 99.999% uptime promise, and a small team keeping it all running. Infrastructure here decides how fast our products feel, how well they scale, and how quickly we can launch new ones.

This year we're wrapping up a big move. We're simplifying where and how we host, and rebuilding our data layer on self-hosted, open-source systems that are faster and cheaper to run. That's the groundwork, not the job. Next comes building the infrastructure every Stream product runs on, including the ones we haven't built yet. That means real SLOs, clear costs, calmer on-call, and a platform that helps teams ship faster.

What you build will shape how much Stream spends and how fast it grows, and we'll trust you with big decisions from your first few weeks.

This is a Software Engineering role first. Attitude, ambition and a willingness to experiment across the stack matter more to us than years of Kubernetes or database operations. We expect AI to cover a good share of the DevOps-heavy work, and we expect you to use it that way. What a tool can't give us is system design judgement, strong fundamentals and good code, so that is what we look for.

Platform & Infrastructure is a small, senior team without the support structures of a large organisation. It moves fast and it gets hectic. We would rather say that now than have you find out in month two.

 
What team is focusing on:
  • Runs Stream's global platform. Keeps the edge and backend that serve customers worldwide fast and up, at single-digit millisecond latency on a 99.999% uptime SLA.

  • Shapes Stream's margins. Our hosting and infrastructure decisions directly drive the company's profitability. You will decide where and how we run, and what we stop paying for.

  • Builds tools people actually want to use. Internal tools that change how engineers deploy and debug, and external ones that give customers a clear view of their own traffic and health.

  • Makes on-call boring. Builds a rotation people can sustain, and cuts alert noise until a page means something real is broken.

  • Makes cost visible. Gives every team a clear view of what it spends on infrastructure and why, so cost becomes part of every engineering decision.

What you will do:
  • Define what "good" looks like for the platform. Set latency and uptime SLOs for every product team, and decide what we measure, from error rates and delivery times to how healthy our customers' apps really are.

  • Make the big infrastructure calls. Lead the design reviews and decisions on where and how we run, working with backend, video and moderation engineers on tradeoffs that cross service boundaries.

  • Build the platform tooling. Write the internal tools and services that help engineers ship, deploy and debug faster, plus the docs and AI skills that make launching a new product repeatable.

  • Turn incidents into fixes that stick. Take part in on-call, and make sure every root cause ends in a durable fix, not a workaround.

  • Make cost part of every design. Weigh cost alongside reliability in every decision, and give teams the numbers to do the same.

What we are looking for:
  • A software engineer who has built systems, not only configured them. You write production code in Go, Python or a similar language. Scripting-only backgrounds are not a fit.

  • Strong system design. You can reason about failure modes, capacity and cost in distributed systems, and explain why you chose what you chose.

  • Solid fundamentals in networking, storage and concurrency, and the habit of asking why a system behaves the way it does instead of accepting the default.

  • You have owned high-scale production systems and set technical direction beyond your own work, through design reviews, RFCs and decisions other engineers build on.

  • You experiment. When you hit Kubernetes, a database or a cloud service you have not run before, you dig in, use AI to close the gap, and ship.

  • You already use AI tools every day in your engineering work.

  • You treat cost and reliability as engineering problems, and you can put a number on an outcome you drove.

  • You write clearly. Much of this role is docs and standards that other teams will follow.

  • You are comfortable in a small team, setting direction and reviewing a PR in the same week.

Nice to have experience:
  • Kubernetes in production: cluster architecture, workload design or a migration you led.

  • PostgreSQL at scale, ideally self-hosted or with CloudNativePG: sharding, replication, partitioning tradeoffs.

  • Valkey or Redis at scale, ideally on Kubernetes.

  • Hands-on AWS and GCP, including multi-cloud setups or migrations between providers.

  • Cloud commitment and reservation strategy (committed use discounts, savings plans), or formal FinOps practice.

  • SLO design, alert hygiene and a Prometheus and Grafana based observability stack.

  • Real-time systems: WebSockets, WebRTC, streaming or other persistent-connection workloads.

  • Open source contributions, or writing and talks on cloud, platform or distributed systems.

Our stack
  • Go, gRPC, RocksDB, Python

  • PostgreSQL, RabbitMQ

  • GCP

  • Grafana, Prometheus, ELK (Elasticsearch and Kibana)

  • Jaeger and Tempo for distributed tracing, Datadog

  • Redis, Memcached

  • Claude Code, Cursor

You will thrive here if
  • You want platform and infrastructure problems at a scale most engineers never touch, and the autonomy to set the direction yourself.

  • You would rather try something, measure it and change course than wait for a perfect plan.

  • You ship fast and learn fast, including when it is hectic.

  • You are self-directed and comfortable working with a globally distributed team across time zones.

You probably will not if
  • You want to stay a specialist in one layer of the stack.

  • You want tightly scoped tickets and step-by-step direction.

  • You need a calm, highly predictable environment.

Compensation and benefits

For employees based in the Netherlands:

  • 28 days paid time off plus Dutch public holidays

  • Pension plan

  • Commute covered: an NS Business Card or a company Swapfiets

  • Fitness contribution of up to €60 a month toward a gym or sports club membership

  • Learning and development budget

  • Work from anywhere: up to four weeks a year

  • Full Dutch statutory leave, with Stream topping up paternity leave to 100% pay

  • Catered team lunches in the office, and snacks

  • MacBook Pro and peripherals

  • Relocation support, visa sponsorship and 30% ruling support

  • An office in the centre of Amsterdam

  • Conferences and meetups, as an attendee or a speaker

  • The chance to work on our open source SDKs

Why join Stream?

We're a Series B company with global presence and a team of around 145 people from more than 35 countries.

We're backed by Felicis Ventures, GGV Capital, 01 Advisors, Techstars, and Arthur Ventures, with angels including Dick Costolo (ex-CEO of Twitter), Olivier Pomel (CEO of Datadog), Tom Preston-Werner (co-founder of GitHub), and Nicolas Dessaigne (co-founder of Algolia).

We'll be straight with you: a startup is more demanding than a large company. There's no fixed playbook, you'll own things end to end, and you'll sometimes pick up work outside your title. That's also what makes it a fast place to grow. If you want real ownership and high scale more than structure and a set career ladder, you'll feel at home here.

Hybrid office policy: applicants based (or relocating to) one of our office locations are expected to work according to the applicable local office attendance policy.

Equal opportunity employer statement: Stream provides equal employment opportunities to all employees and applicants for employment and prohibits discrimination and harassment of any type without regard to race, color, religion, age, sex, national origin, disability status, genetics, protected veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by federal, state or local laws.

This policy applies to all terms and conditions of employment, including recruiting, hiring, placement, promotion, termination, layoff, recall, transfer, leaves of absence, compensation and training.

Note for external recruiters: We currently have this role covered and do not accept unsolicited agency resumes. We are not responsible for any fees related to unsolicited resumes.

Skills Required

  • Production software engineering experience with Go, Python, or a similar programming language; scripting-only backgrounds are not suitable.
  • Strong system design skills, including reasoning about failure modes, capacity, cost, and distributed systems tradeoffs.
  • Strong fundamentals in networking, storage, and concurrency.
  • Experience owning high-scale production systems.
  • Experience setting technical direction through design reviews, RFCs, and decisions adopted by other engineers.
  • Ability to learn unfamiliar Kubernetes, database, and cloud technologies and use AI tools to close knowledge gaps.
  • Daily use of AI tools in engineering work.
  • Ability to evaluate and quantify cost and reliability outcomes.
  • Strong written communication skills for documentation and engineering standards.
  • Comfort working autonomously on a small team and contributing to both technical direction and code reviews.
  • Kubernetes production experience, including cluster architecture, workload design, or migrations.
  • PostgreSQL at scale, ideally self-hosted or using CloudNativePG, including sharding, replication, or partitioning.
  • Valkey or Redis at scale, ideally on Kubernetes.
  • Hands-on AWS and GCP experience, including multi-cloud environments or migrations.
  • Cloud commitment and reservation strategy or formal FinOps experience.
  • SLO design, alert hygiene, and Prometheus and Grafana observability experience.
  • Experience with real-time systems such as WebSockets, WebRTC, streaming, or persistent-connection workloads.
  • Open-source contributions or technical writing and speaking on cloud, platform, or distributed systems.

Stream Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Stream and has not been reviewed or approved by Stream.

  • Fair & Transparent Compensation — Salary ranges are openly shared by role and location, providing clarity into pay bands. Compensation is often characterized as market-aligned for many technical and sales roles.
  • Healthcare Strength — Healthcare offerings include comprehensive medical, dental, and vision coverage. In some locations, employer contributions toward premiums are described as strong.
  • Leave & Time Off Breadth — Paid time off is described as generous, with substantial vacation and holiday allowances. Extended PTO totals and observed national holidays are highlighted across multiple locations.

Stream Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Boulder, CO
140 Employees
Year Founded: 2015

What We Do

APIs and SDKs to Build In-App Chat Video & Feeds Faster. Stream's platform empowers developers with the flexibility and scalability they need to build rich conversations and engaging communities. Stream’s enterprise-grade cloud components make it easy for software teams to build In-App Chat, Video, Moderation & Feeds Faster. Serving more than a billion end users, Stream’s scalable APIs and SDKs come with all the building blocks to ship a custom white-label experience that rivals today’s leading social platforms. Stream also launched Vision Agents - an open-source Video AI framework for building real-time voice and video applications: https://visionagents.ai/. Stream’s back-end infrastructure, beautiful UI kits, and front-end SDKs for iOS, Android, React, React Native, and Flutter combine to form the fastest, most reliable and feature-rich component solution on the market. Stream is headquartered in Boulder, Colora,do with an office in Amsterdam.

Why Work With Us

We are motivated to achieve mastery of our domains, build lasting relationships and be transparent team players.

Gallery

Gallery

Similar Jobs

Affirm Logo Affirm

Staff Software Engineer

Big Data • Fintech • Mobile • Payments • Financial Services
Easy Apply
Remote
Canada
2200 Employees
181K-241K Annually
Remote
2 Locations
498 Employees
Remote
2 Locations
29 Employees
175K-230K Annually

Similar Companies Hiring

Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Revel Thumbnail
Aerospace • Hardware • Robotics • Software
Marina Del Rey, California
60 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software
New York, New York
30 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account