Senior Software Engineer - Reliability, Infrastructure, and Tooling

Reposted 6 Days Ago
30 Locations
In-Office or Remote
135K-300K Annually
Senior level
Artificial Intelligence • Cloud • Information Technology • Software
The Role
The Senior Infrastructure Engineer will manage and scale LiveKit's core infrastructure, focusing on reliability and performance, while working on Golang code and automation for distributed systems.
Summary Generated by Built In
About LiveKit

LiveKit is building the infrastructure layer for the agentic era of computing. Our platform gives developers everything they need to build, test, deploy, scale, and observe AI agents in production. Founded in 2021, LiveKit powers voice and agentic AI applications for OpenAI, Salesforce, Spotify, Meta, and tens of thousands of other developers, collectively facilitating billions of calls each year.

About This Role

We’re hiring Senior Software Engineers to join our team focused on infrastructure development and reliability engineering. This is not an “ops” team by any means. We partner with product dev teams to co-design and develop along with them to ensure that LiveKit systems are reliable, maintainable, and secure. We work on internal tooling to provide a smooth experience for our product dev teams to own and run their workloads on top of our infrastructure and to meet our strict reliability requirements.

Like all teams, we own the ops and maintenance for the systems we work on, but where possible we automate away what we can. Our team also facilitates a healthy oncall rotation and incident management practices, but the rotation is shared with product dev team members to ensure the important production perspective that oncall provides isn’t isolated to just our team.

We support the full range of LiveKit products which provide a fascinating landscape of problems to be solved because they are much more demanding than a simple web app. It includes real time media workloads, hosting of customer agent code in secure sandboxes, and advanced networking requirements, all of which keep us on our toes.

You'll Thrive Here If You:
  • You are tenacious with investigating tricky system level issues.

  • You like the art of observability including quantifying reliability and visualizing it efficiently.

  • You understand the delicate balance between moving fast now and moving fast later.

  • You are able to communicate effectively with partner teams and tactfully handle sometimes contentious topics.

  • You get satisfaction from clean, DRY, error-resistant configuration even when the underlying systems are complex and diverse.

  • You look at the world in terms of signals and control systems.

What You'll Do
  • Ramp on LiveKit's global architecture — CockroachDB, NATS, Nebula, Kubernetes — and map where reliability debt lives

  • Ship product SRE work directly in the product codebase: load balancing, load shedding, instrumentation, scalability, efficiency

  • Build and extend common tooling so product teams can self-service reliability without Infra as a bottleneck

  • Participate in the on-call rotation and help resolve recurring reliability patterns

  • Bring informed systems opinions that improve how the team makes architectural decisions

Who You Are
  • Experience building non-trivial applications (high concurrency, complex control loops, etc).

  • Strong experience with Kubernetes (or equivalent, Borg, etc).

  • Experience with Linux internals and networking.

  • Experience making use of observability tools to debug tricky problems.

  • Experience running large scale globally distributed systems and working with the complex configuration management problems and technical debt that come along with it.

  • Experience with complex production incident handling.

  • Experience running open source tooling (e.g. Kafka, Clickhouse, etc).

Nice to Have
  • Data engineering and analytics.

  • Global layer 3 networking.

  • Experience working with systems that handle long lived load like media.

  • Google SRE or equivalent high-scale background.

  • Dealing with compliance frameworks (e.g. PCI).

Our Commitment to You
  • An opportunity to build something truly impactful to the world

  • Contribute to open source alongside world-class engineers

  • Competitive salary and equity package

  • Health, dental, and vision benefits

  • Flexible vacation policy

LiveKit is an equal opportunity employer and does not discriminate on the basis of any characteristic protected by applicable law. If you require a reasonable accommodation during the application or interview process, please contact [email protected].

Skills Required

  • Experience managing complex multi-region distributed systems
  • Strong software engineering skills with a focus on reliability
  • Familiarity with container orchestration systems, especially Kubernetes
  • Experience with incident management and being an Incident Commander
  • Low level knowledge for troubleshooting latency sensitive workloads
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: San Francisco, CA
83 Employees
Year Founded: 2021

What We Do

LiveKit is an open-source framework and cloud platform for building voice, video, and physical AI agents, providing real-time communication infrastructure.

Similar Jobs

LiveKit Logo LiveKit

Senior Software Engineer

Artificial Intelligence • Information Technology • Internet of Things
In-Office or Remote
30 Locations
34 Employees
135K-300K Annually

Pfizer Logo Pfizer

Audit Lead

Artificial Intelligence • Healthtech • Machine Learning • Natural Language Processing • Biotech • Pharmaceutical
Remote
29 Locations
121990 Employees
163K-272K Annually

Pfizer Logo Pfizer

Senior Manager, IT Service Management (ITSM) Product Lead

Artificial Intelligence • Healthtech • Machine Learning • Natural Language Processing • Biotech • Pharmaceutical
In-Office or Remote
5 Locations
121990 Employees
125K-232K Annually

Mondelēz International Logo Mondelēz International

Cloud Engineer

Big Data • Food • Hardware • Machine Learning • Retail • Automation • Manufacturing
Remote or Hybrid
Greece
90000 Employees

Similar Companies Hiring

Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees
Kepler  Thumbnail
Artificial Intelligence • Fintech • Software
New York, New York
9 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account