Senior Software Engineer, Infrastructure

Posted 8 Days Ago
Be an Early Applicant
San Francisco, CA
In-Office
200K-375K Annually
Senior level
Artificial Intelligence • Software
The Role
Design, build, and operate production infrastructure for high-scale, low-latency systems. Improve system reliability and performance, ensuring documentation and telemetry are in place.
Summary Generated by Built In

About Decagon

Decagon is the leading conversational AI platform empowering every brand to deliver concierge customer experience. Our AI agents provide intelligent, human-like responses across chat, email, and voice, resolving millions of customer inquiries across every language and at any time.

Since coming out of stealth, Decagon has experienced rapid growth. We partner with industry leaders like Hertz, Eventbrite, Duolingo, Oura, Bilt, Curology, and Samsara to redefine customer experience at scale. We've raised over $200M from Bain Capital Ventures, Accel, a16z, BOND Capital, A*, Elad Gil, and notable angels such as the founders of Box, Airtable, Rippling, Okta, Lattice, and Klaviyo.

We’re an in-office company, driven by a shared commitment to excellence and velocity. Our values—customers are everything, relentless momentum, winner’s mindset, and stronger together—shape how we work and grow as a team.

About the Team

The Infrastructure team builds and operates the foundations that power Decagon: networking, data, ML serving, developer platform, and real‑time voice. We partner closely with product, data, and ML to deliver high‑scale, low‑latency systems with clear SLOs and great developer ergonomics.

We organize around five focus areas:

  • Core Infra: The foundational cloud stack—networking, compute, storage, security, and infrastructure‑as‑code—to ensure reliability, scale, and cost efficiency.

  • Data Infra: Streaming/batch data platforms powering analytics/BI and customer‑facing telemetry, including for customer‑managed and on‑prem environments.

  • ML Infra: GPU and model‑serving platforms for LLM inference with multi‑provider routing and support for on‑prem/air‑gapped deployments.

  • Platform (DevEx): CI/CD, paved paths, and core services that make shipping fast, safe, and consistent across teams.

  • Voice Infra: Telephony/WebRTC stack and observability enabling ultra‑low‑latency, high‑quality voice experiences.

Our mission is to deliver magical support experiences — AI agents working alongside humans to resolve issues quickly and accurately.


About the Role

We’re hiring a Senior Infrastructure Engineer to design, build, and operate production infrastructure for high‑scale, low‑latency systems. You’ll own critical services end‑to‑end, improve reliability and performance, and create paved‑paths that let every Decagon engineer ship confidently.


In this role, you will
  • Design and implement critical infrastructure services with strong SLOs, clear runbooks, and actionable telemetry.

  • Partner with research and product teams to architect solutions, set up prototypes, evaluate performance, and scale new features.

  • Tune service latencies: optimize networking paths, apply smart caching/queuing, and tune CPU/memory/I/O for tight p95/p99s.

  • Evolve CI/CD, golden paths, and self‑service tooling to improve developer velocity and safety.

  • Support various deployment architectures for customers with robust observability and upgrade paths.

  • Lead infrastructure‑as‑code (Terraform) and GitOps practices; reduce drift with reusable modules and policy‑as‑code.

  • Participate in on‑call and drive down toil through automation and elimination of recurring issues.


Your background looks something like this
  • 6+ years building and operating production infrastructure at scale.

  • Depth in at least one area across Core/Data/AI‑ML/Platform/Voice, with curiosity to learn the rest.

  • Proven track record meeting high availability and low latency targets (owning SLOs, p95/p99, and load testing).

  • Excellent observability chops (OpenTelemetry, Prometheus/Grafana, Datadog) and incident response (PagerDuty, SLO/error budgets).

  • Clear written communication and the ability to turn ambiguous requirements into simple, reliable designs.


Even better
  • Experience being an early backend/platform/infrastructure engineer at another company

  • Strong Kubernetes experience (GKE/EKS/AKS) and experience across multiple cloud providers (GCP, AWS, and Azure)

  • Experience with customer‑managed deployments


Benefits
  • Medical, dental, and vision

  • Flexible time off

  • Daily lunch/dinner & snacks in the office

Top Skills

Aks
AWS
Azure
Datadog
Eks
GCP
Gke
Grafana
Kubernetes
Opentelemetry
Prometheus
Terraform
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: San Francisco, California
49 Employees

What We Do

Trusted by world-class companies, Decagon is the most advanced AI platform for customer support.

Similar Jobs

NVIDIA Logo NVIDIA

Software Engineer

Artificial Intelligence • Computer Vision • Hardware • Robotics • Metaverse
In-Office
5 Locations
21960 Employees
184K-357K Annually

Retell AI Logo Retell AI

Senior Software Engineer

Artificial Intelligence • Information Technology • Software
In-Office
7 Locations
73 Employees
215K-290K Annually
In-Office
Mountain View, CA, USA
2359 Employees
204K-259K Annually

Jerry Logo Jerry

Senior Software Engineer

Artificial Intelligence • Automotive • Machine Learning • Financial Services
In-Office
7 Locations
296 Employees
160K-200K Annually

Similar Companies Hiring

Standard Template Labs Thumbnail
Software • Information Technology • Artificial Intelligence
New York, NY
10 Employees
PRIMA Thumbnail
Travel • Software • Marketing Tech • Hospitality • eCommerce
US
15 Employees
Scotch Thumbnail
Software • Retail • Payments • Fintech • eCommerce • Artificial Intelligence • Analytics
US
25 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account