Staff Engineer - Distributed Systems

Posted 3 Days Ago
Be an Early Applicant
Hiring Remotely in India
Remote
Expert/Leader
Information Technology • Internet of Things • Marketing Tech
The Role
Own the architecture health and resilience of HighLevel’s billion-scale distributed systems. Identify failure modes, capacity constraints, consistency risks, and cross-service dependencies; approve critical designs; prototype and ship architectural fixes; and improve backpressure, degradation, idempotency, and fault isolation. Partner across teams through design reviews, incident remediation, and postmortems while raising engineering standards for systems operating at massive scale.
Summary Generated by Built In
About HighLevel:
HighLevel is an AI-powered business operating system that gives agencies, entrepreneurs and SMBs the infrastructure to build, automate and scale. Today, HighLevel supports SMBs across 150+ countries, fueling community-driven growth rooted in real customer outcomes.
To date, businesses operating on HighLevel have generated over $7 billion in ecosystem value, demonstrating the impact of shared infrastructure at scale. By centralizing conversations, automation and intelligence into one system, we help businesses move faster, reduce complexity and execute efficiently.
Behind the platform, HighLevel powers more than 4 billion API hits and 2.5 billion message events daily. With 250 terabytes of distributed data, 250+ microservices and over 1 million domain names supported, our architecture is built for performance, resilience and long-term scalability.
Our people
With over 2,000 team members across 10+ countries, HighLevel operates as a global, remote-first organization built for speed and ownership. We value initiative, clarity and execution, creating space for ambitious people to build systems that support millions of businesses worldwide. Here, innovation thrives, ideas are celebrated and people come first, no matter where they call home.
Our impact
Every month, HighLevel enables more than 1.5 billion messages, 200 million leads and 20 million conversations for the more than 1 million businesses we support. Behind those numbers are real people building independence, expanding opportunity and creating measurable impact. We’re proud to be a part of that.
Learn more about us on our YouTube Channel or Blog Posts
 

What are we hiring for?

Thousands of pods. 50+ kinds of deployments. Four databases. Three teams shipping fast. And until now, nobody whose whole job is the system itself.

Our teams are excellent at building products. Each service has owners, each feature has engineers, each database has experts. What we don't have is the person who holds the entire distributed system in their head, who sees that the retry policy in one service and the queue configuration two hops away are, together, a cascading failure waiting for Black Friday.

That's the seat. Your job is to find the failure before it finds production.

You start with Workflows, HighLevel's automation engine, and one of the largest systems in the company: 3.1 billion enrollments and 21.5 billion action executions every month, traffic peaking at 28,000+ requests per second, running across thousands of pods on GCP with Pub/Sub, Cloud Tasks, Redis, and multiple database engines underneath. From there, your blast radius grows, into Conversations (2.6 billion messages a month) and a brand-new Ticketing system being built right now, where you get to make sure it's born right instead of fixed later.

This is not an architect role where you draw boxes and hand them to someone else. And it's not a feature role where you own a backlog. It's the role in between that most companies never create, and most staff engineers spend their careers wishing existed.

Responsibilities:

  • Own the architecture health of a billion-scale distributed system, its failure modes, capacity limits, consistency guarantees, and the interactions between 50+ deployments that no single team can see

  • Approve critical-path designs. Changes that touch the system's core go through you, not as bureaucracy, but as the person accountable for the whole staying sound. When there's a disagreement, you make your case on merit

  • Hunt gaps proactively, single points of failure, unbounded queues, missing idempotency, thundering herds, quiet data-loss windows, and drive the fixes before they become incidents

  • Build the parts nobody else can. No sprint tickets. You prototype the risky architectural bets yourself, ship the remediation after serious incidents, and pair into the gnarliest cross-team bugs, roughly a quarter to a third of your time in code, all of it on the hardest problems

  • Make resilience a property of the system, not a heroic act, degradation strategies, backpressure, isolation boundaries, capacity models that survive 10%+ month-over-month growth

  • Raise the teams around you. Design reviews that teach, post-mortems that change architecture (not just add alerts), and patterns that 80+ engineers build on

  • Set the standard for how AI-assisted engineering works safely on systems this critical, where a bad merge doesn't cost a demo, it costs real businesses their revenue

  • The terrain:

  • Runtime: Node.js (TypeScript), Go, thousands of pods on GKE

  • Messaging & async: GCP Pub/Sub, Cloud Tasks, Redis

  • Storage: MongoDB, Firestore, ClickHouse, ElasticSearch

  • Scale: 21.5B automation actions/month, 2.6B messages/month, 28.5K req/s peaks, ~226B async events across the org

Requirements:

  • 10+ years of engineering experience with deep, hands-on ownership of large-scale distributed systems, comparable scale strongly preferred: hundreds of services or thousands of instances, billions of daily events

  • You've carried sole accountability for a production system through real failures, not adjacent to it, not advising on it. Owned it

  • Deep command of queueing and async architectures, delivery semantics, ordering, backpressure, idempotency, exactly-once myths and at-least-once realities

  • Strong with multiple storage engines (SQL and NoSQL), you reason about consistency models, indexing at scale, and when each engine is the wrong choice

  • Expert-level depth in Redis or comparable in-memory systems, including their failure modes under memory pressure and network partition

  • Production experience on Kubernetes at scale, resource limits, autoscaling behavior, what actually happens when a node pool dies

  • Exceptional design communication, docs, diagrams, and RCAs that drive decisions across multiple teams

  • Fluent in Node.js and/or Go, enough to prototype your own proposals and ship fixes on the critical path

What Success Looks Like

  • The critical paths of Workflows have named owners, capacity models, and tested failure modes, because you made it so

  • Incident count trends down while traffic grows double-digit percent month over month

  • Your design reviews are the ones engineers want their proposals to survive

  • Your scope has expanded on results, more systems, more surface, more trust


What you're built for

  • You've been injured in production. A lot. You've owned distributed systems at serious scale, through the outages, the migrations, the 3 AM discoveries, and every scar changed how you design

  • You've operated at staff scope, whatever your title said, the engineer everyone routed the hardest systems questions to

  • You think in failure modes by default: when you see a design, you instinctively ask what happens at the tail, under partition, at 10x load

  • You write design docs and RCAs that people reference years later, clear trade-offs, honest risks, real recommendations

  • You can disagree with a team and still make them better, influence through rigor and respect, not title

  • You'd rather prevent ten incidents quietly than be the hero of one loudly

Bonus Points

  • You've made AI agents genuinely productive on complex systems, and know how to keep AI-generated code from becoming AI-generated incidents

  • GCP-native experience: Pub/Sub, Cloud Tasks, GKE, Firestore

  • You've done this job before under another name, "the systems person," principal engineer, architect-who-still-codes

  • Experience taking a 0→1 system to production alongside hardening mature ones

 
#LI-Remote #LI-HB1

EEO Statement:

The company is an Equal Opportunity Employer. As an employer subject to affirmative action regulations, we invite you to voluntarily provide the following demographic information. This information is used solely for compliance with government recordkeeping, reporting, and other legal requirements. Providing this information is voluntary and refusal to do so will not affect your application status. This data will be kept separate from your application and will not be used in the hiring decision.
We encourage you to review our Privacy Policy before submitting your application

Skills Required

  • 10+ years of engineering experience
  • Deep hands-on ownership of large-scale distributed systems
  • Sole accountability for a production system through real failures
  • Deep knowledge of queueing and asynchronous architectures, delivery semantics, ordering, backpressure, idempotency, and at-least-once delivery
  • Experience with multiple SQL and NoSQL storage engines, consistency models, and indexing at scale
  • Expert-level experience with Redis or comparable in-memory systems and their failure modes
  • Production Kubernetes experience at scale, including resource limits, autoscaling, and node pool failures
  • Exceptional design communication, including documentation, diagrams, and root-cause analyses
  • Fluency in Node.js and/or Go sufficient to prototype proposals and ship critical-path fixes
  • Comparable-scale experience with hundreds of services or thousands of instances and billions of daily events
  • Experience with GCP-native technologies including Pub/Sub, Cloud Tasks, GKE, and Firestore
  • Experience making AI agents productive on complex systems while controlling risks from AI-generated code
  • Experience taking a system from 0 to 1 into production while hardening mature systems

HighLevel Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about HighLevel and has not been reviewed or approved by HighLevel.

  • Healthcare Strength — Employer materials highlight employer-paid medical and vision for employees, with mental health support and short‑term disability included. Feedback suggests health coverage is a relative bright spot that can raise overall satisfaction even when base pay is not top‑tier.
  • Leave & Time Off Breadth — Flexible PTO, paid family leave, and paid holidays are emphasized alongside a remote‑first setup. Feedback suggests this time‑off approach supports work–life balance and is frequently cited as part of the value proposition.
  • Retirement Support — A company‑matched 401(k) is presented as part of the core package. Feedback suggests retirement support contributes meaningful long‑term value and helps offset tradeoffs in cash compensation.

HighLevel Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Dallas, OR
974 Employees
Year Founded: 2018

What We Do

https://www.gohighlevel.com/quick-links One white-labeled marketing app to rule them all. HighLevel is everything your business needs to succeed! Capture leads using our landing pages, surveys, forms, calendars, inbound phone system & more! Automatically message leads via voicemail, forced calls, SMS, emails, FB Messenger & more! Use our built in tools to collect payments, schedule appointments, and track analytics

Similar Jobs

Remote or Hybrid
3 Locations
3661 Employees
Remote or Hybrid
3 Locations
3661 Employees
Remote or Hybrid
3 Locations
3661 Employees
Remote or Hybrid
3 Locations
3661 Employees

Similar Companies Hiring

NODA AI Thumbnail
Artificial Intelligence • Information Technology • Software • Cybersecurity
Sydney, AU
54 Employees
Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software
New York, New York
30 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account