Engineering Manager, Site Reliability Engineering

Posted 22 Hours Ago
Be an Early Applicant
Foster City, CA, USA
Hybrid
250K-325K Annually
Expert/Leader
Artificial Intelligence • Cloud • Machine Learning • Software • Database • App development • Generative AI
Introducing Replit Agent 4 - built to unlock your creativity. Plan. Design. Build. All at once.
The Role
Lead and grow an SRE team responsible for observability, incident management, load and failure testing, performance engineering, cloud cost and capacity, and rollout infrastructure. Build reliable production platforms, guide safe changes, investigate complex failures, improve system performance, and partner cross-functionally. The role also includes coaching engineers, hiring, developing technical leaders, measuring reliability outcomes, and staying hands-on with software, infrastructure, Kubernetes, telemetry, and AI-assisted development.
Summary Generated by Built In

Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation.

About the Role

Replit enables people to build software with AI. The systems underneath that experience must support safe production changes, measurable reliability, and predictable performance as usage grows.

This Engineering Manager will lead SRE across observability, incident management, load testing, performance engineering, cloud cost and capacity, and rollout infrastructure. You'll lead and grow an existing team that builds and operates production platforms and works hands-on across application and infrastructure boundaries.

This is a software-building leadership role, not simply an incident-management function. You'll help teams ship safely, understand production behavior, and remove performance bottlenecks through concrete engineering improvements. You should be comfortable going deep on a rollout failure or performance investigation while developing technical leaders and sustainable ownership across a distributed team.

What You'll Do
  • Observability. Build and operate metrics, logs, traces, and alerting capabilities. Help teams establish meaningful SLOs and use production telemetry to diagnose problems and verify improvements.

  • Incident Management. Own incident tooling and practices, coordinate cross-team response, and turn incident reviews into engineering improvements that reduce recovery time and repeat failures.

  • Load Testing. Build and maintain load/failure testing capabilities. Validate critical paths under expected demand, quantify headroom, and test recovery and production readiness with service owners.

  • Performance Engineering. Lead deep engagements with internal teams on SLOs and end-to-end performance. Use profiling, telemetry, and load tests to identify bottlenecks and deliver improvements with service owners—not just recommendations.

  • Stay technically engaged. Review designs and production changes, debug difficult failure modes, and use AI coding tools—including Replit—to prototype and automate. Apply rigorous review and verification to AI-generated changes.

  • Build and grow a high-ownership engineering team. Coach engineers, develop technical leaders, manage performance, and hire against agreed needs. Make distributed collaboration, mentoring, and backup coverage deliberate rather than relying on a few permanent escalation points.

  • Measure outcomes and close the loop. Track rollout safety, recovery time, repeat incidents, critical-path latency/throughput, test coverage, and improvements arising from cost/capacity analysis. Agree success measures and continuing ownership with partner teams.

What You'll Bring
  • Demonstrated engineering management. You have led and developed engineers, made prioritization and performance decisions, hired thoughtfully, and delivered through a team—not only acted as its strongest individual contributor.

  • Software-oriented production systems depth. You have built and operated distributed systems or reliability platforms and can reason across deployment behavior, Kubernetes, telemetry, service dependencies, and recovery mechanisms.

  • Safe-change and performance judgment. You have led consequential migrations or incidents and used measurement to diagnose reliability or performance problems. You can distinguish symptoms from causes and validate fixes under realistic conditions.

  • Platform-product and cross-team judgment. You can build capabilities other teams adopt, lead hands-on engagements without absorbing every service's operations, and make clear tradeoffs among reliability, performance, engineering effort, and cost.

Nice to Have
  • Experience with GitOps or progressive-delivery platforms such as Harness, ArgoCD, or Kargo.

  • Experience with observability, profiling, load-testing, and failure-testing systems, including OpenTelemetry or comparable tooling.

  • Experience with cloud cost attribution, capacity planning, and provider coordination, particularly on GCP.

  • Experience growing distributed teams and using AI tools to increase engineering output while preserving production safeguards.

Full-Time Employee Benefits Include:

💰 Competitive Salary & Equity

💹 401(k) Program with a 4% match (US Only)

⚕️ Health, Dental, Vision and Life Insurance

🩼 Short Term and Long Term Disability

🚼 Paid Parental, Medical, Caregiver Leave

🏝 Flexible Time Off (FTO) + Holidays

🚗 Commuter Benefits (In-Office & US Only)

📱 Monthly Wellness Stipend

🧑‍💻 Autonomous Work Environment

🖥 In Office Set-Up Reimbursement (In-Office Only)

🚀 Quarterly Team Gatherings

☕ In Office Amenities (In-Office Only)

Want to learn more about what we are up to?

  • Self-driving Company

  • Replit Agent at Scale

  • AI Adoption

  • Build Open-Source Apps

     

Interviewing + Culture at Replit

  • Operating Principles

  • Reasons not to work at Replit

To achieve our mission of making programming more accessible around the world, we need our team to be representative of the world. We welcome your unique perspective and experiences in shaping this product. We encourage people from all kinds of backgrounds to apply, including and especially candidates from underrepresented and non-traditional backgrounds.

Skills Required

  • Demonstrated experience managing, developing, hiring, and performance-managing software engineers
  • Experience building and operating distributed systems or software-oriented reliability platforms
  • Strong understanding of deployment behavior, Kubernetes, telemetry, service dependencies, and recovery mechanisms
  • Experience leading consequential migrations or incidents and diagnosing reliability or performance problems using measurement
  • Ability to build platform capabilities adopted by other teams and make tradeoffs among reliability, performance, engineering effort, and cost
  • Experience with GitOps or progressive-delivery platforms such as Harness, ArgoCD, or Kargo
  • Experience with observability, profiling, load-testing, and failure-testing systems, including OpenTelemetry or comparable tools
  • Experience with cloud cost attribution, capacity planning, and provider coordination, particularly on GCP
  • Experience growing distributed teams and using AI tools while preserving production safeguards

Replit Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Replit and has not been reviewed or approved by Replit.

  • Parental & Family Support — Parental leave spans up to 24 weeks for birth parents with additional paid medical, bonding, and family‑care leave. These policies are characterized as generous compared to common U.S. tech norms.
  • Healthcare Strength — Medical, dental, and vision coverage are paired with an employer‑funded HSA plus life and disability insurance. Options like FSA/DFSA and mental‑health support further reinforce the health offering.
  • Wellbeing & Lifestyle Benefits — Monthly wellness stipends, learning and development funds, and recognition programs add ongoing value beyond cash compensation. Office meals, commuter benefits, and onsite amenities enhance day‑to‑day experience for those near the HQ.

Replit Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: San Francisco, CA
300 Employees
Year Founded: 2016

What We Do

Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide and over 500,000 business users, Replit is democratizing software development by removing traditional barriers to application creation.

Similar Jobs

In-Office
Berkeley, CA, USA
577 Employees
170K-223K Annually
Hybrid
Mountain View, CA, USA
2359 Employees
298K-368K Annually
In-Office
4 Locations
6000 Employees
204K-306K Annually
In-Office
4 Locations
6000 Employees
232K-319K Annually

Similar Companies Hiring

Revel Thumbnail
Aerospace • Hardware • Robotics • Software
Marina Del Rey, California
60 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software
New York, New York
30 Employees
Vega Thumbnail
Artificial Intelligence • Automotive • Insurance • Transportation
US
43 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account