Senior Platform Reliability Engineer

Reposted One Month Ago
3 Locations
Hybrid
182K-250K Annually
Senior level
Healthtech • Social Impact • Software
The mental health platform built for personal and professional growth.
The Role
Define and scale reliability practices across the company by creating SLO/SLA frameworks, improving observability, evolving incident response, building self-service tooling and scorecards, and driving cross-team adoption to enable teams to build and operate reliable production systems at scale.
Summary Generated by Built In

About Us:

Grow was born to tackle a critical challenge: making mental health care more effective and accessible for everyone. Since launching in 2020, 15 million sessions have started on Grow, and we've barely scratched the surface. We do it through a three-sided marketplace that empowers providers, augments insurance payors, and serves patients. Every solution we build reveals ten more waiting to be delivered, and that's just what gets us going. We've backed that ambition with more than $328M in funding, including our Series D at a $3B valuation from Sequoia Capital, Transformation Capital, TCV, SignalFire, Menlo Ventures, Goldman Sachs Alternatives, and others.

At Grow, the work you do can literally change someone's life. Your work here doesn't stay on a screen; it impacts whether someone can find, afford, and access quality care. That ambition sets us apart, and offers you an endless runway to impact one of the world's most urgent issues. We don’t have passenger seats here. Everyone is driving something that matters.

This work deserves real commitment. So if you thrive on meeting problems head on, untangling serious complexity, and working at pace with the sharpest yet kindest people around, you've come to the right place.

About the Role

We’re hiring a Senior Platform Reliability Engineer to help define and scale reliability as a first-class capability at Grow. In this role you’ll operate horizontally across the organization, shaping how reliability is understood, measured, and built into the developer experience.

You’ll work closely with other members of the platform team as well as our product engineering teams to establish standards around observability, SLOs/SLAs, and incident response—while also helping translate those standards into self-service tooling and “golden paths” that make it easy for teams to adopt them.

This is a high-impact, highly autonomous role where you’ll drive both cultural and technical change, ultimately enabling teams to independently build and operate reliable systems at scale.

What You'll Work On

You’ll help us establish and scale reliability as a discipline at Grow by:

  • Defining Reliability Standards Establishing frameworks for SLOs/SLAs, error budgets, and operational readiness; helping teams understand what to measure and why it matters.

  • Improving Observability & Measurement Identifying gaps in metrics, logging, and tracing; ensuring services are measurable, debuggable, and aligned with reliability goals.

  • Evolving Incident Response Developing and improving incident response practices, from detection to post-incident learning, and helping teams build sustainable on-call and escalation patterns.

  • Enabling Self-Service Reliability Partnering with the platform team to build tooling and abstractions (e.g., service scorecards, dashboards, templates, golden paths) that make it easy for teams to adopt and stay compliant with reliability standards.

  • Driving Adoption Across Teams Working cross-functionally to educate, influence, and guide engineering teams—scaling reliability practices through a combination of clear standards, strong communication, and developer-friendly systems

Who You Are

  • Experienced in production systems: You have 6+ years of experience operating and improving reliability of production systems at scale.

  • Strong foundation in cloud and infrastructure: You have hands-on experience with AWS, Kubernetes (e.g., EKS), and infrastructure as code tools like Terraform.

  • Deep understanding of reliability principles: You’ve defined or worked with SLOs/SLAs, understand error budgets, and have experience improving reliability through measurement and iteration.

  • Observability expertise: You’ve worked with modern observability tooling (we use DataDog) and understand how to build actionable monitoring systems across metrics, logs, and traces.

  • Systems thinker: You’re able to zoom out, identify patterns across teams and services, and design solutions that scale beyond a single system.

  • Impact-oriented: You focus on outcomes over output and care deeply about improving real reliability outcomes—not just adding processes.

  • Strong communicator and influencer: You can drive change across teams without direct authority, balancing pragmatism with long-term vision.

  • Self-directed: You thrive in ambiguous environments and are comfortable defining problems, proposing solutions, and executing independently.

  • Team player: You collaborate well, communicate with empathy, and enjoy mentoring and learning from others.

Bonus Points

  • You’ve helped introduce or scale reliability practices in a growing organization.

  • You’ve built internal tooling or platforms used by multiple teams.

  • You have experience designing service-level scorecards or compliance/reporting systems.

  • You’ve worked with both SaaS (e.g., DataDog) and self-managed observability stacks.

  • You were previously a product engineer and bring empathy for developer experience.

  • You have experience with database reliability and performance (we use PostgreSQL)

Why This Role Is Exciting

This is a rare opportunity to define what reliability looks like at a growing, scaling engineering organization—and to do it in a way that actually sticks.

You won’t just be responding to incidents or working within a single team. You’ll be shaping how reliability is measured, enforced, and experienced across the entire company. You’ll work alongside your team mates to turn best practices into intuitive, self-service systems that engineers rely on every day.

Your work will directly improve system reliability, reduce incidents, and enable teams to move faster with confidence, ultimately making reliability a built-in property of how we build software at Grow.

Role Details
  • Employment Type: Full Time, Exempt

  • Base Compensation: The base compensation range for this position is $182,000–$250,000 USD Annually.

This is a hybrid role with the expectation to work onsite from our San Francisco, NYC, or Seattle hub location three days per week (Tuesday, Wednesday, and Thursday) and travel 2–3 times per year (e.g., company and department offsites).
The base compensation for this role will vary depending on several factors, including relevant experience, qualifications, and the candidate’s working location.

Full Time Employee Benefits:

  • Health Benefits: Comprehensive medical, dental, vision, life, and disability coverage.

  • Grow for Grow: No cost access to therapy through the Grow platform (available to US employees)

  • Financial Wellness: Retirement savings programs and equity opportunities to help you invest in your future.

  • Flexible Time Off, Paid Holidays & Winter Break: Flexible time off, company paid holidays (which vary by country), and a full company wide Winter Break to rest and recharge

  • Parental Leave: Up to 18 weeks of paid parental leave to support you and your growing family and a new child stipend.

  • Mental Health Mornings/Afternoons: Dedicated weekly flexible time for self care, whether that's therapy, exercise, journaling, time with family, or simply taking a break.

  • Wellness & Development Stipend: Annual stipend to support your personal wellbeing and professional growth, from fitness and education to books and learning resources.

  • Commuter Benefits: Pre tax commuter benefits to support teammates working from one of our hub locations.

  • Home Office & Meals: Support for your home workspace plus meal benefits, with offerings tailored to both remote and hub based employees.

  • Additional Perks: A variety of wellbeing benefits, including wellness memberships, virtual care, pet insurance discounts (where available), and global travel assistance.


Research shows that some groups hesitate to apply unless they meet every qualification. If you’re excited about this role but don’t check every box, we encourage you to apply. At Grow, we value diverse experiences, transferable skills, and the unique strengths each person brings.

Grow Therapy is proud to be an equal opportunity workplace and is an affirmative action employer. We are committed to equal employment opportunity regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity or Veteran status. We also consider qualified applicants regardless of criminal histories, consistent with legal requirements.

Use of AI Tools: We use certain AI and automated tools to support recruitment. These tools help our team review applications, organize interviews, detect potential application integrity issues, and maintain interview records; they do not make final hiring decisions.

Application Review: Ashby's AI-Assisted Application Review helps review application materials against role-specific criteria and may flag potential fraud or integrity issues for recruiter review. All advancement decisions are made by our human recruiting team after independent review. See Ashby's AI Bias Audit Report.

Interview Recording: We use BrightHire to record interviews, supporting note-taking and consistency across candidates. Learn more. Interviews are confidential; you may ask your interviewer to stop recording anytime, or opt out in advance. Choosing not to be recorded will not, by itself, affect your candidacy.

AI-Assisted Recruiter Screen: For select roles, BrightHire's AI features may help structure initial recruiter screens and organize interview notes. All responses are reviewed by our recruiting team; no hiring decisions are made by AI alone.

State-specific notices: Depending on your location, you may have additional rights. Illinois: If AI analyzes a video interview, we'll provide notice, explain how it works and what it evaluates, and obtain consent beforehand. New York City: Where required, we'll provide at least 10 business days’ notice before using an automated employment decision tool and explain how to request an alternative process. Maryland: We do not use facial recognition to create facial templates during interviews unless required notice and written consent are obtained.

Questions, accommodation requests, requests for an alternative selection process, objections to recording, or concerns about AI use? Contact [email protected]. We’re happy to help or offer an alternative way to participate.

Skills Required

  • 6+ years operating and improving reliability of production systems at scale
  • Hands-on experience with AWS
  • Hands-on experience with Kubernetes (e.g., EKS)
  • Experience with infrastructure as code tools like Terraform
  • Defined or worked with SLOs/SLAs and error budgets
  • Observability expertise and experience with DataDog
  • Experience improving reliability through measurement across metrics, logs, and traces
  • Strong communication, influencing, and cross-functional collaboration skills
  • Self-directed, systems thinker, mentors and collaborates with others
  • Hybrid onsite expectation: work onsite three days per week in San Francisco, NYC, or Seattle
  • Travel 2-3 times per year for company and department offsites
  • Experience introducing or scaling reliability practices in a growing organization
  • Built internal tooling or platforms used by multiple teams
  • Experience designing service-level scorecards or compliance/reporting systems
  • Worked with both SaaS and self-managed observability stacks
  • Previously a product engineer (empathy for developer experience)
  • Experience with database reliability and performance (PostgreSQL)

What the Team is Saying

Caroline
Lydia
Sami
Chris
Sasha

Grow Therapy Compensation & Benefits Highlights

  • Healthcare Strength Job postings describe comprehensive medical, dental, and vision coverage for full‑time employees, along with life and disability insurance. Materials also call out no‑cost access to therapy through the platform.
  • Leave & Time Off Breadth Listings and postings mention flexible PTO and a company‑wide winter break.
  • Wellbeing & Lifestyle Benefits Company materials highlight wellness supports such as weekly “Mental Health Mornings/Afternoons,” wellness stipends, in‑office meals/snacks, and home‑office support. Several postings also note a culture of flexibility and connection.

Grow Therapy Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: New York, NY
650 Employees
Year Founded: 2020

What We Do

We bring together clinical expertise, always-on tools, and deep insights so clients can improve their lives, providers can build practices they're proud of, and businesses can offer more effective, more efficient care. We believe mental health is a never-ending pursuit, and we build toward that progress every day. For our team, that means a day's work can literally change someone's life. So, if you thrive on directly impacting lives, meeting complex problems head on, and working at a pace with some of the sharpest, kindest people, you've come to the right place. What We Value Build for our loved ones. This work is personal to us. When you deeply care about someone, you'd run through walls for them. Everything we build is based on that urgency and the standard we'd want for the people closest to us. High bar, high care. We expect excellence from ourselves and each other. Our mission demands it. We also expect ego to take a backseat, humanity to drive a culture of growth and feedback, and everyone to support each other as we chase ambitious goals. Change is the job. We see ambiguity as an opportunity. Everything is changing all the time: the world, the landscape, our company, and ourselves. For us, success is driving change, and resilience is essential to succeed. We act and adapt fast, then adapt again. Our Team Grow was founded by a team from Harvard Medical School, Stripe, and Blackstone, pairing clinical depth with experience building infrastructure at scale. That combination of bold vision and rigorous care shapes how Grow leaders show up today: with a humility that gives everyone else permission to do the same. That same spirit runs through the rest of the team. We've grown fast — from 5 to 600+ teammates in 5 years — and while we're a long way from day one, the mindset hasn't changed: no egos, no silos, no "sounds like a you" problem. People here are sharp, humble, and invested in getting it right, and they show up for each other with the same care they bring to the work. The Backing to Build What's Next We're backed by investors who believe in what we're building, including Sequoia Capital, Goldman Sachs Alternatives, TCV, Transformation Capital, SignalFire, and Plus Capital. With $328 million raised at a $3 billion valuation, this backing reflects something bigger than capital: it's a vote of confidence that Grow is the team equipped to take on one of the hardest, most consequential problems — getting people the mental health care they need.

Gallery

Gallery
Gallery
Gallery
Gallery
Gallery
Gallery
Gallery
Gallery

Grow Therapy Offices

Hybrid Workspace

Employees engage in a combination of remote and on-site work.

We value hybrid and remote work, with our headquarters in New York City and offices in San Francisco, Seattle, and Waterloo, Canada. Our teammates also work from various states from coast to coast.

Typical time on-site: 3 days a week
HQNew York, NY
San Francisco, CA
Seattle, WA
Waterloo, ON
Learn more

Similar Jobs

Grow Therapy Logo Grow Therapy

Senior Manager, Training, Content, and Quality

Healthtech • Social Impact • Software
Remote or Hybrid
USA
650 Employees
130K-190K Annually

Grow Therapy Logo Grow Therapy

Specialist II, Release of Information Operations

Healthtech • Social Impact • Software
Remote or Hybrid
USA
650 Employees
25-26 Hourly

Grow Therapy Logo Grow Therapy

Recruiter

Healthtech • Social Impact • Software
Remote or Hybrid
USA
650 Employees
70-75 Hourly

Grow Therapy Logo Grow Therapy

Staff Engineer

Healthtech • Social Impact • Software
Remote or Hybrid
4 Locations
650 Employees
182K-288K Annually

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account