Senior Site Reliability Engineer

Posted Yesterday
Hiring Remotely in USA
Remote
Senior level
Artificial Intelligence • Software • Conversational AI • Automation
The Role
Build and improve Replicant’s AI-native platform, including site reliability, CI/CD, developer tooling, observability, incident management, cloud infrastructure, and autonomous-agent harnesses. Own production infrastructure reliability at scale, reduce operational toil, improve deployment workflows, participate in on-call rotation, and shape platform engineering patterns. The role uses TypeScript/Node.js, Python, Terraform, Kubernetes, Helm, GCP, and modern monitoring tools.
Summary Generated by Built In

At Replicant, we believe AI should work for people, starting with customer service. That’s why we built a platform that helps contact centers resolve more requests, proactively identify issues, and improve agent performance with AI-powered conversation intelligence and AI agents that act like your best reps.

Our AI agents handle millions of calls every month for Fortune 500 companies and high-growth innovators. From processing payments to booking appointments and authenticating users, they help customers get what they need instantly, 24/7. Meanwhile, our real-time conversation insights help contact center leaders coach better and improve every interaction.

We are leading the shift from legacy systems to AI-first service, powered by large language models (LLMs) and designed for enterprise scale, security, and empathy. If you’re excited by the potential of LLMs, voice AI, and building category-defining technology with a kind, ambitious team, you’ll love it here.

Our SRE team builds - not just supports - an AI-native platform, and they own exciting domains like platform and harness engineering, site reliability, cloud infrastructure, CI/CD, DevEx, observability, incident management, and COGS (e.g. cloud spend visibility). We're looking for a Site Reliability Engineer who has opinions about how these domains should work and wants agency in shaping where they go. If you are energized by enabling teams to succeed through systems- and patterns-level work, come build Replicant’s platform with us!

What You’ll Do

  • Contribute to patterns, design, and implementation of our domains; help shape the future of platform engineering at Replicant.

  • Build and improve systems that help reduce toil and enable Replicant's production infrastructure to remain available and operable under large-scale, real-time conversational AI traffic.

  • Extend and iterate our agent harness: Unsupervised AI agents are currently used by about 10% of the dev team - help us grow that number. The agent harness includes CI, sandboxes, guardrails, and validation (e.g. agent-first eval loops).

  • Own and improve our CI/CD pipelines and surrounding developer tooling: build and test performance, deployment ergonomics, and paved paths for new services.

  • Participate in on-call rotation and incident management to ensure platform uptime and quality. (SRE owns the base infrastructure, not the applications; non-business-hours pages are rare)

What You'll Bring

  • 6+ years’ experience in software development enablement roles.

  • Solid experience owning CI/CD platforms end to end - including domains like caching, architecture, and developer self-service.

  • Effective use of AI tools such as Claude and Cursor for coding, troubleshooting, and reasoning. You pair these skills with a defensible opinion on where to avoid using AI tools.

  • Familiarity with Node/TypeScript including making code changes (e.g. exposing new metrics), Python and Terraform for automation, and developing in a Kubernetes/Helm ecosystem.

  • Practical experience with observability: logs/metrics/tracing, monitoring/alerting, incident management process, and tooling.

  • Experience working in fully remote teams - tell us how you’ve made one work better.

  • Bonus:

    • Harness engineering experience - building platforms for autonomous agents.

    • Production-at-scale experience with GCP.

    • Telephony and SIP architectures, FreeSWITCH in particular.

Our stack is TypeScript/Node and Python running on Kubernetes - primarily on GCP (we are multi-cloud), with GitLab CI, Helm, Terraform, Datadog, Prometheus, and Grafana.

For all full-time employees, we offer:

🌴 In-person connection that counts: company-wide offsites and smaller team gatherings designed to make remote work feel personal

🖥️ Tech & learning stipend: Conferences, books, courses — interested? We’ll fund them

📍 Remote by design: We’re distributed — no guilt about life events, we trust you to manage your calendar

🏋️ Health & wellness: Flexible vacations, paid sabbatical after 5 years, comprehensive benefits, plus a stipend to support your physical and mental well-being

💸 Compensation that matches your impact: competitive salaries in the company you’re helping to build

📈 Equity with upside: We believe in shared ownership—You’ll own a real piece of a fast-growing AI company

Our Values

Replicant has three core values. It is critical that everyone who joins the team feels excited and moved by these values as every new team member makes an impact on our culture.

Blade Runners: We take ownership and pride to influence the outcomes of our goals. We are successful, and like a Blade Runner, use the tools at our disposal to reach our objectives. We value open and honest communication and proactively seek feedback along the way. We are a company driven to grow and achieve both individually and as a team.

Bread Makers: We are humble and strive toward an egalitarian culture. No task is too big or too small. We work together to achieve our goals and develop our company mission. We believe that the whole is greater than the sum of its parts in everything that we do.

Självdistans (Self-Distance): Självdistans is Swedish for self-distance. It's the ability to critically reflect on oneself and one's relations from an external perspective. With this in mind, we act with objectivity and always remember that we are not our work. There's no perfect science to growing a team or business, but we trust everyone at Replicant to point out our blind spots and humbly admit their own.

Replicant is proud to be an equal opportunity employer. We are committed to fostering an inclusive, diverse and equitable workplace that is built on trust, support and respect. We welcome all individuals and do not discriminate on the basis of gender identity and expression, race, ethnicity, disability, sexual orientation, colour, religion, creed, gender, national origin, age, marital status, pregnancy, sex, citizenship, education, languages spoken or veteran status. Accommodation is available upon request at any point during our recruitment process. If you require an accommodation, please speak to your talent acquisition partner or email us at [email protected] and we’ll work to meet your needs.

Skills Required

  • 6+ years of experience in software development enablement roles
  • End-to-end ownership of CI/CD platforms, including caching, architecture, and developer self-service
  • Effective use of AI tools such as Claude and Cursor for coding, troubleshooting, and reasoning
  • Familiarity with Node.js and TypeScript, including making code changes
  • Experience with Python and Terraform for automation
  • Experience developing in Kubernetes and Helm ecosystems
  • Practical experience with observability, including logs, metrics, tracing, monitoring, alerting, and incident management
  • Experience working in fully remote teams
  • Experience building platforms or harnesses for autonomous agents
  • Production-at-scale experience with GCP
  • Experience with telephony and SIP architectures, particularly FreeSWITCH
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: San Francisco, CA
176 Employees
Year Founded: 2017

What We Do

Replicant is a company that uses AI to automate customer service requests and resolve conversations, enabling contact centers to deliver exceptional customer service and empowering agents to focus on more complex issues.

Similar Jobs

Garner Health Logo Garner Health

Senior Site Reliability Engineer

Big Data • Healthtech • HR Tech • Machine Learning • Software • Telehealth • Big Data Analytics
Easy Apply
Remote
USA
350 Employees
191K-226K Annually

Optum Logo Optum

Senior Site Reliability Engineer

Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
In-Office or Remote
Eden Prairie, MN, USA
160000 Employees
92K-164K Annually

Optum Logo Optum

Senior Site Reliability Engineer

Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
In-Office or Remote
Eden Prairie, MN, USA
160000 Employees
92K-164K Annually

Optum Logo Optum

Senior Site Reliability Engineer

Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
In-Office or Remote
Eden Prairie, MN, USA
160000 Employees
92K-164K Annually

Similar Companies Hiring

LTX Thumbnail
Robotics • Conversational AI • Generative AI
Jerusalem, Israel
200 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Revel.io Thumbnail
Aerospace • Hardware • Robotics • Software
US
50 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account