Staff Site Reliability Engineer (AI Platform)

Posted 10 Days Ago
Be an Early Applicant
Barcelona, Cataluña, ESP
Hybrid
Senior level
Marketing Tech • Social Media
Manychat makes marketing and growth easier for business owners so they grow as big as their dreams.
The Role
Lead SRE role owning and hardening AWS infrastructure and EKS clusters, migrating services to Kubernetes, codifying infra with Terraform and Ansible, improving CI/CD, observability, and platform reliability while partnering with engineering and occasionally supporting rare off-hours incidents.
Summary Generated by Built In
WHO WE ARE 🌍

Creating content that resonates is great — turning that attention into growth is even better. That's what Manychat does.

Our AI-powered automations help creators and brands engage with audiences across Instagram, Messenger, WhatsApp, and TikTok — at the right moment, with the right message, minus the manual work.

We're 400+ people across three continents behind the leading platform for conversational growth, with AI skills shaping how we build, how we work, and how we grow.

WHO WE’RE LOOKING FOR 🌟

Want to shape an AI platform's architecture from day one, instead of just maintaining what someone else built?

Manychat runs AI features for businesses worldwide, and we're hiring an SRE to own the reliability, performance, and cost of that AI infrastructure, and to raise the bar for how our whole engineering org builds on LLMs.

This isn't a classical SRE role with a bit of AI sprinkled on top. We need someone AI-native: you already understand how modern LLM systems behave in production, things like token throughput, provider rate limits, degraded model quality, and inference latency tails, and you treat all of that as a first-class reliability concern, not an afterthought.

Why this role is worth your time:

  • The AI Platform is still young, so you'll shape its architecture, standards, and roadmap from the ground up.
  • The stakes are real. AI features sit right in the critical path of customer-facing automation, so your work actually matters to the business, not just to a dashboard.
  • You'll partner directly with the Head of Infrastructure, with real autonomy and visibility into how your decisions play out.

Sound like the kind of ownership you've been looking for?

WHAT YOU'LL DO 🚀
  • Own reliability and performance of our AI infrastructure: AI Gateway, inference services, and integrations with Amazon Bedrock, Azure OpenAI, and other LLM providers.
  • Design and evolve the AI Gateway: routing, failover between providers, rate limiting, caching, and guardrails.
  • Build observability for AI systems: latency/throughput/error SLOs per model and provider, token-level metrics, quality and drift signals.
  • Drive cost optimization and FinOps for AI workloads: per-feature cost visibility, model right-sizing, caching strategies, provider mix.
  • Run capacity planning and incident response for inference services; write and improve runbooks and postmortems.
  • Scale AI expertise across the org: set standards, review designs, and coach teams shipping LLM-backed features.
TO SHINE IN THIS ROLE 💥

You’ll need:

  • 5+ years in SRE / platform / infrastructure engineering, including production ownership at significant scale.
  • Hands-on experience operating LLM-backed systems in production: provider APIs (Bedrock, OpenAI, Anthropic, or similar), inference pipelines, self-hosted or managed model serving.
  • Deep cloud-native background: AWS, Kubernetes, Terraform/IaC, CI/CD.
  • Strong observability practice (Prometheus/Grafana, OpenTelemetry, or equivalent) and experience defining SLOs for non-deterministic systems.
  • Proven cost-optimization work: you can show where you cut cloud or inference spend and how you made cost visible.
  • Staff-level influence: you've set technical direction beyond your own team and brought others along

It would be great if you have:

  • Experience building or operating an LLM gateway/proxy (e.g., LiteLLM, Kong AI Gateway, custom).
  • Experience with GPU workload optimization, quantization, or serving frameworks (vLLM, TGI, Triton).
  • Experience with eval pipelines and quality monitoring for LLM outputs.
Why this role
  • Green-field ownership: the AI Platform is young; you'll shape its architecture, standards, and roadmap.
  • Real scale and real stakes: AI features sit in the critical path of customer-facing automation.
  • Direct partnership with the Head of Infrastructure; high autonomy and visibility.
WHAT WE OFFER 🤗 

We care deeply about your growth, well-being, and comfort:

  • 🌍 Hybrid onboarding to start work remotely and relocation support for you and your family.
  • 💙 Comprehensive health insurance for both you and your family.
  • 📚 Professional development budget for conference tickets, online courses, and other relevant resources to help you grow.
  • 🫶 Flexible benefits package: no one-size-fits-all perks here. You get a budget and you decide where it goes, from health and wellbeing to family and setting up your home office.
  • 🪴 Hybrid work and generous, flexible time off — planned with your team, not rationed by a rigid quota.
  • 🍽️ In-office perks, including free meals and snacks.
  • 🤝 Company-funded sport activities, annual offsites and team-building events.
  • 🤖 AI isn't a perk here — it's how we work. Claude, OpenAI, and more are on by default from day one, and we back teams in adopting whatever makes them faster.

Manychat is an Equal Opportunity Employer. We’re committed to building a diverse and inclusive team. We do not discriminate against qualified employees or applicants because of race, color, religion, gender identity, sex, sexual preference, sexual identity, pregnancy, national origin, ancestry, citizenship, age, marital status, physical disability, mental disability, medical condition, military status, or any other characteristic protected by local law or ordinance.

This commitment is also reflected through our candidate experience. If you have individual needs that may require an accommodation during the interview process, please indicate this in your application. We will do our best to provide assistance throughout your interview process to ensure you’re set up for success.

Skills Required

  • 5+ years managing Linux in production (Ubuntu, Amazon Linux)
  • Strong experience with Kubernetes (ideally EKS) and Helm
  • Experience with Terraform for infrastructure codification
  • Comfort running and debugging Python workloads in containers
  • AWS infrastructure experience (EC2, ALB/NLB, WAF, IAM, CloudWatch)
  • Hands-on Nginx experience (Ingress and reverse proxy setups)
  • Solid understanding of networking and cloud security best practices
  • Experience building and improving CI/CD pipelines (GitHub Actions)
  • Ownership of observability: Prometheus, Grafana, alerting, on-call readiness
  • Excellent communication skills to explain complex infrastructure to developers
  • Strong Ansible skills
  • PostgreSQL or Amazon RDS tuning and operations experience
  • Deep familiarity with observability tools (Loki, advanced Prometheus/Grafana)
  • Familiarity with PHP production environments
  • Experience with TDD, CI/CD best practices, and agile development
  • Previous SRE-like exposure (resilience, automation, incident tooling)

Manychat Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Manychat and has not been reviewed or approved by Manychat.

  • Fair & Transparent Compensation Compensation is framed as fair and transparent in company materials, with signals that pay levels are broadly viewed positively. Posted ranges and employer-provided bands for multiple roles point to a market-competitive stance.
  • Healthcare Strength Health coverage is described as comprehensive, including medical, dental, and vision plans with tax-advantaged account options. Well-being programs and wellness reimbursements further reinforce the health-focused offering.
  • Leave & Time Off Breadth Time off is portrayed as generous, with flexible time off, paid holidays, sick time, and paid parental leave. These components are highlighted alongside a broader well-being focus.

Manychat Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Austin, TX
232 Employees
Year Founded: 2015

What We Do

Manychat is a leading Chat Marketing platform. We enable businesses and creators to drive more sales and conversions on messaging apps, such as Instagram, WhatsApp, Facebook Messenger, and Telegram, using automation. Trusted by over 1 million brands in 170+ countries, we're an official Meta Business Partner, backed by top investors, including Bessemer Venture Partners.

Why Work With Us

We are a global company with offices in Austin, Barcelona, Yerevan, São Paulo, and Amsterdam. Founded in 2016, we still embrace a startup work culture — making fast decisions, setting ambitions goals, fostering open feedback and transparency.

Gallery

Gallery

Similar Jobs

Cloudflare Logo Cloudflare

Account Executive

Cloud • Information Technology • Security • Software • Cybersecurity
Remote or Hybrid
Spain
4400 Employees

Mondelēz International Logo Mondelēz International

Industrial Engineer Trainee (M/F/D) - 12 months Internship - Granollers (near Barcelona) - December 2025

Big Data • Food • Hardware • Machine Learning • Retail • Automation • Manufacturing
Hybrid
Granollers, Barcelona, Cataluña, ESP
90000 Employees
1K-1K Annually

Perk Logo Perk

Sales Development Representative

Artificial Intelligence • Fintech • Greentech • Sales • Software • Travel • Hospitality
Hybrid
Barcelona, Cataluña, ESP
1800 Employees

Perk Logo Perk

Staff Product Designer

Artificial Intelligence • Fintech • Greentech • Sales • Software • Travel • Hospitality
Hybrid
Barcelona, Cataluña, ESP
1800 Employees

Similar Companies Hiring

ClickMint Thumbnail
AdTech • eCommerce • Marketing Tech • Generative AI
Malibu, CA
9 Employees
PRIMA Thumbnail
Travel • Software • Marketing Tech • Hospitality • eCommerce
US
15 Employees
Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account