Senior Site Reliability Engineer

Posted 14 Days Ago
Be an Early Applicant
DKI Jakarta, IDN
In-Office
Senior level
Machine Learning • Analytics
The Role
Own and improve Rabbit’s SRE foundation across GCP, Terraform, CI/CD, observability, incident response, and deployment safety. Build automated validation, progressive rollouts, recovery mechanisms, SLOs, and AI-agent workflows that reduce operational effort while preserving reliability. Strengthen cloud infrastructure, networking, IAM, capacity management, and cost efficiency, while developing maintainable tooling in Go or Python and driving lasting improvements from production incidents.
Summary Generated by Built In
The Role

Help Rabbit build and ship faster with AI — safely, securely and reliably.

Rabbit runs automated cost optimization across enterprise Google Cloud environments. When we change a customer's BigQuery reservations or rightsize their GKE clusters, those changes need to be correct and dependable. Reliability is central to the trust customers place in our product.

Our foundation is already in place: logging, alerting, automated deployment and Terraform-managed infrastructure. Your mission is to evolve that foundation for an AI-accelerated engineering team: turn faster implementation into faster, dependable delivery through automated validation, safe releases and rapid feedback.

You'll apply proven SRE practices — SLOs, observability, incident response and deployment safety — to AI-assisted development and agent-driven workflows. The goal is to increase how quickly the team can deliver verified improvements, while controlling production risk and reducing manual operational work.

What You'll Do
  • Make AI-assisted delivery faster and safer. Build automated validation, progressive rollout and recovery mechanisms that let engineers and agents move quickly with clear checks before and after changes reach production.
  • Make reliability measurable. Define and operationalize SLOs, SLIs and error budgets, and use them to guide practical decisions about delivery speed, stability and reliability work.
  • Automate workflows with AI agents. Identify repetitive operational work and build reusable agent-driven workflows for alert triage, incident investigation, routine maintenance and reporting. Add verification and human approval where needed, and measure the reduction in manual effort.
  • Improve observability and feedback. Evolve logging, metrics, tracing and alerting so failures are detected early and changes can be traced, investigated and verified.
  • Turn incidents into lasting improvements. Improve runbooks, investigation and blameless postmortems, and translate recurring problems into tests, safeguards and automation.
  • Keep infrastructure reproducible. Extend our Terraform and delivery tooling so environments remain consistent and changes stay reviewable as the platform grows.
  • Improve GCP reliability and efficiency. Strengthen our cloud infrastructure, networking, access controls and capacity management, balancing performance, reliability and cost.
  • Use AI to accelerate reliability engineering itself. Build maintainable tooling in Go, Python or a comparable language, and use agents to accelerate investigation, implementation, testing and documentation while verifying their outputs.
How We Work — AI-First, Agentic by Default

AI-assisted engineering is an expectation of this role, not an optional experiment. Claude Code, Cursor and agent-driven workflows are part of how we work, including infrastructure and reliability engineering.

We want someone who actively looks for ways to increase engineering speed with AI and makes those improvements safe to repeat. That means shorter feedback loops, automated checks, traceable changes and recovery paths — not simply generating more code.

You remain accountable for engineering judgment: what to automate, how to verify it, when human approval is needed and when a change should be stopped or rolled back. Success means faster delivery of reliable improvements, less repetitive work and a platform the team can trust.

What You'll BringMust-have
  • 6+ years in SRE, production engineering or infrastructure-heavy backend roles, with hands-on ownership of production systems.
  • Strong production GCP experience, including Cloud Run, networking and IAM. Hands-on Google Cloud experience is required and will be assessed during the interview process.
  • Infrastructure-as-code fluency with Terraform, plus solid experience in CI/CD and deployment safety.
  • Strong observability and troubleshooting skills: you can make systems debuggable, identify root causes and verify that a fix works.
  • Coding ability in Go, Python or a comparable language, with experience building maintainable operational tooling.
  • Experience leading production incident investigation and driving follow-up improvements that prevent recurrence.
  • Practical fluency with AI coding agents and the ability to critically review, test and validate their work. You are motivated to make AI-assisted engineering faster and more dependable.
  • Strong written English and the discipline to collaborate asynchronously with a distributed team.
Nice to have
  • Kubernetes / GKE experience, including deploying, operating, debugging and scaling containerized services.
  • Hands-on Datadog experience, including dashboards, monitors, logs, APM and distributed tracing.
  • Experience improving delivery speed and safety through progressive delivery, policy-as-code and automated rollback.
  • GCP cost management or FinOps experience.
  • Security experience is a plus.
Why This Role
  • Shape how an AI-first engineering team scales. Build the reliability practices and automation that let Rabbit turn faster development into dependable customer outcomes.
  • Work on systems with real customer impact. Rabbit operates inside enterprise GCP environments, where reliability and security directly affect customer trust.
  • Own meaningful improvements. Work in a small team with short decision paths and end-to-end ownership, supported by appropriate review and production safeguards.
  • Use AI as an engineering multiplier. Apply agentic tools to infrastructure, delivery and operations, and help define how we measure and improve their impact.

Skills Required

  • 6+ years in SRE, production engineering, or infrastructure-heavy backend roles with hands-on production-system ownership
  • Strong production Google Cloud experience, including Cloud Run, networking, and IAM
  • Infrastructure-as-code fluency with Terraform
  • Solid CI/CD and deployment-safety experience
  • Strong observability and troubleshooting skills, including root-cause identification and fix verification
  • Coding ability in Go, Python, or a comparable language, with experience building maintainable operational tooling
  • Experience leading production incident investigations and driving preventive follow-up improvements
  • Practical fluency with AI coding agents and ability to review, test, and validate their work
  • Strong written English and ability to collaborate asynchronously with a distributed team
  • Kubernetes or GKE experience deploying, operating, debugging, and scaling containerized services
  • Hands-on Datadog experience with dashboards, monitors, logs, APM, and distributed tracing
  • Experience with progressive delivery, policy-as-code, and automated rollback
  • GCP cost management or FinOps experience
  • Security experience
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Budapest
182 Employees
Year Founded: 2009

What We Do

Let’s become partners, and move your business to the cloud together! Aliz helps you reach your business goals with Machine Learning and data solutions – powered by Google Cloud. Because of our strong partnership with Google, we have access to the most modern technologies and developments in the industry; and those resources combined with our dedicated, professional team is ensured to give your company a competitive edge. Data has become the most valuable resource – if you use it well. We help you put your company’s data to use, gain valuable insights, predict future trends, and optimise your business to increase ROI. Join those who are already in the cloud, taking full advantage of their automated processes – and make your business even more successful than it is now. We are committed to delivering quality to our clients. To ensure that your company is provided the best possible solution tailored to your business needs, we use cutting-edge technologies, and we’re very proud to be able to call ourselves the only Premier Google Cloud partners in the CEE region. So far we have helped airlines, retailers, telco companies, and many more high-profile clients to make their business more successful. We are specialised in Machine Learning and Big Data – with our analytics and data warehousing solutions, the sky is the limit.

Similar Jobs

Mondelēz International Logo Mondelēz International

Sr. Tax Analyst, ID

Big Data • Food • Hardware • Machine Learning • Retail • Automation • Manufacturing
Hybrid
DKI Jakarta, IDN
90000 Employees

Tapestry - Coach and Kate Spade Logo Tapestry - Coach and Kate Spade

Senior Manager, Onsite Product Engineering

eCommerce • Fashion • Retail • Sales • Wearables • Design
Remote or Hybrid
Indonesia
16000 Employees

Pfizer Logo Pfizer

Calibration Specialist

Artificial Intelligence • Healthtech • Machine Learning • Natural Language Processing • Biotech • Pharmaceutical
In-Office
DKI Jakarta, IDN
121990 Employees

Takeda Logo Takeda

Training Lead (BioLife Plasma)

Healthtech • Software • Analytics • Biotech • Pharmaceutical • Manufacturing
Hybrid
DKI Jakarta, IDN
50000 Employees

Similar Companies Hiring

Northslope Thumbnail
Artificial Intelligence • Information Technology • Software • Analytics • Consulting • Generative AI
London, GB
100 Employees
Scotch Thumbnail
Artificial Intelligence • eCommerce • Fintech • Payments • Retail • Software • Analytics
US
35 Employees
Milestone Systems Thumbnail
Artificial Intelligence • Security • Software • Analytics • Big Data Analytics
Lake Oswego, OR
1500 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account