Senior Site Reliability Engineer

Posted 22 Days Ago
Be an Early Applicant
Kuala Lumpur, Wilayah Persekutuan Kuala Lumpur, MYS
Hybrid
Senior level
Marketing Tech • Software
The Role
Operate and improve production reliability through SLO management, incident response, observability, safe deployments, infrastructure as code, capacity planning, security controls, and resilience testing. Build automation to reduce operational toil and manage Kubernetes, cloud infrastructure, CI/CD, and secrets. Support production MCP and AI tooling, applying AI-assisted automation to incident triage, log analysis, and runbook generation.
Summary Generated by Built In
Company Description

Carousell Group is the leading recommerce group in Greater Southeast Asia on a mission to inspire the world to start selling, and to make secondhand the first choice. Founded in August 2012 in Singapore, the Group has a leading presence in eight markets under the brands Carousell, Cho Tot, Laku6, Mudah.my, OneKyat, Ox Street, and Refash, serving tens of millions of monthly active users. Carousell is backed by leading investors including Telenor Group, Rakuten Ventures, Naver, STIC Investments and Sequoia Capital India. 

As a team of passionate individuals working together to solve meaningful problems, there is so much more for you to discover in a career with Carousell. Our culture is made up of hiring, developing, and promoting people who embody our values of solving problems for our users; having a mission-first mindset; being relentlessly resourceful; caring deeply; and staying humble to constantly improve. Together as an organisation, we make magic happen.

About Mudah

Mudah.my Sdn. Bhd is Malaysia’s largest digital platform for selling and finding almost anything - from Cars to Cameras, Properties to Pets, Mobile phones to Motorcycles, Treadmills to Textbooks, Bicycles to Beds, Guitars to Golf sets, Plants to Printers, Watches to Washing machines, Tyres to Tablets, Dresses to Drums, Shoes to Shops, Collectibles to Computers, Jobs and more – Semua Pun Mudah! Mudah’s mission is to democratize commerce by empowering everyone, especially individuals and budding entrepreneurs, with a platform of equal opportunity.

Job Description

Responsibilities

  • Operate against our existing SLIs, SLOs and error budgets,  hold services to them, review and adjust targets as systems evolve, and use budget burn to arbitrate between reliability work and feature velocity.

  • Own the incident lifecycle end to end   detection, response, mitigation, blameless postmortems with tracked follow-through   troubleshooting across the whole stack (OS, application, database, cache, network), and mature the on-call rotation around it: alerts tuned for signal, runbooks kept current, MTTD and MTTR trending down.

  • Extend and improve our observability stack   instrument new services, close coverage gaps in metrics, logs and traces, and raise dashboard and alert quality so teams can diagnose their own services.

  • Design and maintain safe release processes: canary, progressive rollout, automated rollback across dev, staging and production.

  • Eliminate toil through automation; treat repetitive manual operations as bugs to be engineered away.

  • Own infrastructure as code   provisioning, configuration, policy, and the documentation around it.

  • Capacity planning, performance tuning and cloud cost efficiency   forecast growth, model headroom, upgrade before saturation.

  • Implement and maintain infrastructure security controls   secrets and credential management in Vault, access control, monitoring and response.

  • Conduct production readiness reviews and resilience testing for new and existing services.

  • Operate our MCP gateway and internal AI tooling as production infrastructure   availability, access control, rate limiting, cost and usage visibility   and apply AI-assisted automation to operational work such as incident triage, log summarisation and runbook generation.

Qualifications

Reliability and infrastructure

  • 5+ years operating production systems at scale in SRE, DevOps or infrastructure engineering.

  • Kubernetes in production   deployment, upgrades, troubleshooting and ongoing maintenance.

  • Strong Linux fundamentals and performance tuning (RHEL / CentOS / Debian / Ubuntu), plus Docker and container networking.

  • Hands-on Google Cloud Platform experience, with infrastructure as code and declarative provisioning (Terraform or equivalent).

  • Secrets and credential management with HashiCorp Vault (or equivalent)   policies, rotation and least-privilege access.

  • CI/CD delivery pipelines on GitHub Actions (or equivalent), including progressive delivery and automated rollback.

  • Bash plus working proficiency in a general-purpose language (Go, Python or similar) for building real tooling.

  • Monitoring and observability with Prometheus, Grafana and Google Cloud Operations   instrumentation, dashboards, alerting and tracing.

  • Able to work independently on large, complex projects with minimal guidance.

AI-assisted operations

  • Practical experience applying AI and LLM-based tooling to engineering or operational workflows, with a clear view of where it helps and where it doesn't.

  • Working familiarity with MCP (Model Context Protocol) and agent harnesses   tool-calling loops, context management, guardrails and failure handling   or the appetite and fundamentals to pick them up quickly.

Additional Information

Why Join Us?

At Mudah, we are evolving towards an AI-first engineering organisation. This role is an opportunity to go beyond traditional technical leadership and help shape how engineering teams build software with AI and agentic workflows.

You will have the opportunity to influence both the technology we build and the way we build it.

 

By proceeding with your application, you are adhering to our PDPA policies. In case you are interested to know more, read about our Candidates Personal Data Privacy Statement. 

Skills Required

  • 5+ years operating production systems at scale in SRE, DevOps, or infrastructure engineering
  • Production Kubernetes deployment, upgrades, troubleshooting, and maintenance experience
  • Strong Linux fundamentals and performance tuning experience
  • Experience with Docker and container networking
  • Hands-on Google Cloud Platform experience
  • Infrastructure as code and declarative provisioning experience using Terraform or equivalent
  • Secrets and credential management experience with HashiCorp Vault or equivalent
  • CI/CD pipeline experience with GitHub Actions or equivalent, including progressive delivery and automated rollback
  • Bash proficiency and working proficiency in Go, Python, or a similar general-purpose language
  • Monitoring and observability experience with Prometheus, Grafana, and Google Cloud Operations
  • Ability to work independently on large, complex projects with minimal guidance
  • Practical experience applying AI and LLM-based tooling to engineering or operational workflows
  • Familiarity with MCP and agent harnesses, or ability to learn them quickly
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Kuala Lumpur
234 Employees

What We Do

Mudah.my is Malaysia’s leading marketplace that offers a free and convenient platform for people to buy and sell new and preloved goods just with a simple post of an ad. Today, more than 8 million unique visitors visit Mudah to sell and buy everything from Cars to Cameras, Properties to Pets, Mobile phones to Motorcycles, Treadmills to Textbooks, Bicycles to Beds, Guitars to Golf sets, Plants to Posters, Watches to Washing machines, Tyres to Tablets, Dresses to Drums, Shoes to Shops, Collectibles to Computers, Jobs and more – Everything Also Mudah!

Similar Jobs

Guidewire Software Logo Guidewire Software

Senior Site Reliability Engineer

Cloud • Information Technology • Insurance • Software • Analytics
In-Office
Kuala Lumpur, Wilayah Persekutuan Kuala Lumpur, MYS
3400 Employees

Hewlett Packard Enterprise Logo Hewlett Packard Enterprise

Tech Arch - Storage Channel Presales

Artificial Intelligence • Cloud • Information Technology • Consulting
Hybrid
Kuala Lumpur, Wilayah Persekutuan Kuala Lumpur, MYS
85422 Employees

Pfizer Logo Pfizer

Brand Manager

Artificial Intelligence • Healthtech • Machine Learning • Natural Language Processing • Biotech • Pharmaceutical
In-Office
Kuala Lumpur, Wilayah Persekutuan Kuala Lumpur, MYS
121990 Employees
In-Office or Remote
2 Locations
121228 Employees

Similar Companies Hiring

Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Revel Thumbnail
Aerospace • Hardware • Robotics • Software
Marina Del Rey, California
60 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software
New York, New York
30 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account