Senior Site Reliability Engineer

Posted Yesterday
Be an Early Applicant
Chapitala, Muradnagar, BGD
In-Office
Senior level
Aerospace • Transportation • Travel
The Role
Design, build, and operate highly available, secure cloud platforms on GCP using Terraform, GitOps (Argo CD), and GitLab CI/CD. Implement observability, automate operations, manage Cloudflare services, perform on-call incident response and RCA, drive reliability and cost/control governance, mentor engineers, and improve developer productivity through automation and best practices.
Summary Generated by Built In


Job Description

We are looking for a highly motivated Senior Site Reliability Engineer (SRE) to join our Platform Engineering team. In this role, you will design, build, and operate highly available, secure, and scalable cloud platforms while driving automation across the software delivery lifecycle.

You will partner closely with Engineering, Security, DevOps, and Product teams to improve platform reliability, developer productivity, operational excellence, and cloud governance. This is a hands-on engineering role requiring strong expertise in cloud infrastructure, Infrastructure as Code (IaC), GitOps, observability, and incident management.


Key Responsibilities
  • Design, build, and operate highly available production platforms on Google Cloud Platform (GCP).
  • Develop Infrastructure as Code (IaC) using Terraform to provision and manage cloud infrastructure.
  • Implement and maintain GitOps workflows using Argo CD and GitLab.
  • Build and enhance CI/CD pipelines using GitLab to enable secure, reliable, and automated software delivery.
  • Develop automation solutions to eliminate repetitive operational tasks using scripting and APIs.
  • Manage and optimize Cloudflare services including DNS, WAF, CDN, Load Balancing, Zero Trust, and security controls.
  • Build and maintain observability platforms including monitoring, logging, alerting, tracing, dashboards, and SLO/SLI reporting.
  • Drive platform reliability through proactive monitoring, capacity planning, performance tuning, resilience testing, and automation.
  • Participate in an on-call rotation, troubleshoot production incidents, lead incident response, perform root cause analysis (RCA), and implement permanent corrective actions.
  • Improve operational excellence by reducing toil through automation and self-service capabilities.
  • Collaborate with development teams to improve application reliability, deployment strategies, and operational readiness.
  • Ensure platform security by implementing infrastructure best practices, policy enforcement, secrets management, and least-privilege access.
  • Create and maintain technical documentation, operational runbooks, and standard operating procedures.
  • Mentor junior engineers and promote SRE best practices across engineering teams.

Required Qualifications
  • Bachelor's degree in Computer Science, Information Technology, Engineering, or equivalent practical experience.
  • 5+ years of experience in Site Reliability Engineering, Platform Engineering, Cloud Engineering, or DevOps.
  • Strong hands-on experience with Google Cloud Platform (GCP).
  • Strong experience with Terraform and Infrastructure as Code.
  • Hands-on experience with GitLab CI/CD.
  • Experience implementing GitOps using Argo CD.
  • Experience managing Cloudflare services including DNS, WAF, CDN, and Load Balancing.
  • Strong Linux administration and troubleshooting skills.
  • Experience with container technologies including Docker and Kubernetes.
  • Strong scripting skills using Bash, Python, or Go.
  • Experience with monitoring, logging, and observability platforms.
  • Experience with incident management, production support, and on-call operations.
  • Excellent troubleshooting and root cause analysis skills.
  • Strong communication and stakeholder management skills.

Preferred Qualifications
  • Experience operating Kubernetes platforms such as GKE.
  • Experience with service mesh technologies (Istio, Linkerd, or Envoy).
  • Knowledge of SRE principles including SLIs, SLOs, Error Budgets, and Toil Reduction.
  • Experience implementing platform security and DevSecOps practices.
  • Experience with FinOps and cloud cost optimization.
  • Experience with policy-as-code and infrastructure governance.
  • Google Cloud Professional certifications are an advantage.
  • Knowledge in API’s and gateways is added advantages

What Success Looks Like

Within your first 12 months, you will:

  • Improve platform reliability and availability through automation and engineering improvements.
  • Reduce operational toil by automating manual processes.
  • Improve deployment reliability using GitOps and CI/CD best practices.
  • Enhance observability with actionable monitoring and alerting.
  • Strengthen platform security and operational governance.
  • Enable engineering teams to deliver software faster and more reliably.

Skills Required

  • Bachelor's degree in Computer Science, Information Technology, Engineering, or equivalent practical experience
  • 5+ years experience in Site Reliability Engineering, Platform Engineering, Cloud Engineering, or DevOps
  • Hands-on experience with Google Cloud Platform (GCP)
  • Experience with Terraform and Infrastructure as Code
  • Hands-on experience with GitLab CI/CD
  • Experience implementing GitOps using Argo CD
  • Experience managing Cloudflare services (DNS, WAF, CDN, Load Balancing)
  • Strong Linux administration and troubleshooting skills
  • Experience with container technologies including Docker and Kubernetes
  • Strong scripting skills using Bash, Python, or Go
  • Experience with monitoring, logging, tracing, and observability platforms (SLO/SLI reporting)
  • Experience with incident management, production support, and on-call operations
  • Excellent troubleshooting and root cause analysis skills
  • Strong communication and stakeholder management skills
  • Experience developing automation solutions to eliminate repetitive operational tasks using scripting and APIs
  • Experience building and enhancing CI/CD pipelines to enable secure, reliable, automated software delivery
  • Experience operating Kubernetes platforms such as GKE
  • Experience with service mesh technologies (Istio, Linkerd, Envoy)
  • Knowledge of SRE principles including SLIs, SLOs, Error Budgets, and Toil Reduction
  • Experience implementing platform security and DevSecOps practices
  • Experience with FinOps and cloud cost optimization
  • Experience with policy-as-code and infrastructure governance
  • Google Cloud Professional certifications
  • Knowledge in APIs and gateways
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Sepang, Selangor Darul Ehsan
13,132 Employees
Year Founded: 2001

What We Do

It all starts here. 23 years ago, a dream took flight - shaping and forever changing the travel industry in Asia. The idea was simple: Make flying affordable for everyone. We made that dream happen. We started an airline in 2001. Today, we’ve evolved to become something much bigger. We’re now a world-class brand, a leading Asean airline, a digital travel and lifestyle platform; and we’re not stopping. If you’re passionate about connecting people and transforming lives, we want you onboard. When it comes to your career, your Allstar journey will be an adventure. Find your dream career destination with us

Similar Jobs

DBS Bank Ltd Logo DBS Bank Ltd

Full-stack Engineer

Fintech • Information Technology • Software • Financial Services
In-Office or Remote
17 Locations
41000 Employees

Snap! Mobile Logo Snap! Mobile

Sales Representative

Edtech • Fintech • Sports
Easy Apply
In-Office
North Tula, Shahrasti, BGD
350 Employees

Turner & Townsend Logo Turner & Townsend

Associate Director, Cost Management - Infrastructure

Professional Services • Real Estate • Consulting
In-Office or Remote
17 Locations
17263 Employees
In-Office or Remote
17 Locations
17263 Employees

Similar Companies Hiring

Toro TMS Thumbnail
Cloud • Enterprise Web • Sales • Software • Transportation
Chicago, IL
80 Employees
PRIMA Thumbnail
Travel • Software • Marketing Tech • Hospitality • eCommerce
US
15 Employees
Outpost Space Thumbnail
Aerospace • Defense
US
24 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account