System Reliability Engineer, Consultant

Reposted 2 Months Ago
Be an Early Applicant
Kuala Lumpur, Wilayah Persekutuan Kuala Lumpur, MYS
In-Office
60K-85K Annually
Mid level
Insurance • Financial Services
The Role
The System Reliability Engineer will ensure system reliability, manage incidents, automate tasks, optimize performance, and enhance security compliance while collaborating with various teams and maintaining documentation.
Summary Generated by Built In

At AIA we’ve started an exciting movement to create a healthier, more sustainable future for everyone.

As pioneering innovators for over 100 years, we’re now transforming our organisation to be faster, simpler and more connected. Because we want to be even better equipped to develop digital solutions and experiences that help more people live Healthier, Longer, Better Lives.

To get there, we need people with tech/digital/analytics expertise and passion to help develop positive, sustainable change through digitally enhanced experiences that will impact the lives of millions of people and create a healthier future for everyone.

If you believe in developing a better tomorrow, read on. 

About the Role

We are looking for a System / Site Reliability Engineer (SRE) to help ensure the reliability, scalability, and performance of our enterprise systems and services. In this role, you will apply software engineering principles to operations, partner closely with development and infrastructure teams, and build automation that strengthens system stability and efficiency. You will play a pivotal role in bridging the gap between software development and IT operations, driving a culture of resilience, observability, automation, and proactive problem‑solving.

Key Responsibilities

1. Ensure System Reliability & Availability

  • Monitor and report on application performance, and highlight any deviations or issues.

  • Collaborate with application engineers and developers to identify root causes and implement durable fixes.

2. Incident Management & Root Cause Analysis

  • Participate as a Subject Matter Advisor during production incidents and outages.

  • Provide insights backed by system monitoring, code review, and database analysis.

  • Support post‑mortem reviews and drive follow‑up actions.

3. Automation & Tooling

  • Automate operational tasks such as monitoring, alerts, and recovery processes.

  • Build scripts and internal tools to eliminate manual toil and improve operational efficiency.

4. Monitoring & Observability

  • Implement telemetry and observability practices to track system health, latency, and error rates.

  • Manage the Dynatrace platform and its integrations with application services.

  • Support teams in designing dashboards and visualization setups.

5. Security & Compliance

  • Work with Security teams to ensure systems comply with regulatory and industry standards (e.g., PCI‑DSS, GDPR).

  • Implement necessary access controls, encryption, and audit capabilities within SRE scope.

6. Capacity Planning & Performance Optimization

  • Analyze usage trends to forecast demand and support scaling decisions.

  • Contribute to cost‑performance optimization efforts across infrastructure and applications.

  • Collaborate closely with development, QA, and infrastructure teams to embed reliability into the SDLC.

7. Documentation & Knowledge Sharing

  • Maintain clear and up‑to‑date operational documentation, runbooks, and architecture diagrams.

  • Champion SRE principles across the organization to foster resilience and accountability.

Job Requirements

Education

  • Bachelor’s degree in Computer Science, Software Engineering, IT, or related fields.

Experience

  • 3–5 years of experience in SRE, DevOps, or Software Engineering roles.

  • Experience supporting front‑end applications in production environments, ideally within financial services or other regulated industries.

Technical Skills

  • Strong understanding of front‑end performance monitoring and instrumentation.

  • Hands‑on experience with Real User Monitoring (RUM), Synthetic Monitoring, and APM tools (e.g., Dynatrace, New Relic, Datadog).

  • Proficiency in building dashboards and alerts using Dynatrace, Grafana, Prometheus, Elastic Stack, or Splunk.

  • Familiarity with OpenTelemetry for distributed tracing.

  • Scripting skills in Python, Bash, or JavaScript.

  • Experience with CI/CD pipelines (e.g., GitHub Flow).

  • Practical experience with cloud technologies (AWS or Azure).

  • Knowledge of Docker and Kubernetes.

  • Understanding of secure coding practices for front‑end applications.

  • Awareness of financial compliance standards such as PCI‑DSS.

Why Join Us?
  • Be part of a high‑impact team shaping system resilience across the enterprise.

  • Work with modern observability and automation technologies.

  • Influence engineering culture through SRE best practices.

  • Opportunities to innovate and drive real improvements in system reliability.

Build a career with us as we help our customers and the community live Healthier, Longer, Better Lives.

You must provide all requested information, including Personal Data, to be considered for this career opportunity. Failure to provide such information may influence the processing and outcome of your application. You are responsible for ensuring that the information you submit is accurate and up-to-date.

Skills Required

  • Bachelor's degree in Computer Science, Software Engineering, IT, or related fields
  • 3-5 years of experience in SRE, DevOps, or Software Engineering roles
  • Strong understanding of front-end performance monitoring and instrumentation
  • Hands-on experience with monitoring and APM tools
  • Proficiency in building dashboards and alerts
  • Scripting skills in Python, Bash, or JavaScript
  • Experience with CI/CD pipelines
  • Practical experience with cloud technologies
  • Knowledge of Docker and Kubernetes
  • Understanding of secure coding practices
  • Awareness of financial compliance standards
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Hong Kong
25,938 Employees
Year Founded: 1919

What We Do

AIA Group Limited is a multinational insurance and financial services corporation headquartered in Hong Kong, providing life insurance, savings, and health protection products across the Asia-Pacific region.

Similar Jobs

Circle (circle.so) Logo Circle (circle.so)

Lead Engineer, AI Platform

Artificial Intelligence • Consumer Web • Digital Media • Information Technology • Social Impact • Software
In-Office or Remote
18 Locations
250 Employees

Circle (circle.so) Logo Circle (circle.so)

Senior Quality Engineer

Artificial Intelligence • Consumer Web • Digital Media • Information Technology • Social Impact • Software
In-Office or Remote
18 Locations
250 Employees
120K-130K Annually

Circle (circle.so) Logo Circle (circle.so)

Designer

Artificial Intelligence • Consumer Web • Digital Media • Information Technology • Social Impact • Software
In-Office or Remote
18 Locations
250 Employees
100K-120K Annually

Circle (circle.so) Logo Circle (circle.so)

Lead Engineer, AI Platform

Artificial Intelligence • Consumer Web • Digital Media • Information Technology • Social Impact • Software
In-Office or Remote
43 Locations
250 Employees

Similar Companies Hiring

Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Vega Thumbnail
Artificial Intelligence • Automotive • Insurance • Transportation
US
65 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account