Senior Manager, Cloud Solution Architect (Observability)

Posted 16 Days Ago
Be an Early Applicant
2 Locations
In-Office or Remote
Senior level
Insurance • Financial Services
The Role
Lead design and implementation of Azure-based solution architectures and governance. Implement Azure Well-Architected practices, API management, PaaS/container platforms, and IaC (Terraform/Ansible/CloudFormation). Define application design standards, monitoring/logging, and policy-based governance. Work with security, evaluate data tools (MSSQL, PostgreSQL, Cosmos DB, Databricks, Redis, Azure Synapse), and drive DevOps practices to meet business needs.
Summary Generated by Built In
Are you ready to shape a better tomorrow?

AIA Digital+ is a Technology, Digital and Analytics innovation hub dedicated to powering AIA to be more efficient, connected and innovative as it fulfils its Purpose to help millions of people across Asia-Pacific live Healthier, Longer, Better Lives.

If you are hungry and driven to play an active role in shaping a better tomorrow, we want to hear from you. Because the work we do at AIA Digital+ makes a difference in the lives of millions of people, every day. We will equip you with the critical skills, tools and technology, and endless opportunities to learn, contribute and thrive in a dynamic and exciting environment.

If you want to shape a brighter future at AIA Digital+, please read on.

About the Role

The Observability Architect is responsible for defining, designing, implementing, and governing enterprise observability capabilities across AIA multi-cloud, hybrid, and containerized environments. The role will establish consistent monitoring, logging, tracing, event correlation, service health, and operational intelligence practices across Azure, Alibaba Cloud, and related enterprise platforms.
Dynatrace will be the primary observability platform, integrated with ServiceNow ITOM to support event management, incident enrichment, service mapping, root cause analysis, and operational automation. The role will also define complementary patterns for logging and open-source monitoring platforms such as Elastic, OpenSearch, Prometheus, Grafana, and OpenTelemetry.
A key objective of this role is to drive AIOps-enabled operations and self-healing automation to reduce alert noise, improve Mean Time To Detect (MTTD), reduce Mean Time To Resolve (MTTR), and improve overall platform and application reliability.

Roles and Responsibilities:

Observability Architecture & Design
•    Define and maintain enterprise observability reference architecture, standards, patterns, and governance across AIA Group and Business Units.
•    Design end-to-end observability for infrastructure monitoring, application performance monitoring, distributed tracing, logging, digital experience monitoring, network observability, and service health dashboards.
•    Establish observability KPIs, SLOs, SLIs, error budgets, alerting principles, and service reliability reporting standards.
•    Review solution designs and ensure observability requirements are embedded from architecture and delivery stages.

Dynatrace Platform Architecture
•    Lead architecture and governance of Dynatrace as the enterprise observability platform.
•    Design monitoring standards for applications, Kubernetes, virtual machines, databases, middleware, APIs, network services, and cloud-native services.
•    Define Dynatrace tagging standards, management zones, dashboard patterns, service mapping, synthetic monitoring, real user monitoring, and Davis AI adoption.
•    Drive platform configuration, onboarding patterns, operational dashboards, reporting, and observability data retention standards.

ServiceNow ITOM Integration
•    Architect integration between Dynatrace and ServiceNow ITOM Event Management / ITOM Health capabilities.
•    Design event correlation, alert enrichment, topology-aware service impact analysis, and automated incident creation workflows.
•    Define noise reduction, deduplication, priority mapping, escalation, and operational ownership models.
•    Support integration of observability insights into ITSM processes, command center dashboards, and operational reporting.

AIOps & Auto-Healing Automation
•    Define and implement AIOps operating patterns for anomaly detection, predictive alerting, root cause identification, and closed-loop remediation.
•    Design auto-healing workflows integrated with Dynatrace, ServiceNow ITOM, automation platforms, and cloud-native services.
•    Identify repeatable operational failure patterns and convert them into automated remediation runbooks where appropriate.
•    Drive continuous improvement to reduce manual intervention, alert fatigue, incident recurrence, MTTD, and MTTR.

Logging & Observability Data Platforms
•    Architect centralized logging and log analytics solutions using Elastic, OpenSearch, cloud-native log services, and enterprise logging standards.
•    Define log collection, normalization, enrichment, indexing, retention, access control, and cost governance standards.
•    Establish patterns for correlation across logs, metrics, traces, events, configuration, and service topology.
•    Ensure logging platforms support operational troubleshooting, security visibility, auditability, and compliance requirements.

Open Source Monitoring & Cloud Native Observability
•    Design monitoring and visualization solutions using Prometheus, Grafana, OpenTelemetry, and related open-source monitoring platforms.
•    Define observability patterns for AKS, Alibaba Cloud ACK, containers, microservices, APIs, service mesh, and DevOps pipelines.
•    Establish metrics federation, dashboarding, alerting, and integration approaches with Dynatrace and enterprise platforms.
•    Promote consistent instrumentation and telemetry standards across modern application architectures.

Azure & Alibaba Cloud Native Monitoring
•    Define monitoring standards for Azure platform services including Azure Monitor, Log Analytics, Application Insights, Network Watcher, Azure Managed Prometheus, Azure Managed Grafana, Azure Advisor, and Azure Resource Health.
•    Define monitoring standards for Alibaba Cloud services including CloudMonitor, Log Service (SLS), ActionTrail, ARMS, Managed Service for Prometheus, Security Center, and related observability and alarm services.
•    Ensure cloud-native monitoring capabilities are integrated with enterprise observability, ITOM, incident, and reporting processes.
•    Drive cost-aware telemetry design across Azure and Alibaba Cloud environments.

Security, Governance & Compliance
•    Embed security-by-design, least-privilege access, data protection, and compliance requirements into observability architecture.
•    Support audits, risk assessments, regulatory reviews, and evidence requirements related to monitoring, logging, and service reliability.
•    Define guardrails for observability platform access, retention, data classification, dashboard sharing, and operational reporting.
•    Maintain observability documentation, runbooks, standards, design patterns, and operational controls.

Stakeholder & Technical Leadership
•    Act as the enterprise subject matter expert for observability, monitoring, logging, AIOps, and self-healing automation.
•    Collaborate with Cloud Architecture, Cloud Engineering, Operations, Security, DevOps, Application, Service Management, and Business Unit teams.
•    Mentor engineering and operations teams on observability best practices, platform onboarding, dashboarding, and incident reduction techniques.
•    Drive observability transformation initiatives and promote SRE-aligned operational practices across the organization.
 

Minimum Job Requirements:

Skills:
•    10+ years relevant experience in Cloud Architecture, Infrastructure, Operations, Observability, or Application Performance Management.
•    5+ years practical experience designing and governing enterprise observability platforms.
•    Strong hands-on experience with Dynatrace, including APM, infrastructure monitoring, Kubernetes monitoring, dashboards, management zones, service mapping, and Davis AI capabilities.
•    Experience integrating observability platforms with ServiceNow ITOM / Event Management / ITSM processes.
•    Strong experience with logging platforms such as Elastic Stack, OpenSearch, or equivalent enterprise log analytics solutions.
•    Experience with open-source monitoring and visualization platforms such as Prometheus, Grafana, and OpenTelemetry.
•    Hands-on experience with Azure native monitoring services including Azure Monitor, Log Analytics, Application Insights, Network Watcher, Azure Managed Prometheus, and Azure Managed Grafana.
•    Hands-on experience with Alibaba Cloud native monitoring services including CloudMonitor, Log Service (SLS), ActionTrail, ARMS, Managed Service for Prometheus, and related observability services.
•    Strong understanding of Kubernetes, containers, microservices, APIs, network monitoring, distributed tracing, cloud networking, IAM, and security principles.
•    Experience implementing AIOps, event correlation, anomaly detection, automated remediation, and self-healing operations.
•    Experience working in enterprise or regulated environments is highly desirable.
•    Relevant professional certifications such as Dynatrace Associate/Professional, Microsoft Azure Solutions Architect Expert, Alibaba Cloud Professional Architect, CKA, ITIL, SRE Foundation, or TOGAF will be an advantage.
•    Sound understanding of IT partner ecosystem and partner collaboration in a multinational corporation.
•    Experience in top-tier multinational corporation will be an advantage.
•    Strong problem-solving and analytical skills.
•    Excellent communication, stakeholder management, and collaboration skills.
•    Ability to translate complex operational telemetry into actionable service reliability insights.
•    Ability to operationalize disruptive technology services, including building implementation roadmaps for observability, AIOps, and self-healing automation.
 

Build a career with us as we help our customers and the community live healthier, longer, better lives.

You must provide all requested information, including Personal Data, to be considered for this career opportunity. Failure to provide such information may influence the processing and outcome of your application. You are responsible for ensuring that the information you submit is accurate and up-to-date.

Skills Required

  • Degree in IT, Computer Science, or equivalent discipline
  • At least 12 years IT work experience with at least 5 years in cloud architecture and design
  • Solution architecture experience on Azure Cloud and SaaS services
  • Experience on API management
  • Design and implement Azure Well-Architected framework
  • Design technology platforms on Azure using PaaS services and container platforms
  • Implement and maintain Infrastructure as Code (IaC) using Terraform, Ansible, or CloudFormation
  • Design and implement policy-based governance using Azure Policies
  • Define monitoring and logging for application software
  • Knowledge of data products: MSSQL, PostgreSQL, Cosmos DB, Databricks, Redis, Azure Synapse
  • Knowledge of DevOps practices
  • Collaborate with security teams to ensure compliance and implement security measures
  • Ability to work under pressure
  • Strong sense of ownership and self-driven
  • Excellent communication skills in English
  • Cooperative, good teamwork and able to work independently
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Hong Kong
25,938 Employees
Year Founded: 1919

What We Do

AIA Group Limited is a multinational insurance and financial services corporation headquartered in Hong Kong, providing life insurance, savings, and health protection products across the Asia-Pacific region.

Similar Jobs

Micron Technology Logo Micron Technology

Program Manager

Artificial Intelligence • Hardware • Information Technology • Machine Learning
Remote
Batu Kawan, Selatan, Pulau Pinang, MYS
45000 Employees

Micron Technology Logo Micron Technology

Development Engineer

Artificial Intelligence • Hardware • Information Technology • Machine Learning
Remote
Batu Kawan, Selatan, Pulau Pinang, MYS
45000 Employees

Micron Technology Logo Micron Technology

General Deposit - Engineer, Supervisor & Other Professionals

Artificial Intelligence • Hardware • Information Technology • Machine Learning
Remote
Batu Kawan, Selatan, Pulau Pinang, MYS
45000 Employees

Micron Technology Logo Micron Technology

Principal, Global Leadership & Organization Capability

Artificial Intelligence • Hardware • Information Technology • Machine Learning
Remote
Batu Kawan, Selatan, Pulau Pinang, MYS
45000 Employees

Similar Companies Hiring

Granted Thumbnail
Artificial Intelligence • Healthtech • Insurance • Mobile • Financial Services
New York, New York
23 Employees
Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account