Data Engineer (Observability Engineer) - A26323

Posted 7 Days Ago
Be an Early Applicant
Singapore, SGP
In-Office
Mid level
Information Technology
The Role
Designs and operates observability across on-premises, cloud, hybrid infrastructure, applications, and networks. The role establishes OpenTelemetry instrumentation standards, telemetry collection, dashboards, SLIs, SLOs, alerting, and operational health indicators. Responsibilities include integrating telemetry sources, supporting incident investigation, improving service reliability, managing retention and security requirements, automating deployments with Infrastructure as Code, and maintaining runbooks. The engineer collaborates with application, infrastructure, network, security, platform, and operations teams and participates in on-call support.
Summary Generated by Built In

Activate Interactive Pte Ltd (“Activate”) is a leading technology consultancy headquartered in Singapore with a presence in Malaysia and Indonesia. Our clients are empowered with quality, cost-effective, and impactful end-to-end application development, like mobile and web applications, and cloud technology that remove technology roadblocks and increase their business efficiency.

We believe in positively impacting the lives of people around us and the environment we live in through the use of technology. Hence, we are committed to providing a conducive environment for all employees to realise their full potential, who in turn have the opportunity to continuously drive innovation.

We are searching for our next team members to join our growing team.

If you love the idea of being part of a growing company with exciting prospects in mobile and web technologies that create positive impact on people’s lives, then we would love to hear from you.

Co-Development Business Unit is looking for Data Engineer (Observability Engineer)

This is a 1 - year contract role.

Internal Code: A26323

Digital Excellence & Products Division (DXD) is a GovTech team within the Ministry of Education (MOE). DXD sits at the intersection of technology, design, and education, building meaningful products, platforms, and digital services that improve teaching, learning, school operations, and the experience of students, teachers, and school leaders.

What will you do?

We are looking for an Observability Engineer to help build and operate the observability capabilities of the future SSOE platform.

You will help provide end-to-end visibility across MOE's technology environment, spanning on-premises infrastructure, networks, applications, cloud platforms, and hybrid environments. You will enable engineering and operations teams to understand system health, identify issues early, diagnose incidents quickly, and continuously improve service reliability.

As an Observability Engineer, you will establish and operate consistent observability capabilities across SSOE infrastructure and applications.

You will work across metrics, events, logs, and traces to provide a unified view of service health and performance. You will define instrumentation standards, service-level indicators and objectives, alerting strategies, dashboards, and operational health signals.

You will work closely with application, infrastructure, network, security, and platform teams to ensure observability is built into services from the outset rather than added after deployment.

Observability Engineering

  • Design and operate end-to-end observability across on-premise infrastructure, networks, applications, cloud platforms, and hybrid environments
  • Collect, aggregate, and correlate metrics, events, logs, and traces across infrastructure and application workloads
  • Define and maintain observability standards that work consistently across legacy, on-premise, containerised, and cloud-native systems
  • Establish application and infrastructure instrumentation standards using OpenTelemetry and other appropriate technologies
  • Support engineering teams with instrumentation, SDK, agent, and telemetry integration
  • Define common conventions for service naming, metadata, tagging, correlation IDs, and telemetry enrichment
  • Identify observability gaps and continuously improve end-to-end visibility across SSOE services

Service Reliability & Monitoring

  • Define SLIs, SLOs, alerting rules, and service health indicators for critical services
  • Build operational dashboards covering infrastructure health, application performance, user experience, availability, capacity, and service reliability
  • Develop leadership-level views that provide meaningful visibility into service performance and operational trends
  • Design actionable alerting that enables teams to identify and respond to issues while minimising unnecessary alert noise
  • Establish monitoring and operational-readiness requirements for new applications, infrastructure, and platform components
  • Use observability data to support capacity planning, performance analysis, reliability improvements, and operational decision-making

Telemetry & Integration

  • Define secure telemetry collection and routing across on-premise environments, GCC, cloud platforms, and approved SaaS services
  • Work with infrastructure and platform teams to integrate telemetry from servers, network devices, applications, containers, databases, and managed cloud services
  • Design observability approaches that account for network boundaries, security zones, data residency, and connectivity constraints
  • Define telemetry retention, lifecycle, and cost-management requirements
  • Ensure logs, metrics, and traces can be correlated across distributed and hybrid systems

Incident Management & Continuous Improvement

  • Support operational teams during incidents by using observability data to identify symptoms, dependencies, and potential root causes
  • Participate in incident investigation, root-cause analysis, and post-incident reviews
  • Identify recurring operational issues and recommend improvements to instrumentation, alerting, architecture, or operational processes
  • Define appropriate SLOs and operational health indicators for observability services
  • Participate in operational support and on-call responsibilities for owned services
  • Maintain architecture documentation, operational procedures, and runbooks

Requirements

What are we looking for?

  • Minimum 3–5 years of experience in observability engineering, Site Reliability Engineering (SRE), platform engineering, infrastructure engineering, or a related discipline
  • At least 2 years of hands-on experience implementing or operating observability and monitoring capabilities in production environments
  • Demonstrated experience working with metrics, logging, tracing, dashboards, alerting, and incident troubleshooting
  • Experience monitoring and supporting production infrastructure, applications, or distributed systems
  • Experience working with on-premise and/or cloud environments, with an understanding of hybrid infrastructure
  • Experience working with engineering or operations teams to implement instrumentation and improve service reliability
  • Treats observability configuration, instrumentation, dashboards, and platform components as version-controlled engineering artefacts
  • Uses automation and Infrastructure as Code for repeatable and auditable deployments
  • Designs observability for reliability, scalability, security, and operational sustainability
  • Understands the difference between collecting telemetry and creating useful operational signals
  • Designs monitoring and alerting around service and user impact rather than individual infrastructure metrics alone
  • Builds observability into services from the beginning of the engineering lifecycle
  • Uses code review, testing, and CI/CD for observability-related changes where appropriate
  • Collaborates effectively across application, infrastructure, network, security, data, and platform teams
  • Apply MOE and Government data-classification requirements
  • Ensure telemetry is collected, transmitted, stored, and accessed according to applicable security and data-residency requirements
  • Prevent sensitive information, credentials, and secrets from being unnecessarily captured in telemetry
  • Implement appropriate access controls for observability platforms and operational data
  • Maintain auditability and traceability of observability configuration and operational activities
  • Participate in security, architecture, and operational-readiness reviews
  • Experience with Singapore Government platforms such as TechPass, SHIP-HATS, SEED, and GCC
  • Familiarity with OC/SN data-classification requirements
  • AWS or Azure cloud certifications
  • Experience implementing OpenTelemetry at scale
  • Experience operating observability platforms across large or distributed environments
  • Experience monitoring hybrid infrastructure spanning data centres and cloud environments
  • Familiarity with SRE practices such as error budgets, SLO management, incident response, and reliability engineering
  • Dynatrace, Elastic, Grafana, Prometheus or equivalent
  • OpenTelemetry
  • AWS-native observability and monitoring capabilities
  • Docker, ECS, CI/CD, SHIP-HATS, Terraform / OpenTofu, Ansible
  • Cloud-native storage and telemetry data services
  • Kafka, MQ, event-driven telemetry patterns

Benefits

What do we offer in return?

Competitive Compensation: Market competitive salary and variable performance bonus aligned to your skills, impact, and contribution. 

Benefits: Outpatient medical, specialist medical coverage and generous customisable flexi benefits, or flexi allowance; life and health insurance, thoughtful perks like special occasions “red packet” and CNY goodies, etc. 

Employee Wellness: Support for your physical, mental, and overall well-being through year-round initiatives. 

Growth & Development: Learning programmes, certification support, and a dedicated staff development budget. (We are a “SHRI 2025 Gold winner” in “Learning & Development; Coaching & Mentoring”) 

Career Progression: Structured career pathways that enable you to grow along a technical/domain expert track or a leadership track. 

Competency Framework: A structured and practical framework to support you to develop, perform, and succeed. 

Flexible Work Arrangement: Staff may choose to work flexi-place, flexi-time, and flexi-load based on existing framework 

Why you'll love working with us? 

If you are looking for opportunities to collaborate with leading industry experts and be surrounded by highly motivated and talented peers, we welcome you to join us. We provide all employees with equal opportunities to grow and develop with us. We believe your success is our success. 

Does it sound like something you are interested in exploring further? Please be in touch with our team for an initial chat.

Activate Interactive Singapore is an equal opportunity employer. Employment decisions will be based on merit, qualifications and abilities. Activate Interactive Pte Ltd does not discriminate in employment opportunities or practices on the basis of race, colour, religion, gender, sexuality, national origin, age, disability, marital status or any other characteristics protected by law. 

Protecting your privacy and the security of your data are longstanding top priorities for Activate Interactive Pte Ltd. 

Your personal data will be processed for the purposes of managing Activate Interactive Pte Ltd’s recruitment related activities, which include setting up and conducting interviews and tests for applicants, evaluating and assessing the results, and as is otherwise needed in the recruitment and hiring processes. 

Please consult our Privacy Notice (https://www.activate.sg/privacy-policy) to know more about how we collect, use, and transfer the personal data of our candidates. Here you can find how you can request for access, correction and/or withdrawal of your Personal Data. 

Skills Required

  • Minimum 3-5 years of experience in observability engineering, Site Reliability Engineering, platform engineering, infrastructure engineering, or a related discipline
  • At least 2 years of hands-on experience implementing or operating observability and monitoring capabilities in production environments
  • Experience with metrics, logging, tracing, dashboards, alerting, and incident troubleshooting
  • Experience monitoring and supporting production infrastructure, applications, or distributed systems
  • Experience with on-premises and/or cloud environments, including hybrid infrastructure
  • Experience collaborating with engineering or operations teams to implement instrumentation and improve service reliability
  • Experience using automation and Infrastructure as Code for repeatable and auditable deployments
  • Experience with security, data classification, data residency, access control, auditability, and telemetry protection requirements
  • Experience participating in security, architecture, and operational-readiness reviews
  • Experience with Singapore Government platforms such as TechPass, SHIP-HATS, SEED, and GCC
  • Familiarity with OC/SN data-classification requirements
  • AWS or Azure cloud certification
  • Experience implementing OpenTelemetry at scale
  • Experience operating observability platforms across large or distributed environments
  • Experience monitoring hybrid infrastructure spanning data centres and cloud environments
  • Familiarity with SRE practices including error budgets, SLO management, incident response, and reliability engineering
  • Experience with Dynatrace, Elastic, Grafana, Prometheus, or equivalent platforms
  • Experience with OpenTelemetry, AWS-native observability, Docker, ECS, CI/CD, SHIP-HATS, Terraform or OpenTofu, and Ansible
  • Experience with cloud-native storage and telemetry data services
  • Experience with Kafka, MQ, and event-driven telemetry patterns
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Singapore
143 Employees
Year Founded: 1997

What We Do

Founded in 1997, Activate Interactive Pte Ltd is a leading technology consultancy in Singapore that fuses strategy consulting, creativity, and engineering to drive digital innovations. ​ ​We offer quality, cost-effective and impactful end-to-end application development, including mobile and web applications, cloud technology, UI/UX design and more.​ ​We integrate digital technology into all areas of a business, helping clients remove technology roadblocks and increase their efficiency to better serve and deliver value to their communities, regardless of their business size and type.​ At Activate, we also believe in providing a conducive environment and developing our employees to realise their full potential. Today, we have a team of more than 150 employees. ​We aspire to help people live better and healthier with technology by providing holistic solutions to improve population health. ​

Similar Jobs

Atlassian Logo Atlassian

Machine Learning Engineer

Cloud • Information Technology • Productivity • Security • Software • App development • Automation
In-Office or Remote
Singapore, SGP
11000 Employees

Coinbase Logo Coinbase

Software Engineer

Artificial Intelligence • Blockchain • Fintech • Financial Services • Cryptocurrency • NFT • Web3
Easy Apply
In-Office or Remote
Singapore, SGP
4700 Employees
144K-144K Annually

Cloudflare Logo Cloudflare

Customer Engineer, Digital Native Business, Indonesia

Cloud • Information Technology • Security • Software • Cybersecurity
Remote or Hybrid
Singapore, SGP
4400 Employees

Cloudflare Logo Cloudflare

Customer Engineer, Digital Native Bsuiness, Singapore

Cloud • Information Technology • Security • Software • Cybersecurity
Remote or Hybrid
Singapore, SGP
4400 Employees

Similar Companies Hiring

Standard Template Labs Thumbnail
Artificial Intelligence • Information Technology • Software
New York, NY
25 Employees
NODA AI Thumbnail
Artificial Intelligence • Information Technology • Software • Cybersecurity
Sydney, AU
54 Employees
Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account