#127606 - Senior Software Engineer / SRE (Observability Focus)

Posted One Month Ago
Be an Early Applicant
Hiring Remotely in Singapore, SGP
In-Office or Remote
Senior level
Enterprise Web • HR Tech • Professional Services • Software
The Role
Build and operate observability capabilities for containerized microservices and Kubernetes platforms. Design and implement software, APIs, integrations, dashboards, alerts, tracing, metrics, and automation (scripting preferred in Python). Integrate observability into AWS and CI/CD pipelines, install/configure agents, manage access and API keys, and drive proactive reliability, scalability, and performance improvements.
Summary Generated by Built In
Job Description

We are seeking a Senior Software Engineer / SRE with a strong observability focus to support platform reliability, monitoring, and modernization initiatives. This role combines approximately 60–70% software engineering with 30–40% site reliability engineering and requires hands-on experience working with Kubernetes, cloud infrastructure, observability platforms, APIs, and operational automation.

 

The successful candidate will help build and operate observability capabilities across containerized and microservices-based environments while proactively improving platform reliability, scalability, and performance.

 

Enterprise experience strongly preferred.

 

Key Responsibilities:

 

- Design, build, and maintain software, APIs, integrations, and automation that support platform reliability and observability.

- Support monitoring, reliability, maintenance, and continuous improvement across internal platforms and systems.

- Work in Kubernetes-based environments across deployment, operations, and monitoring activities.

- Build and maintain observability solutions with a focus on Datadog.

- Configure dashboards, alerts, application performance monitoring, tracing, metrics, and logging.

- Monitor containerized and microservices-based applications.

- Integrate observability platforms and monitoring capabilities into AWS environments.

- Integrate observability capabilities into CI/CD pipelines and deployment processes.

- Automate monitoring and operational tasks through scripting, with Python preferred.

- Install and configure Datadog agents and integrations.

- Manage observability API keys and secure configurations.

- Manage user roles, permissions, and access controls within observability platforms.

- Lead proactive maintenance efforts and platform improvements.

- Drive improvements in reliability, scalability, performance, and operational efficiency.

Qualifications

Must-Have Skills:

 

- Strong proficiency in at least one of Python, JavaScript using Node.js, or Java.

- Hands-on experience designing, consuming, and implementing API integrations.

- Strong Kubernetes experience covering deployment, operations, and monitoring.

- Hands-on experience with Datadog or a comparable observability platform such as Prometheus or Grafana.

- Experience configuring dashboards, alerts, application performance monitoring, tracing, metrics, and logging.

- Experience monitoring containerized and microservices-based architectures.

- Hands-on AWS experience.

- Experience integrating observability tooling into cloud environments.

- Experience integrating observability capabilities into CI/CD pipelines.

- Ability to automate monitoring and operational work through scripting.

 

Nice-to-Have Skills:

 

- Strongly preferred experience owning and operating an internal engineering platform.

- Strongly preferred experience owning reliability, scalability, and performance outcomes.

- Strongly preferred experience proactively leading maintenance efforts and platform improvements rather than providing only reactive support.

- Familiarity with Go or Golang.

- Experience with New Relic, Dynatrace, Elastic, or Splunk Observability.

- Experience working across multiple observability and monitoring platforms.

Additional Information

Required Tools & Platforms:

 

- Python, JavaScript using Node.js, or Java

- Kubernetes

- Datadog, Prometheus, Grafana, or a comparable observability platform

- AWS

- CI/CD pipelines

- APIs and integration tooling

- Application performance monitoring, tracing, metrics, logging, dashboards, and alerting tools

 

Location, Time & Engagement:

 

- Remote contract opportunity.

- Candidates must be located in APAC.

- Ability to provide overlap with Japan Standard Time is preferred.

- Full-time allocation of 40 hours per week.

- Expected contract end date is March 31, 2027.

Skills Required

  • Proficiency in Python
  • Proficiency in JavaScript / Node.js
  • Proficiency in Java
  • Hands-on experience designing, consuming, and implementing API integrations
  • Strong Kubernetes experience (deployment, operations, monitoring)
  • Hands-on experience with Datadog or comparable observability platform (Prometheus, Grafana)
  • Experience configuring dashboards, alerts, APM, tracing, metrics, and logging
  • Experience monitoring containerized and microservices-based architectures
  • Hands-on AWS experience
  • Experience integrating observability tooling into cloud environments and CI/CD pipelines
  • Ability to automate monitoring and operational tasks through scripting (Python preferred)
  • Install and configure Datadog agents and integrations
  • Manage observability API keys, secure configurations, user roles, and access controls
  • Located in APAC
  • Overlap with Japan Standard Time
  • Enterprise platform experience
  • Familiarity with Go / Golang
  • Experience with New Relic, Dynatrace, Elastic, or Splunk Observability
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
283 Employees
Year Founded: 2025

What We Do

Lifted, an Upwork Company, is a B2B SaaS platform that helps enterprise organizations source, contract, manage, and pay contingent talent globally and compliantly. It supports multiple engagement models, including independent contractors, staff augmentation, and employer-of-record, while integrating with MSPs and VMSs to provide centralized visibility, spend control, and audit-ready compliance, delivering a white-labeled talent experience for hiring managers and contingent workforce programs.

Similar Jobs

Atlassian Logo Atlassian

Machine Learning Engineer

Cloud • Information Technology • Productivity • Security • Software • App development • Automation
In-Office or Remote
Singapore, SGP
11000 Employees

Coinbase Logo Coinbase

Software Engineer

Artificial Intelligence • Blockchain • Fintech • Financial Services • Cryptocurrency • NFT • Web3
Easy Apply
In-Office or Remote
Singapore, SGP
4700 Employees
144K-144K Annually

Cloudflare Logo Cloudflare

Customer Engineer, Digital Native Business, Indonesia

Cloud • Information Technology • Security • Software • Cybersecurity
Remote or Hybrid
Singapore, SGP
4400 Employees

Cloudflare Logo Cloudflare

Customer Engineer, Digital Native Bsuiness, Singapore

Cloud • Information Technology • Security • Software • Cybersecurity
Remote or Hybrid
Singapore, SGP
4400 Employees

Similar Companies Hiring

Kepler  Thumbnail
Artificial Intelligence • Fintech • Software
New York, New York
9 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Revel.io Thumbnail
Aerospace • Hardware • Robotics • Software
US
50 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account