We are looking for an experienced Platform Engineer to design, build, and implement a centralized observability platform across a large enterprise technology environment.
The role will focus on improving visibility across critical applications, APIs, integration services, and infrastructure using Grafana, Prometheus, Grafana Loki, and OpenTelemetry.
The successful candidate will work across Microsoft Azure and OpenShift/Kubernetes environments to establish centralized logging, metrics, tracing, dashboards, and automated alerting.
Key responsibilities include:
- Design, deploy, and maintain a centralized observability platform using Grafana, Prometheus, and Grafana Loki.
- Implement and manage OpenTelemetry collectors for logs, metrics, and distributed tracing.
- Build telemetry pipelines for applications, APIs, infrastructure, integration services, and security-related logs.
- Configure operational and management dashboards to monitor system health, performance, availability, and incidents.
- Set up automated alerts for application errors, API latency, service timeouts, infrastructure issues, and abnormal system behaviour.
- Define and maintain standards for structured logging, error codes, trace IDs, correlation IDs, and application telemetry.
- Ensure internally developed and third-party applications comply with agreed observability and logging standards.
- Develop troubleshooting guides and operational runbooks for support teams.
- Support L1 and L2 teams in using dashboards, logs, alerts, and traces to diagnose incidents.
- Collaborate with software engineering, infrastructure, security, QA, and operations teams.
- Integrate synthetic monitoring and application health checks into the wider monitoring platform.
- Help improve Mean Time to Detect (MTTD) and Mean Time to Resolve (MTTR) across critical technology services.
Requirements
- 3 - 5+ years of experience in Platform Engineering, DevOps, Site Reliability Engineering, Cloud Engineering, or a similar role.
- Strong hands-on production experience with Grafana, Prometheus, and Grafana Loki.
- Strong experience with OpenTelemetry, including collectors, instrumentation, distributed tracing, and trace propagation.
- Experience designing and operating centralized logging, monitoring, metrics, and alerting platforms.
- Strong experience with Kubernetes and/or OpenShift.
- Hands-on experience with Microsoft Azure infrastructure and services.
- Experience working in hybrid cloud and on-premise environments.
- Experience with Infrastructure as Code and automation tools such as Terraform, Ansible, GitHub Actions, Azure DevOps, or similar.
- Good understanding of application and infrastructure logging, log parsing, and structured logging.
- Experience monitoring APIs, microservices, and enterprise applications.
- Ability to troubleshoot complex application and infrastructure issues using logs, metrics, and traces.
- Experience defining technical standards and ensuring engineering teams follow them.
- Strong communication skills and the ability to work with engineering, infrastructure, security, and support teams.
Experience with Java/Spring Boot, Node.js, enterprise integration platforms, WAF/security logging, or regulated enterprise environments would be an added advantage.
Skills Required
- 3-5+ years of experience in Platform Engineering, DevOps, Site Reliability Engineering, Cloud Engineering, or a similar role
- Production experience with Grafana, Prometheus, and Grafana Loki
- Experience with OpenTelemetry collectors, instrumentation, distributed tracing, and trace propagation
- Experience designing and operating centralized logging, monitoring, metrics, and alerting platforms
- Experience with Kubernetes and/or OpenShift
- Hands-on experience with Microsoft Azure infrastructure and services
- Experience in hybrid cloud and on-premises environments
- Experience with Infrastructure as Code and automation tools such as Terraform, Ansible, GitHub Actions, or Azure DevOps
- Understanding of application and infrastructure logging, log parsing, and structured logging
- Experience monitoring APIs, microservices, and enterprise applications
- Ability to troubleshoot complex application and infrastructure issues using logs, metrics, and traces
- Experience defining technical standards and ensuring engineering teams follow them
- Strong communication and cross-functional collaboration skills
- Experience with Java/Spring Boot, Node.js, enterprise integration platforms, WAF/security logging, or regulated enterprise environments
What We Do
FinSense Africa is a Nairobi-based financial technology company that specializes in digital transformation and open banking solutions. The firm focuses on accelerating innovation within the financial services industry across Africa by providing API integration, modernizing core systems, and offering experienced tech consultants to help banks and financial institutions overcome talent shortages and scale their digital capabilities.









