Lead DevOps / Observability Engineer

Posted 3 Days Ago
Be an Early Applicant
Hiring Remotely in Poland
Remote
Entry level
Information Technology • Consulting
The Role
Lead a hands-on migration from New Relic and Datadog to Grafana using Prometheus, Loki, and Tempo. Design observability architecture, deploy infrastructure with AWS, EKS, Kubernetes, Terraform, and Docker, instrument application code, rebuild dashboards and alerts, support CI/CD automation, address security issues, and coordinate cutover across engineering teams.
Summary Generated by Built In

We are looking for a Lead DevOps / Observability Engineer to lead the migration of the client's observability stack from New Relic and Datadog to Grafana, using technologies such as Prometheus, Loki, and Tempo. 

This is a hands-on role combining:
- DevOps & platform engineering - deploying and configuring the new observability stack using AWS, Kubernetes, and Terraform.
- Application instrumentation - modifying application code to add or adapt metrics, logs, and traces.
- Migration & collaboration - rebuilding existing dashboards and alerts in Grafana and working with multiple engineering teams to ensure full monitoring coverage.

We are looking for someone who is comfortable working across both infrastructure and application code and can independently drive a cross-team observability migration.

Responsibilities:

  • Lead the migration of monitoring, alerting, and dashboards from New Relic and Datadog to Grafana.
  • Design and implement the target observability architecture (metrics, logs, traces) and the infrastructure-as-code (Terraform) needed to support it.
  • Audit existing New Relic and Datadog dashboards, alerts, and integrations to build a complete migration inventory and ensure feature/coverage parity in Grafana.
  • Make application-level code changes across services to add, adjust, or replace instrumentation (metrics exporters, logging, tracing libraries) as needed for the new stack.
  • Work with AWS infrastructure, including EKS, Kubernetes, and Docker, to deploy and operate the new observability tooling.
  • Configure AWS networking and access as needed to support the observability platform (VPCs, security groups, IAM).
  • Contribute to CI/CD pipelines and deployment automation using GitHub Actions to support the rollout.
  • Collaborate with development, platform, and infrastructure teams to coordinate the cutover from New Relic/Datadog to Grafana with minimal disruption.
  • Identify and remediate any security issues (secrets, keys, dependencies) encountered in the course of this work.
  • Document the new observability architecture, migration procedures, and dashboard/alert mappings.
  • Use AI-assisted development tools, including Codex, to improve engineering productivity and delivery speed.

Requirements: 

  • Strong hands-on experience with Grafana, including dashboard design, alerting, and data source configuration (e.g. Prometheus, Loki, Tempo, or similar).
  • Practical experience migrating away from commercial observability platforms such as New Relic and/or Datadog.
  • Full-stack development experience — comfortable reading, modifying, and instrumenting application code across front-end and back-end services, not just infrastructure.
  • Solid experience with AWS in production environments, including EKS and Kubernetes.
  • Practical experience with Docker and containerized workloads.
  • Strong Terraform and infrastructure-as-code skills.
  • Experience with GitHub Actions or comparable CI/CD tools.
  • Ability to work across multiple teams, services, and codebases to coordinate a cross-cutting migration.
  • Fluency with Codex or comparable AI-assisted coding tools.

Nica to have:

  • Prior experience running a similar New Relic/Datadog-to-Grafana (or equivalent) observability migration end to end.
  • Grafana-specific certifications or demonstrated community contributions (plugins, dashboards, etc.).
  • Experience with the broader Prometheus/Loki/Tempo (LGTM) ecosystem.
  • AWS certification, such as AWS Solutions Architect Associate/Professional.
  • Experience with RDS and other AWS data services.
  • Prior contract or consulting experience with clearly defined deliverables.
  • At least an Upper-Intermediate level of English 

We offer*:

  • Flexible working format - remote, office-based or flexible
  • A competitive salary and good compensation package
  • Personalized career growth
  • Professional development tools (mentorship program, tech talks and trainings, centers of excellence, and more)
  • Active tech communities with regular knowledge sharing
  • Education reimbursement
  • Memorable anniversary presents
  • Corporate events and team buildings
  • Other location-specific benefits

*not applicable for freelancers

Skills Required

  • Strong hands-on experience with Grafana, including dashboard design, alerting, and data source configuration
  • Practical experience migrating from New Relic and/or Datadog
  • Full-stack development experience with application code instrumentation across front-end and back-end services
  • Production experience with AWS, including EKS and Kubernetes
  • Practical experience with Docker and containerized workloads
  • Strong Terraform and infrastructure-as-code skills
  • Experience with GitHub Actions or comparable CI/CD tools
  • Ability to coordinate a cross-cutting migration across multiple teams, services, and codebases
  • Fluency with Codex or comparable AI-assisted coding tools
  • Prior experience running a similar observability migration end to end
  • Grafana-specific certifications or demonstrated community contributions
  • Experience with the Prometheus, Loki, and Tempo ecosystem
  • AWS certification, such as AWS Solutions Architect Associate or Professional
  • Experience with RDS and other AWS data services
  • Prior contract or consulting experience with clearly defined deliverables
  • At least Upper-Intermediate English proficiency
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Valletta
2,135 Employees
Year Founded: 2002

What We Do

N-iX is a global software solutions and engineering services company that helps world’s leading organizations turn challenges into lasting business value, operational efficiency, and revenue growth using advanced technology. Whether you need to build a custom solution, modernize your digital product or acquire extra tech expertise - we have the experience and capabilities to ensure your success. With over 2,000 professionals in 25 countries across Europe and the Americas, N-iX offers expert solutions in cloud, data analytics, embedded software, IoT, AI, machine learning, and other tech domains. Being in business for over two decades, we have worked with dozens of industry-leading enterprises and Fortune 500 companies creating value across a wide variety of sectors, including finance, manufacturing, supply chain, retail, e-commerce, healthcare, and more. Our unique combination of business domain expertise and technical know-how enables us to effectively collaborate with ISVs, tech companies, and enterprises of all sizes. Thanks to the strong tech ecosystem and partnerships with AWS, GCP, Microsoft, SAP, OpenText, Snowflake, and others, we bring extra speed, scale and efficiency to more than 160 organizations across the globe. N-iX is recognized by numerous industry awards, such as CRN Solution Provider 500, Global Outsourcing 100 by IAOP, ISG Provider Lens™, Modern Application Development services providers by Forrester, etc

Similar Jobs

Akamai Technologies Logo Akamai Technologies

Penetration Tester

Cloud • Security • Software • Cybersecurity
In-Office or Remote
2 Locations
10285 Employees

Akamai Technologies Logo Akamai Technologies

Senior User Experience Designer

Cloud • Security • Software • Cybersecurity
In-Office or Remote
2 Locations
10285 Employees

Akamai Technologies Logo Akamai Technologies

Senior Software Engineer

Cloud • Security • Software • Cybersecurity
In-Office or Remote
2 Locations
10285 Employees

Akamai Technologies Logo Akamai Technologies

Senior Site Reliability Engineer

Cloud • Security • Software • Cybersecurity
In-Office or Remote
2 Locations
10285 Employees

Similar Companies Hiring

Standard Template Labs Thumbnail
Artificial Intelligence • Information Technology • Software
New York, NY
25 Employees
NODA AI Thumbnail
Artificial Intelligence • Information Technology • Software • Cybersecurity
Sydney, AU
54 Employees
Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account