Platform Reliability Engineer

Posted 7 Days Ago
Be an Early Applicant
Montréal, QC, CAN
In-Office
Senior level
Information Technology • Consulting
The Role
Build and operate reliability and observability capabilities for an internal AWS Bedrock AgentCore platform. Responsibilities include OpenTelemetry instrumentation, CloudWatch and enterprise observability integrations, distributed tracing, SLOs, alerting, token and compute cost monitoring, dashboards, runbooks, incident response, and post-incident improvements. The role requires strong AWS, SRE, observability, FinOps, scripting, and coding experience, with preferred exposure to LLM monitoring, infrastructure as code, and regulated industries.
Summary Generated by Built In

About us

Appnovation is a global, full-service digital partner that combines Strategy, Experience & Design, Engineering and Managed Services. We build digital solutions that deliver real impact today and serve as foundations for future growth.  Bold ambition. Practical action. Endless possibilities.

We’re putting together a dedicated delivery pod to build and run an internal agent platform on AWS Bedrock AgentCore for a global life sciences client. The pod works as one team with the client’s engineers to deliver the platform other teams will build their agents on.

In this role you make the platform visible and dependable. You’ll build the OpenTelemetry instrumentation, observability integrations, tracing, SLOs, alerting, and cost monitoring for token and compute spend. When an agent is slow, wrong or expensive, your work is how the team finds out and fixes it. This is a named-team engagement, so the person we propose is the person who starts.

ROLE RESPONSIBILITIES
  • Instrumentation: Build OpenTelemetry instrumentation standards for agents, tools and platform services, and make it easy for teams to adopt.
  • Observability Integration: Connect AgentCore Observability and CloudWatch with the client’s existing observability tools (e.g., Datadog, Splunk, Grafana).
  • Tracing: Set up end-to-end tracing across agent steps, model calls, tool calls and multi-agent handoffs so issues can be traced to their source.
  • SLOs and Alerting: Define SLOs for latency, availability and error rates, and build alerting that is useful and not noisy.
  • Cost Monitoring: Track token usage and compute spend by team, agent and environment, with dashboards, budgets and alerts for unexpected spikes.
  • Incident Readiness: Write runbooks, support incident response and run post-incident reviews that lead to real fixes.
QUALIFICATIONS
  • 6+ years in SRE, platform reliability or observability engineering, with strong hands-on AWS experience.
  • Hands-on experience with OpenTelemetry (SDKs, collectors, exporters) and distributed tracing.
  • Experience with Amazon CloudWatch and at least one enterprise observability platform (Datadog, Splunk, Grafana, New Relic or similar).
  • Experience defining and running SLOs, error budgets and alerting strategies.
  • Experience building cost visibility and FinOps reporting on AWS (Cost Explorer, CUR, tagging strategies).
  • Scripting and coding skills in Python, Go or TypeScript.
  • Clear communication and able to turn data into decisions for technical and non-technical audiences.
PREFERRED QUALIFICATIONS
  • Experience monitoring LLM or agent-based applications, including token usage, latency and quality signals.
  • Familiarity with LLM observability tools (e.g., Langfuse, Arize, LangSmith) or AgentCore Observability.
  • Experience with infrastructure as code (Terraform or AWS CDK) for monitoring resources.
  • Experience in pharma, life sciences or another regulated industry.
WHO YOU ARE
  • You want to know what’s happening in a system before a user tells you
  • You build alerts people trust and dashboards people actually open
  • You treat cost as a reliability concern, not an afterthought
  • You stay calm in incidents and focus on learning afterward
  • You work well inside a client team and build trust quickly
  • You have prior experience in consulting
  • Prior experience and connections in the Life Sciences industry is preferred
Thank you for your interest in a career with Appnovation Technologies! Please note that only those selected for an interview will be contacted.
 
At Appnovation, we recognize that diverse teams are the strongest teams. Diversity, Equity & Inclusion is not only something that we embrace - we celebrate it! We are proud to be an Equal Opportunity Employer and we encourage applicants from all backgrounds, lived experiences and industries to apply. Come join us at Appnovation, and learn more about how we stay true to our company values as we build better lives through better digital.

Accommodations are available upon request throughout the recruitment process.

Skills Required

  • 6+ years of experience in SRE, platform reliability, or observability engineering
  • Strong hands-on AWS experience
  • Hands-on experience with OpenTelemetry SDKs, collectors, exporters, and distributed tracing
  • Experience with Amazon CloudWatch and at least one enterprise observability platform such as Datadog, Splunk, Grafana, or New Relic
  • Experience defining and running SLOs, error budgets, and alerting strategies
  • Experience building AWS cost visibility and FinOps reporting using Cost Explorer, CUR, and tagging strategies
  • Scripting and coding skills in Python, Go, or TypeScript
  • Clear communication skills and ability to turn data into decisions for technical and non-technical audiences
  • Experience monitoring LLM or agent-based applications, including token usage, latency, and quality signals
  • Familiarity with LLM observability tools such as Langfuse, Arize, LangSmith, or AgentCore Observability
  • Experience with infrastructure as code using Terraform or AWS CDK
  • Experience in pharma, life sciences, or another regulated industry
  • Prior experience in consulting
  • Prior experience and connections in the life sciences industry
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Vancouver
372 Employees
Year Founded: 2007

What We Do

Inspiring Possibility Appnovation is a full service digital consultancy with experience and capacity to meet the needs of even the largest most complex of organizations in the world. Our services portfolio enables us to offer clients the best of experiences when working with our teams so as to make sure we keep the focus on their needs, customers and delivering tangible value to the business. End to end services; endless ideas.

Similar Jobs

Hadrian Logo Hadrian

Site Reliability Engineer

Aerospace • Hardware • Software • Defense • Manufacturing
In-Office or Remote
9 Locations
550 Employees
175K-285K Annually

Hewlett Packard Enterprise Logo Hewlett Packard Enterprise

Cloud Engineer

Artificial Intelligence • Cloud • Information Technology • Consulting
In-Office
Saint-Laurent, Montréal, QC, CAN
85422 Employees
65K-108K Annually

NBCUniversal Logo NBCUniversal

Designer

AdTech • Cloud • Digital Media • Information Technology • News + Entertainment • App development
Hybrid
Montréal, QC, CAN

NBCUniversal Logo NBCUniversal

Designer

AdTech • Cloud • Digital Media • Information Technology • News + Entertainment • App development
Remote or Hybrid
Montréal, QC, CAN

Similar Companies Hiring

Axle Health Thumbnail
Artificial Intelligence • Healthtech • Information Technology • Logistics
Santa Monica, CA
25 Employees
NODA AI Thumbnail
Artificial Intelligence • Information Technology • Software • Cybersecurity
Sydney, AU
54 Employees
Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account