Senior AI/ML Engineer

Posted Yesterday
Easy Apply
Be an Early Applicant
Lisbon, PRT
Hybrid
Senior level
Artificial Intelligence • Cloud • Information Technology • Machine Learning • Software • Big Data Analytics • Automation
Empowering teams of all kinds to do the critical work that moves business forward through the PagerDuty Operations Cloud
The Role
Design, build, and operate production AI systems including LLM agents, retrieval pipelines, event intelligence, and real-time inference services. Own architecture, orchestration, evaluation, guardrails, observability, reliability, and cost optimization at scale. Partner with platform, product, and research teams, and provide technical mentorship while developing resilient distributed systems for high-volume event streams.
Summary Generated by Built In

PagerDuty (NYSE:PD) is a leader in Digital Operations Management. In an always-on world, organizations of all sizes trust PagerDuty to help them deliver a perfect digital experience to their customers, every time. Teams use PagerDuty to identify issues and opportunities in real time and bring together the right people to fix problems faster and prevent them in the future. Over 13,000 organizations (including 60 of Fortune 100) rely on PagerDuty to succeed with Digital Transformation, Cloud Migration, and DevOps Modernization. Notable customers include GE, Cisco, Genentech, Electronic Arts, Cox Automotive, Netflix, Shopify, Zoom, DoorDash, Lululemon and more. We are expanding rapidly as a platform for Digital Operations Management using AI/ML and Automation and growing our adoption by Development, IT, Customer Service, Security, and other teams across the organization.

About the role

PagerDuty’s Operations Cloud runs on a platform that ingests billions of signals and turns them into real-time action for thousands of customers. We’re looking for a Senior AI/ML Engineer who lives at the intersection of two disciplines: large-scale distributed systems and applied AI.

In this role you will design and ship AI systems that run in production at PagerDuty’s scale — powering Incident Management AI Agents, event intelligence, and the LLM-powered capabilities embedded across our platform. You’ll own the full lifecycle, from framing the problem to serving reliably at scale.

We are looking for a candidate who is genuinely passionate about building with modern AI — LLMs, agents, and retrieval — but grounded in the realities of building resilient, high-throughput systems.

What you’ll do
  • Design and build AI-powered features — LLM agents, retrieval, and event intelligence — that operate on high-volume, real-time event streams, from problem framing through production deployment and monitoring.
  • Architect and own the systems behind them: agent and prompt orchestration, retrieval pipelines, tool/API integrations, and low-latency inference and evaluation at scale.
  • Reason about consistency, throughput, fault tolerance, and cost across services that must stay reliable under bursty, unpredictable load.
  • Take AI features from prototype to production, establishing the evaluation, guardrail, observability, and improvement loops that keep them accurate and trustworthy over time.
  • Partner with platform, product, and applied-research teams to define what “good” looks like and to integrate AI cleanly into existing services.
  • Raise the bar through example, reviews and mentorship, and help shape the team’s technical direction.
What you’ll bring
  • 5+ years of software engineering experience, with meaningful time spent building and operating production distributed systems (high-throughput services, streaming/event-driven architectures, or large-scale data platforms).
  • Hands-on experience building and shipping AI systems in production — LLM-powered applications, agents, or retrieval — including the surrounding orchestration, serving, and evaluation, not just prototypes.
  • Strong programming fundamentals and comfort moving between systems and AI/application code.
  • Solid grounding in applied AI fundamentals: prompting, retrieval, agent patterns, and how to evaluate and guardrail LLM behavior.
  • Experience with cloud infrastructure (AWS, GCP, or Azure), containers, and orchestration (Kubernetes).
  • A pragmatic, reliability-minded mindset: you optimize for systems that work correctly at scale, and you can articulate the trade-offs behind your choices.
  • Strong communication and collaboration skills, and a track record of raising the quality of the teams and systems around you.
Nice to have
  • Experience with LLMOps tooling and patterns — evaluation harnesses, prompt/version management, tracing and observability for agents, and online/offline eval consistency.
  • Deep experience serving LLM-based systems in production, including retrieval-augmented generation, multi-step agents, and tool use.
  • Background in anomaly detection, event correlation, or applied problems in observability, AIOps, or reliability.
  • Familiarity with the ecosystem — e.g. LLM APIs and frameworks such as LangChain or LlamaIndex, vector databases, and distributed data/compute tools such as Kafka, Airflow, or Spark.
  • Contributions to open-source AI or distributed-systems projects.
Why PagerDuty

At PagerDuty, AI it’s core to how we help the world’s teams keep their digital services running. You’ll work on problems where scale, latency, and correctness genuinely matter, alongside engineers who care about building systems that people depend on in their most critical moments.


Hesitant to apply?

We encourage you to submit your resume even if you don't meet every requirement. We value potential and consider each candidate's full professional story. Whether you're exploring a career change or taking your next step, we look forward to reviewing your application. If this just isn’t the right role or time - sign up for job alerts!

Where we work

PagerDuty operates a hybrid work model with offices in 8 major cities: Atlanta, Lisbon, London, San Francisco, Santiago, Sydney, Tokyo, and Toronto. While we offer flexibility within our established locations, we cannot employ candidates residing in:

Location restrictions:
Australia: Northern Territory, Queensland, South Australia, Tasmania, Western Australia
Canada: Alberta, Manitoba, Newfoundland, Northwest Territories, Nunavut, PEI, Quebec, Saskatchewan, Yukon
United States: Alaska, Hawaii, Iowa, Louisiana, Mississippi, Nebraska, New Mexico, Oklahoma, Rhode Island, South Dakota, West Virginia, Wyoming
Candidates must reside in an eligible location, which vary by role.

How we work

Our values guide how we support customers, collaborate with colleagues, develop products, and foster a culture of belonging. They define not just our actions, but what it means to be Dutonian.

People Leaders at PagerDuty are responsible for creating high performance environments that drive accountability. PagerDuty has four key dimensions that define our Leadership Impact: Lead Self, Lead the Team, Lead the Business, and Lead the Future. Each dimension has three associated competencies to give leaders a shared language for guiding their development, career, promotion, and succession planning discussions. Our Manager Expectations serve as a practical guide for managers to understand their responsibilities, prioritize their efforts, and drive engagement and performance.

What we offer

As a global organization, our total rewards approach is competitive with industry standards and aligned with local laws and regulations. Learn more, including country-specific offerings, on our benefits site.

Your package may include:

  • Competitive salary
  • Comprehensive benefits package 
  • Flexible work arrangements
  • Company equity*
  • ESPP (Employee Stock Purchase Program)*
  • Retirement or pension plan*
  • Generous paid vacation time
  • Paid holidays and sick leave
  • Dutonian Wellness Days & HibernationDuty - companywide paid days off in addition to PTO
  • Paid parental leave: 22 weeks for pregnant parent, 12 weeks for non-pregnant parent (some countries have longer leave standards and we comply with local laws)*
  • Paid volunteer time off: 20 hours per year
  • Company-wide hack weeks
  • Mental wellness programs

*Eligibility may vary by role, region, and tenure

About PagerDuty

PagerDuty, Inc. (NYSE:PD) is a global leader in digital operations management. The PagerDuty Operations Cloud is an AI-powered platform that empowers business resilience and drives operational efficiency for enterprises. With a generative AI assistant at its core, PagerDuty empowers teams to detect and resolve issues in real time, orchestrate complex workflows, and drive continuous improvement across their digital operations. Trusted by nearly half of both the Fortune 500 and the Forbes AI 50, as well as approximately two-thirds of the Fortune 100, PagerDuty is essential for delivering always-on digital experiences to modern businesses

PagerDuty is Great Place to Work-certified™, a Fortune Best Workplace for Millennials, a Fortune Best Medium Workplace, a Fortune Best Workplace in Technology, and a top rated product on TrustRadius and G2. 

Go behind-the-scenes on our careers site and @pagerduty on Instagram.

Additional Information

PagerDuty is an equal opportunity employer. PagerDuty does not discriminate on the basis of race, religion, color, national origin, gender, sexual orientation, age, marital status, parental status, veteran status, or disability status. Your privacy is important to us. By submitting an application, you confirm that you have read and understand PagerDuty's Privacy Policy.

PagerDuty is committed to providing reasonable accommodations for qualified individuals with disabilities in our job application process. Should you require accommodation, please email [email protected] and we will work with you to meet your accessibility needs.

PagerDuty uses the E-Verify employment verification program.

Skills Required

  • 5+ years of software engineering experience
  • Experience building and operating production distributed systems, including high-throughput services, streaming or event-driven architectures, or large-scale data platforms
  • Hands-on experience building and shipping production AI systems, including LLM-powered applications, agents, or retrieval systems
  • Experience with AI orchestration, serving, and evaluation
  • Strong programming fundamentals and ability to work across systems and AI/application code
  • Knowledge of prompting, retrieval, agent patterns, LLM evaluation, and guardrails
  • Experience with cloud infrastructure such as AWS, GCP, or Azure
  • Experience with containers and Kubernetes
  • Reliability-focused approach to scalable systems and technical trade-offs
  • Strong communication and collaboration skills
  • Experience with LLMOps tooling and patterns
  • Experience serving LLM-based systems in production, including RAG, multi-step agents, and tool use
  • Background in anomaly detection, event correlation, observability, AIOps, or reliability
  • Familiarity with LangChain, LlamaIndex, vector databases, Kafka, Airflow, or Spark
  • Contributions to open-source AI or distributed-systems projects

What the Team is Saying

Jen
Suzan
Kyle
Anne
Hannah
Jhanae
Ben
Alan
Kurt
Evelyn Bassett
Caroline Hood
Dormain Drewitz
Mandi Walls
Kurt
Karen
Vince

PagerDuty Compensation & Benefits Highlights

  • Leave & Time Off Breadth Companywide programs like HibernationDuty, Dutonian Wellness Days, and a paid year‑end closure add coordinated rest on top of regular PTO. Careers materials describe extra companywide time off that helps employees disconnect at the same time.
  • Healthcare Strength Medical, dental, and vision coverage start on the hire date, with options such as Aetna nationwide and Kaiser in California. Mental‑health resources like Headspace Care, a global EAP, and health advocacy support are explicitly highlighted.
  • Parental & Family Support Paid parental leave is described as generous for both pregnant and non‑pregnant parents, and family‑planning support is provided via Cleo. Adoption assistance and return‑to‑work support are also noted in company benefits materials.

PagerDuty Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
Atlanta, GA
1,200 Employees
Year Founded: 2009

What We Do

PagerDuty, Inc. (NYSE:PD) is a global leader in digital operations management, enabling customers to achieve operational efficiency at scale with the PagerDuty Operations Cloud. The PagerDuty Operations Cloud combines AIOps, Automation, Customer Service Operations and Incident Management with a powerful generative AI assistant to create a flexible, resilient and scalable platform to increase innovation velocity, grow revenue, reduce cost, and mitigate the risk of operational failure. Half of the Fortune 500 and nearly 70% of the Fortune 100 rely on PagerDuty as essential infrastructure for the modern enterprise. PagerDuty is Great Place to Work-certified™, a Fortune Best Workplace for Millennials, a Fortune Best Medium Workplace, a Fortune Best Workplace in Technology, and a top rated product on TrustRadius and G2. Go behind-the-scenes at careers.pagerduty.com and @pagerduty on Instagram.

Why Work With Us

PagerDuty offers a hybrid, flexible environment where ambition thrives. We champion innovation, learning, and growth through our core values: Champion the Customer, Take the Lead, Run Together, Ack & Own, and Bring Your Self. Join us!

Gallery

Gallery
Gallery
Gallery
Gallery
Gallery
Gallery

PagerDuty Offices

Hybrid Workspace

Employees engage in a combination of remote and on-site work.

We offer a hybrid, flexible workplace, while also providing ample opportunities for connection in-person and virtually with your colleagues.

Typical time on-site: 1 days a week
Company Office Image
Atlanta, GA
Company Office Image
Lisbon, PT
Company Office Image
London, GB
Company Office Image
San Francisco, CA
Company Office Image
Santiago, CL
Company Office Image
Sydney, NSW
Company Office Image
Tokyo, JP
Company Office Image
Toronto, Ontario
Learn more

Similar Jobs

PagerDuty Logo PagerDuty

Machine Learning Engineer

Artificial Intelligence • Cloud • Information Technology • Machine Learning • Software • Big Data Analytics • Automation
Easy Apply
Hybrid
Lisbon, PRT
1200 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account