Director, AI Platform Reliability

Posted 5 Hours Ago
Easy Apply
Be an Early Applicant
San Francisco, CA, USA
Hybrid
248K-275K Annually
Expert/Leader
Artificial Intelligence • Cloud • Information Technology • Machine Learning • Software
Hybrid Observability powered by AI
The Role
Leads multiple engineering teams responsible for highly scalable, reliable distributed systems and data platforms. Defines architecture and technical strategy for Java microservices, Kafka pipelines, data lakes, and cloud-native services. Owns reliability objectives, observability, disaster recovery, performance, capacity planning, and operational excellence. Partners across product, infrastructure, security, data, and SRE teams while recruiting and developing senior engineering leaders, improving developer productivity, and managing modernization, technical debt, and cloud costs.
Summary Generated by Built In

About Us:  

We love going to work and think you should too. Our team is dedicated to trust, customer obsession, agility, and striving to be better everyday. These values serve as the foundation of our culture, guiding our actions and driving us towards excellence. We foster a culture of performance and recognition, allowing us to transform growth as we enable our employees to do the best work of their careers.

This role is open to candidates based in or near San Francisco, CA. At LogicMonitor, we hire within our Centers of Energy—vibrant locations where our teams connect, collaborate, and innovate.

To learn more about life at LogicMonitor, check out our Careers Page.

What You'll Do:

LogicMonitor® is the AI-first hybrid observability platform powering the next generation of digital infrastructure. LogicMonitor delivers complete visibility and actionable intelligence across on-premises, cloud, and edge environments. By anticipating issues before they strike, optimizing resources in real time, and enabling faster, smarter decisions, LogicMonitor helps IT and business leaders protect margins, accelerate innovation, and deliver exceptional digital experiences without compromise.

Our customers love LogicMonitor's ability to bring cloud and traditional IT together into one view, as seen in minimal churn rates, expansion business, and exciting new customer references. In fact, LogicMonitor has received the highest Net Promoter Score of any IT Infrastructure Management provider. LogicMonitor also boasts high employee satisfaction. We have been certified as a Great Place To Work®, and named one of BuiltIn's Best Places to Work for the seventh year in a row! 

We are looking for an accomplished and hands-on Director of AI Platform Reliability to lead the architecture, development, and operation of highly scalable, distributed software platforms.

This leader will be responsible for systems that process hundreds of millions/billions of transactions and events, manage terabytes to petabytes of data, and deliver reliable, low-latency services to enterprise customers. The ideal candidate combines strong engineering depth in Java, Kafka, distributed systems, and cloud-native microservices with a demonstrated ability to build and lead high-performing engineering organizations.

This is a strategic leadership role, but it requires a leader who can remain close to the technology, participate in architecture reviews, challenge design decisions, guide teams through complex production problems, and establish the engineering practices required to operate mission-critical platforms at scale.

Here's a closer look at this key role:

  • Lead and scale multiple engineering teams responsible for high-volume, business-critical distributed systems and data platforms.
  • Define the technical strategy and architecture for platforms processing hundreds of millions of transactions and terabytes of data.
  • Guide the development of Java-based microservices, APIs, Kafka streaming pipelines, batch-processing workflows, and cloud-native services.
  • Build and evolve scalable data lake and Data Lakehouse platforms supporting real-time, near-real-time, and batch analytics workloads.
  • Establish reliable data ingestion, transformation, storage, governance, lineage, retention, and data-quality practices across streaming and batch pipelines.
  • Build low-latency, highly available, fault-tolerant systems with strong scalability, resiliency, and disaster-recovery capabilities.
  • Define and own operational SLAs, SLOs, availability targets, recovery objectives, and performance metrics for critical services and data pipelines.
  • Drive capacity planning, load testing, throughput optimization, and improvements to p95 and p99 latency.
  • Ensure effective Kafka design, including partitioning, consumer groups, ordering, schema evolution, replay, and lag management.
  • Establish engineering standards for architecture, coding, testing, security, observability, and production readiness.
  • Partner with Product, Architecture, SRE, Security, Data, and Infrastructure teams to deliver strategic platform initiatives.
  • Strengthen operational excellence through monitoring, incident management, on-call practices, root-cause analysis, and continuous reliability improvements.
  • Recruit, mentor, and develop engineering managers, architects, and senior technical leaders.
  • Improve developer productivity, CI/CD automation, deployment safety, and release predictability.
  • Manage technical debt, platform modernization, cloud costs, and long-term scalability investments.
What You'll Need:
  • 10+ years of professional software-engineering experience, including significant experience building large-scale distributed systems.
  • Experience leading engineering teams, architects, and staff engineers.
  • Demonstrated success delivering and operating platforms that process hundreds of millions of transactions, requests, or events.
  • Deep technical expertise in Java, JVM performance, concurrency, multithreading, memory management, and application profiling.
  • Strong experience designing and operating microservice-based and event-driven architectures.
  • Extensive production experience with Apache Kafka or a comparable distributed streaming platform.
  • Strong understanding of Kafka partitioning, replication, consumer groups, offset management, ordering, delivery semantics, schema evolution, and reprocessing.
  • Experience designing low-latency, highly available APIs and backend services.
  • Experience managing terabyte- or petabyte-scale datasets across relational, NoSQL, streaming, and object-storage technologies.
  • Strong understanding of distributed-systems concepts, including consensus, replication, partitioning, consistency models, idempotency, backpressure, and fault tolerance.
  • Experience operating cloud-native applications using Kubernetes, containers, infrastructure as code, and automated CI/CD pipelines.
  • Experience with at least one major cloud platform, such as AWS, Google Cloud, or Microsoft Azure.
  • Strong knowledge of observability practices involving metrics, logs, traces, profiling, alerting, dashboards, and service-level objectives.
  • Demonstrated experience improving system reliability, scalability, latency, cost efficiency, and engineering productivity.
  • Strong written and verbal communication skills, including the ability to explain complex technical decisions to engineering teams, executives, and business stakeholders.
  • Proven ability to build inclusive, accountable, and high-performing engineering organizations.

Residents of California, click Here to view our California Applicant Privacy Notice.

Anticipated Application Close Date: 10/26/26

LogicMonitor is an Equal Opportunity Employer
At LogicMonitor, we believe that innovation thrives when every voice is heard and each individual is empowered to bring their unique perspective. We’re committed to creating a workplace where diversity is celebrated, and all employees feel inspired and supported to contribute their best.

For us, equal opportunity means fostering a truly inclusive culture where everyone has the chance to grow and succeed. We don’t just open doors; we invite you to step through and be part of something bigger. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, or status as a protected veteran.

Work Authorization:
At this time, we are able to consider candidates who are authorized to work in the United States on a full-time, permanent basis without requiring new or initial employer-sponsored work authorization.
Candidates who currently hold valid U.S. work authorization that can be transferred to a new employer (such as certain H-1B statuses) may be considered on a case-by-case basis.
We are not able to provide new sponsorship for employment-based visas that require an initial petition or application by the employer.

#LI-JP1 #LI-Hybrid #BI-Hybrid

LogicMonitor is dedicated to fostering a culture of transparency and fairness, including our commitment to pay transparency. We provide the base salary ranges for all positions posted within the United States. 

Compensation packages at LogicMonitor for eligible roles include base salary, a variable plan depending on role, along with comprehensive benefits. The range displayed on each job posting reflects the minimum and maximum base salary target for new hires in the position, determined by work location and additional factors, including job-related skills, experience, interview performance, and relevant education or training. As part of our holistic compensation philosophy, your package will also include, but is not limited to: Comprehensive health, dental and vision coverage, generous parental leave policies, access to our Employee Assistance Program and various Wellness programs, a 401K with company matching, a Lifestyle Spending Account, and an unlimited vacation policy. For more information on our benefits, see our careers page.

The Base Salary range for this role is:
$247,500$275,000 USD

                                               

Our goal is to ensure an accessible and inclusive experience for every candidate.

If you need a reasonable accommodation during the application or interview process under applicable local law, please submit a request via this Accommodation Request Form.

Know your rights: workplace discrimination is illegal. Please click here to review LogicMonitor’s U.S. Pay Transparency Nondiscrimination Provision.

Skills Required

  • 10+ years of professional software engineering experience, including significant large-scale distributed systems experience
  • Experience leading engineering teams, architects, and staff engineers
  • Experience delivering and operating platforms processing hundreds of millions of transactions, requests, or events
  • Deep expertise in Java, JVM performance, concurrency, multithreading, memory management, and application profiling
  • Experience designing and operating microservice-based and event-driven architectures
  • Extensive production experience with Apache Kafka or a comparable distributed streaming platform
  • Strong knowledge of Kafka partitioning, replication, consumer groups, offset management, ordering, delivery semantics, schema evolution, and reprocessing
  • Experience designing low-latency, highly available APIs and backend services
  • Experience managing terabyte- or petabyte-scale datasets across relational, NoSQL, streaming, and object-storage technologies
  • Strong understanding of distributed-systems concepts, including consensus, replication, partitioning, consistency models, idempotency, backpressure, and fault tolerance
  • Experience operating cloud-native applications using Kubernetes, containers, infrastructure as code, and automated CI/CD pipelines
  • Experience with at least one major cloud platform, such as AWS, Google Cloud, or Microsoft Azure
  • Strong knowledge of observability practices involving metrics, logs, traces, profiling, alerting, dashboards, and service-level objectives
  • Experience improving system reliability, scalability, latency, cost efficiency, and engineering productivity
  • Strong written and verbal communication skills for explaining complex technical decisions to technical and business stakeholders
  • Ability to build inclusive, accountable, and high-performing engineering organizations

What the Team is Saying

Kenyon
Franky
Kwame
Gisselle
Antonio
Carly
Chris
Rockel
Rob
David
Peyton

LogicMonitor Compensation & Benefits Highlights

  • Leave & Time Off Breadth Time off includes “unlimited” vacation plus a separate bucket of 12 paid sick/personal days per year, alongside paid volunteer time off and hybrid/remote flexibility. These programs are explicitly listed for U.S. roles and are positioned to support work-life balance.
  • Healthcare Strength Coverage spans multiple medical, dental, and vision options with HSA/FSA access, plus mental‑health resources such as an EAP and Calm. Company‑paid short‑ and long‑term disability at 60% of earnings and basic life/AD&D are also included.
  • Parental & Family Support Support includes Maven for fertility, family planning, parenting/pediatrics, and menopause, along with generous parental leave and adoption assistance. Abortion‑related travel support is also listed.

LogicMonitor Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Santa Barbara, CA
1,100 Employees
Year Founded: 2007

What We Do

LogicMonitor® is the AI-first hybrid observability platform powering the next generation of digital infrastructure. LogicMonitor delivers complete visibility and actionable intelligence across on-premises, cloud, and edge environments. By anticipating issues before they strike, optimizing resources in real time, and enabling faster, smarter decisions, LogicMonitor helps IT and business leaders protect margins, accelerate innovation, and deliver exceptional digital experiences without compromise. For more information, visit www.logicmonitor.com and our blog, or follow us on LinkedIn, X, Facebook, and YouTube.

Why Work With Us

We love going to work and think you should too. We are customer-obsessed, work as one agile team, and strive to be better every day while building trust. These are our core values. So it's no surprise that we work hard and genuinely have fun working with each other as we expand our global presence and achieve record-breaking success.

Gallery

Gallery
Gallery
Gallery
Gallery
Gallery
Gallery
Gallery
Gallery
Gallery
Gallery

LogicMonitor Offices

Hybrid Workspace

Employees engage in a combination of remote and on-site work.

We call our offices Centers of Energy, because they’re where we accelerate work, spark creativity, and ignite our culture of connection and celebration. Our teams coordinate their time in Centers of Energy to reflect how they work best.

Typical time on-site: Flexible
Company Office Image
HQSanta Barbara, CA
Company Office Image
Austin, TX
Company Office Image
Boston, MA
Company Office Image
London, UK
Company Office Image
Pune, IN
Company Office Image
San Francisco
Company Office Image
Singapore
Company Office Image
Sydney, Australia
Learn more

Similar Jobs

LogicMonitor Logo LogicMonitor

Sr. Evaluation Engineer

Artificial Intelligence • Cloud • Information Technology • Machine Learning • Software
Easy Apply
Hybrid
San Francisco, CA, USA
1100 Employees
158K-218K Annually

LogicMonitor Logo LogicMonitor

AI Operations & GTM Engineer

Artificial Intelligence • Cloud • Information Technology • Machine Learning • Software
Easy Apply
Hybrid
San Francisco, CA, USA
1100 Employees
158K-218K Annually

LogicMonitor Logo LogicMonitor

Sr. Forward Deployed Engineer

Artificial Intelligence • Cloud • Information Technology • Machine Learning • Software
Easy Apply
Hybrid
San Francisco, CA, USA
1100 Employees
131K-175K Annually

LogicMonitor Logo LogicMonitor

Architect

Artificial Intelligence • Cloud • Information Technology • Machine Learning • Software
Easy Apply
Remote or Hybrid
San Francisco, CA, USA
1100 Employees
136K-190K Annually

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account