Director, Platform Engineering

Posted 2 Days Ago
Be an Early Applicant
Kraków, Małopolskie, POL
In-Office
Expert/Leader
Healthtech
The Role
Leads platform engineering and end-to-end reliability for AI infrastructure, including cloud and accelerated compute, Kubernetes, CI/CD, infrastructure as code, observability, incident response, application support, and release engineering. Establishes SLOs, error budgets, and operational practices; manages GPU capacity, performance, and cost; supports internal AI teams; and defines reliable infrastructure and governance for LLM and agentic workloads. Builds and scales distributed engineering teams across time zones.
Summary Generated by Built In

Bring more to life.

At Danaher, our work saves lives. And each of us plays a part. Fueled by our culture of continuous improvement, we turn ideas into impact – innovating at the speed of life.

Our 60,000+ associates work across the globe at more than 15 unique businesses within life sciences, diagnostics, and biotechnology. 

Are you ready to accelerate your potential and make a real difference? At Danaher, you can build an incredible career at a leading science and technology company, where we’re committed to hiring and developing from within. You’ll thrive in a culture of belonging where you and your unique viewpoint matter.

Learn about the Danaher Business System which makes everything possible.

The Director, Platform Engineering role is responsible for end-to-end reliability of the platform behind our AI initiatives via accelerated compute, cloud infrastructure, DevOps and release engineering, application support, and incident response. Your customers are the teams building AI into consequential work across the company: molecular design, autonomous labs, supply chain, professional services, and more. You will keep their systems available and performant today, and define what reliability looks like as workloads become increasingly LLM-driven and agentic.

This position reports to the Senior Director, Data and AI Platform, part of the Chief Information Officer (CIO) office and will be located onsite in Kraków, Poland.

This is a Danaher Corporate role, hosted by our Cytiva operating company in Kraków.

In this role, you will have the opportunity to:

  • End-to-end platform reliability. Own availability, performance, and recovery for the compute and cloud platform underpinning the company's AI initiatives — you are accountable for whether the systems our AI teams depend on are working, not just for the components your team builds.

  • Incident management and response. Lead the incident lifecycle across the platform: detection, triage, mitigation, and post-incident review. Establish SLOs, error budgets, and on-call practices that hold across a diverse set of workloads and drive measurable reduction in time-to-detect and time-to-mitigate.

  • Compute and cloud infrastructure for AI workloads. Design and operate the GPU/accelerator fleet, scheduling and capacity management, storage, and networking that support large-scale training, inference, and simulation — balancing utilization, cost, and the reliability expectations of production-facing teams.

  • Application support and operations. Own the operational layer between infrastructure and the applications running on it: deployment, configuration, runtime health, and the support model that gets AI teams unblocked quickly when something breaks.

  • DevOps and delivery pipelines. Own CI/CD, infrastructure as code, environment management, and release engineering so that changes ship rapidly and safely, and so that reliability is enforced in the pipeline rather than discovered in production.

  • Serving internal AI customer teams. Treat corporate-wide AI initiatives — molecular design, autonomous labs, supply chain, professional services, and others — as your customers. Gather their reliability and capacity requirements, translate friction into scoped platform work, and prove impact through before/after measurement.

  • Building toward agentic systems. Set the reliability strategy for the next generation of LLM-based and agentic workloads: observability and evaluation for non-deterministic systems, guardrails and governance for agents acting in production, and the infrastructure patterns that make agentic execution safe, traceable, and dependable at scale.

The essential requirements of this job include:
  • Education: Bachelor's degree in Computer Science, Engineering, or a related technical field, or equivalent practical experience; advanced degree welcome but not required.

  • Functional experience: 10+ years in engineering, with significant time owning production reliability, SRE, or platform operations for large-scale distributed systems, and demonstrated ownership of reliability outcomes at organizational scale — you can point to concrete before/after results in availability, MTTR, incident volume, or error-budget performance from programs you led.

  • Leadership experience: Proven success leading and scaling multidisciplinary engineering organizations, including platform tech leads and senior engineers across distributed time zones — recruiting, coaching, setting technical direction, and holding a team accountable for operational outcomes.

  • Deep, hands-on experience running incident management for business-critical systems: on-call design, escalation paths, blameless post-incident review, and the follow-through that turns findings into permanent fixes.

  • Production experience with major cloud platforms (AWS, Azure, or GCP), containerization and orchestration (Docker, Kubernetes), and infrastructure as code (Terraform / Pulumi / etc), operating infrastructure for ML/AI workloads at scale — accelerated compute, distributed training or high-throughput inference, job scheduling, and the associated capacity and cost management.

  • Strong observability expertise: metrics, logging, tracing, and SLO instrumentation, plus a track record of driving toil reduction and automation rather than adding headcount to absorb operational load.

  • Excellent written and verbal communication, and the ability to work with demanding internal customers — negotiating trade-offs, setting expectations, and representing platform reliability to senior leadership.

Travel, Motor Vehicle Record & Physical/Environment Requirements:

  • Ability to travel – up to 20%

It would be a plus if you also possess previous experience in:

  • Hands-on experience with LLM and agentic systems in production — inference serving, evaluation and guardrails, or observability for non-deterministic workloads.

  • Familiarity with scientific or research computing environments (HPC, lab automation, instrument data pipelines) or other domains where AI workloads sit close to physical systems.

  • Experience operating in compliance-constrained or isolated environments (SOC 2, FedRAMP, GxP, or similar), or across multiple clouds and on-premise infrastructure.

#LI-KK1

Join our winning team today. Together, we’ll accelerate the real-life impact of tomorrow’s science and technology. We partner with customers across the globe to help them solve their most complex challenges, architecting solutions that bring the power of science to life.

For more information, visit www.danaher.com.

Skills Required

  • Bachelor's degree in Computer Science, Engineering, or a related technical field, or equivalent practical experience
  • 10+ years of engineering experience
  • Significant experience owning production reliability, SRE, or platform operations for large-scale distributed systems
  • Demonstrated ownership of organizational-scale reliability outcomes, including availability, MTTR, incident volume, or error-budget performance
  • Experience leading and scaling multidisciplinary engineering organizations across distributed time zones
  • Experience recruiting, coaching, setting technical direction, and holding teams accountable for operational outcomes
  • Hands-on experience with incident management, on-call design, escalation paths, blameless post-incident reviews, and permanent remediation
  • Production experience with AWS, Azure, or GCP
  • Experience with Docker and Kubernetes
  • Experience with infrastructure as code, such as Terraform or Pulumi
  • Experience operating infrastructure for ML/AI workloads at scale, including accelerated compute, distributed training or high-throughput inference, scheduling, capacity, and cost management
  • Strong observability expertise across metrics, logging, tracing, and SLO instrumentation
  • Track record of reducing toil and using automation to reduce operational load
  • Excellent written and verbal communication skills, including managing internal customers and communicating with senior leadership
  • Hands-on experience with production LLM and agentic systems, including inference serving, evaluation, guardrails, or observability
  • Familiarity with scientific or research computing, HPC, lab automation, or instrument data pipelines
  • Experience in compliance-constrained or isolated environments such as SOC 2, FedRAMP, or GxP
  • Experience across multiple clouds and on-premise infrastructure

Danaher Corporation Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Danaher Corporation and has not been reviewed or approved by Danaher Corporation.

  • Healthcare Strength Healthcare coverage is described as comprehensive, including medical plan options alongside dental, vision, life, disability, and mental health support. Wellness initiatives and support programs such as an EAP and vaccination or fitness offerings add breadth beyond core insurance.
  • Retirement Support Retirement benefits include a 401(k) plan with employer matching and options such as pre-tax and Roth contributions. Broader financial rewards such as performance bonuses and access to equity or an employee stock purchase plan are also described.
  • Parental & Family Support Parental leave and family-building support are described as available, including maternity and paternity leave and fertility assistance. Childcare and eldercare support are also highlighted as part of the overall package.

Danaher Corporation Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Washington , DC
57,802 Employees
Year Founded: 1984

What We Do

Danaher is a global science and technology innovator committed to helping our customers solve complex challenges and improve quality of life around the world. A global network of more than 25 operating companies, we drive meaningful innovation in some of today’s most dynamic industries through our operating companies in four strategic platforms: Life Sciences, Diagnostics, Water Quality and Product Identification. The engine at the heart of our success is the Danaher Business System (DBS), a set of tools that enables continuous improvement around lean, growth and leadership. Through the ingenuity of our people, the power of DBS and the impact of our meaningful technologies, we help realize life’s potential in ourselves and for those we serve.

Similar Jobs

Mondelēz International Logo Mondelēz International

Senior Director, Global Supply Chain Excellence Program & Focused Improvement Lead

Big Data • Food • Hardware • Machine Learning • Retail • Automation • Manufacturing
Remote or Hybrid
29 Locations
90000 Employees
174K-287K Annually

Capco Logo Capco

Senior Delivery Lead Manager – Modern Core Banking (She/ He/ They)

Fintech • Professional Services • Consulting • Energy • Financial Services • Cybersecurity • Generative AI
Remote or Hybrid
Poland
6000 Employees

Capco Logo Capco

Senior Manager – Risk & Regulatory (She/ He/ They)

Fintech • Professional Services • Consulting • Energy • Financial Services • Cybersecurity • Generative AI
Remote or Hybrid
Poland
6000 Employees

Capco Logo Capco

Senior Manager – Payments (She/ He/ They)

Fintech • Professional Services • Consulting • Energy • Financial Services • Cybersecurity • Generative AI
Remote or Hybrid
Poland
6000 Employees

Similar Companies Hiring

Sailor Health Thumbnail
Healthtech • Social Impact • Telehealth
New York City, NY
20 Employees
Granted Thumbnail
Artificial Intelligence • Healthtech • Insurance • Mobile • Financial Services
New York, New York
23 Employees
OneImaging Thumbnail
Healthtech
Miami, FL
62 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account