Manager, Platform Operations & Site Reliability Engineer

Posted Yesterday
Be an Early Applicant
2 Locations
In-Office
Senior level
Artificial Intelligence • Healthtech • Machine Learning • Natural Language Processing • Biotech • Pharmaceutical
We’re in relentless pursuit of breakthroughs that change patients’ lives.
The Role
Lead platform reliability and operations for enterprise AI agent platforms: build and maintain observability, IaC, automation, incident response, SLOs/SLAs, cost/performance optimization, and mentor engineers while collaborating with architecture, security, and product teams.
Summary Generated by Built In

ROLE SUMMARY

Pfizer’s purpose, breakthroughs that change patients’ lives, is rooted in being science driven and a patient focused company.  Digital and Technology is the driving force of data and AI innovation at Pfizer. 

 

As part of the Data and AI Platforms organization, you will join a team of Platform Engineers responsible for building and operating the enterprise pro‑code AI Agent Platforms. We are expanding our forward-thinking Platform Engineering team, focused on delivering secure, scalable, and resilient cloud-native infrastructure.

In the role of Site Reliability/Operations Engineer, you will play a critical part in ensuring the reliability, performance, and operational excellence of our platforms. This position is well suited to a high-caliber, well-rounded engineer who thrives in dynamic environments, takes initiative, and enjoys solving complex problems across infrastructure, automation, and observability. You will be part of a collaborative and inclusive team that values curiosity, continuous learning, and shared success, with clear goals, strong mentorship, and a culture designed to help you succeed while taking ownership of meaningful challenges.

You will contribute to the design and development of a reliable, transparent, interoperable, and sustainable enterprise-grade observability platform, with a strong focus on security, governance, cost management, and lifecycle considerations. In this role, you will help shape the platform roadmap, collaborate closely with architecture, security, and product teams, and drive the delivery of a best-in-class developer experience. You will also support and mentor engineers, promoting engineering excellence and best practices across the team.

Success in this role will require building trusted partnerships with technology vendors, delivery partners, and internal stakeholders across the organization. You will apply both emerging and established technologies to enhance analytics and observability capabilities, supporting the broader agentic platform strategy and enabling scalable, high-performing, and well-governed systems.

You will start your day by reviewing system health and platform metrics to ensure everything is operating as expected. From there, you will collaborate with your platform engineer and operations teammates to prioritize work, whether that involves deploying infrastructure, refining automation, or addressing new technical challenges. Some days will see you working deeply within Terraform modules or tuning infrastructure, while others will focus on troubleshooting unexpected issues or supporting colleagues with complex problems. The role provides a balance of focused individual work and collaborative problem-solving, all within a fast-paced yet supportive environment where your contributions directly influence platform reliability, scalability, and security.

 

In this role, you will be responsible for establishing and continuously improving observability and operational practices across the platform for both internal platform engineers and external end-users of the platform services. This includes building robust monitoring, logging, tracing, alerting, and incident response capabilities, alongside defining and maintaining SLOs and SLAs. You will proactively optimize performance, availability, and cost efficiency through techniques such as autoscaling, right-sizing, and resource planning.

 

Key responsibilities include:

  • Monitoring cloud infrastructure and responding to alerts, working closely with senior engineers to ensure stability and rapid resolution
  • Deploying and maintaining infrastructure as code and automation tooling (e.g., Terraform), while identifying opportunities to improve efficiency and scalability
  • Maintaining and enhancing observability systems such as Langfuse, Langsmith, and Grafana, incorporating feedback from engineering teams
  • Participating in incident response, root cause analysis, and remediation efforts to drive continuous improvement
  • Collaborating with team members to evolve and strengthen CI/CD pipelines and delivery processes
  • Contributing to documentation and improving operational processes to support platform maturity
  • Proactively developing your skills while applying security and compliance best practices in day-to-day activities

 

BASIC QUALIFICATIONS

 

  • Education: Bachelor’s degree in a relevant field (e.g., Computer Science, Data Science, Engineering, or related discipline)
  • 5+ years of experience in Site reliability, Operations, or Infrastructure engineering
  • Demonstrated experience with telemetry and observability to instrument, generate, collect, and export telemetry data (metrics, logs, and traces) to help analyze agent’s performance and behavior (e.g., OpenTelemetry, Grafana, Langfuse).
  • Demonstrated handson expertise with Microsoft Azure and/or Google GCP services
  • Familiarity and experience with Terraform and GitHub Actions
  • Understanding of Kubernetes, Docker, and container orchestration
  • Good scripting skills (e.g., Bash, Python, Typescript)
  • Familiarity and experience with Linux/Unix system administration
  • Familiarity with networking, security, and database administration
  • Strong problem-solving skills and eagerness to learn in a collaborative environment
  • Fluent in English; capable of clear technical communication across scientific and engineering disciplines
  • Collaborate with User Success team to maintain key informational content, drive platform adoption, and champion a strong user community.

 

PREFERRED QUALIFICATIONS

 

  • Experience leading distributed platform Operations & Support teams.
  • Experience working in regulated environments or with compliance frameworks (e.g., GxP, SOC2, HIPAA)
  • Experience working in team-based environments, either professionally or academically
  • Certifications: Microsoft Certified, Google Cloud Certified, HashiCorp Terraform Associate.

 

PHYSICAL/MENTAL REQUIREMENTS

  • Ability to perform complex data analysis, architectural and code reviews, incident triage.

 

NON-STANDARD WORK SCHEDULE, TRAVEL OR ENVIRONMENT REQUIREMENTS

  • Occasional business travel to Pfizer sites and service partners

 

Pfizer is an equal opportunity employer and complies with all applicable equal employment opportunity legislation in each jurisdiction in which it operates.

To learn more about acceptable and prohibited uses of AI during the recruitment process, please review our candidate AI-use guidelines available on Pfizer Careers.

Information & Business Tech

Skills Required

  • Bachelor's degree in Computer Science, Data Science, Engineering, or related discipline
  • 5+ years experience in site reliability, operations, or infrastructure engineering
  • Experience with telemetry and observability (OpenTelemetry, Grafana, Langfuse, Langsmith)
  • Hands-on expertise with Microsoft Azure and/or Google Cloud Platform (GCP)
  • Familiarity and experience with Terraform
  • Familiarity and experience with GitHub Actions
  • Understanding of Kubernetes, Docker, and container orchestration
  • Scripting skills (Bash, Python, TypeScript)
  • Familiarity and experience with Linux/Unix system administration
  • Familiarity with networking, security, and database administration
  • Fluent in English with clear technical communication skills
  • Experience leading distributed platform operations & support teams
  • Experience in regulated environments or with compliance frameworks (GxP, SOC2, HIPAA)
  • Experience working in team-based environments
  • Certifications: Microsoft Certified, Google Cloud Certified, HashiCorp Terraform Associate

What the Team is Saying

Daniel
Anna
Esteban
Pfizer

Pfizer Compensation & Benefits Highlights

  • Healthcare Strength Multiple U.S. medical plan options include telehealth, comprehensive mental‑health support, fertility/family‑building benefits, transgender‑inclusive coverage, and certain Pfizer medications at no cost. A Wellbeing Wallet and wellness resources broaden the health and wellbeing offering.
  • Retirement Support A 401(k) with company matching is paired with an additional Pfizer Retirement Savings Contribution, alongside company‑paid life and disability insurance. One‑on‑one financial planning support is provided through Fidelity.
  • Leave & Time Off Breadth Paid time off spans vacation, holidays, and personal days, with additional caregiver and medical leave. U.S. parental leave commonly includes 12 weeks paid with options for additional unpaid bonding time and a return‑to‑work transition.

Pfizer Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: New York, NY
121,990 Employees
Year Founded: 1848

What We Do

Our purpose ensures that patients remain at the center of all we do. We live our purpose by sourcing the best science in the world; partnering with others in the healthcare system to improve access to our medicines; using digital technologies to enhance our drug discovery and development, as well as patient outcomes; and leading the conversation to advocate for pro-innovation/pro-patient policies.

Why Work With Us

We are the inventors, the problem solvers, the big thinkers — those who surmount any hurdle to deliver breakthrough medicines to the people who are counting on them the most.

Gallery

Gallery
Gallery
Gallery
Gallery
Gallery

Pfizer Offices

Hybrid Workspace

Employees engage in a combination of remote and on-site work.

Typical time on-site: 2.5 days a week
Company Office Image
HQHudson Yards
Provincia de Buenos Aires
Andover, MA
Athens, GR
Chennai, IN
Collegeville, PA
Durham, NC
Groton, CT
Madison, NJ
Madrid, ES
Mumbai, Maharashtra
Rochester, MI
San Diego, CA
Seattle, WA
Company Office Image
Tampa, FL
Center for Digital Innovation
Learn more

Similar Jobs

Pfizer Logo Pfizer

Manager, MAPP Developer

Artificial Intelligence • Healthtech • Machine Learning • Natural Language Processing • Biotech • Pharmaceutical
In-Office
Chennai, Tamil Nadu, IND
121990 Employees

Pfizer Logo Pfizer

Manager, Operations Lead, Veeva Clinical

Artificial Intelligence • Healthtech • Machine Learning • Natural Language Processing • Biotech • Pharmaceutical
In-Office
3 Locations
121990 Employees

Pfizer Logo Pfizer

Associate Manager - Indirect Taxation

Artificial Intelligence • Healthtech • Machine Learning • Natural Language Processing • Biotech • Pharmaceutical
In-Office
Chennai, Tamil Nadu, IND
121990 Employees

Pfizer Logo Pfizer

Associate Data Manager, Clinical Data Sciences

Artificial Intelligence • Healthtech • Machine Learning • Natural Language Processing • Biotech • Pharmaceutical
In-Office
Chennai, Tamil Nadu, IND
121990 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account