102723 - Sr Platform/Infrastructure Engineer

Reposted 16 Days Ago
Be an Early Applicant
Kountríon, Trifylia, GRC
In-Office
Senior level
Information Technology • Business Intelligence • Consulting
Leading the agentic AI revolution in IT services and solutions.
The Role
Design, deploy, and operate production Kubernetes clusters; build Python automation; integrate Prometheus monitoring and Ceph storage; migrate services to cloud-native patterns on AWS/Azure; troubleshoot distributed systems; document runbooks; collaborate across teams and participate in on-call incident response.
Summary Generated by Built In
Sr Platform/Infrastructure EngineerSummary

We are seeking a senior Sr Platform/Infrastructure Engineer to strengthen our platform team and drive cloud-native infrastructure initiatives. This role focuses on deploying and maintaining Kubernetes services, integrating monitoring and storage platforms, and troubleshooting distributed systems to ensure resilient, scalable operations.

You will work with Python-driven tooling, Prometheus-based monitoring, Ceph-backed storage, and public cloud environments (AWS and Azure) to modernize and operate our platform. This is an opportunity to shape platform reliability and performance in a hands-on engineering role.

Responsibilities
  • Design, deploy, and maintain production Kubernetes clusters and related services.
  • Build and maintain automation and tooling using Python to support platform operations.
  • Integrate and operate Prometheus for monitoring, alerting, and observability.
  • Deploy and manage Ceph storage solutions for distributed workloads.
  • Support platform modernization initiatives and migrate services to cloud-native patterns.
  • Troubleshoot and resolve issues in distributed systems across compute, storage, and network layers.
  • Collaborate with development, SRE, and operations teams to define platform requirements and SLAs.
  • Document platform designs, runbooks, and operational procedures.
  • Participate in on-call rotations and incident response to maintain platform availability.
Requirements
  • 5+ years of experience in platform, infrastructure, or site reliability engineering roles.
  • Proven experience deploying and operating Kubernetes in production.
  • Strong Python skills for automation, tooling, and operational scripts.
  • Experience implementing and operating Prometheus-based monitoring and alerting.
  • Hands-on experience with Ceph or similar distributed storage systems.
  • Cloud experience with AWS and Azure (designing, deploying, and operating services).
  • Demonstrated ability to troubleshoot distributed systems and resolve production incidents.
  • Experience collaborating across teams to deliver platform improvements and migrations.
Nice to Have
  • Experience with OpenSearch.
  • Proficiency with Bash scripting.
  • Familiarity with Java-based services.
  • Experience with Fluent Bit for log collection.
  • Experience working with PostgreSQL.

Skills Required

  • 5+ years experience in platform, infrastructure, or site reliability engineering roles
  • Proven experience deploying and operating Kubernetes in production
  • Strong Python skills for automation, tooling, and operational scripts
  • Experience implementing and operating Prometheus-based monitoring and alerting
  • Hands-on experience with Ceph or similar distributed storage systems
  • Cloud experience with AWS and Azure (designing, deploying, and operating services)
  • Demonstrated ability to troubleshoot distributed systems and resolve production incidents
  • Experience collaborating across teams to deliver platform improvements and migrations
  • Experience with OpenSearch
  • Proficiency with Bash scripting
  • Familiarity with Java-based services
  • Experience with Fluent Bit for log collection
  • Experience working with PostgreSQL
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: San Francisco, CA
2,263 Employees
Year Founded: 2008

What We Do

Taller is the enterprise accelerator for digital transformation, expertly orchestrating hybrid teams of senior specialists and AI agents under trusted oversight — the "humans in the loop" delivering unparalleled speed, scale, and strategic impact. Subscribe to our monthly newsletter covering the latest breakthroughs in enterprise AI: https://hubs.ly/Q03tqbNy0

Similar Jobs

Pfizer Logo Pfizer

Sustainability Senior Manager

Artificial Intelligence • Healthtech • Machine Learning • Natural Language Processing • Biotech • Pharmaceutical
Remote or Hybrid
30 Locations
121990 Employees
112K-207K Annually

Pfizer Logo Pfizer

Sr. Director, Product Intelligence & Marketing Lead

Artificial Intelligence • Healthtech • Machine Learning • Natural Language Processing • Biotech • Pharmaceutical
In-Office or Remote
30 Locations
121990 Employees
215K-358K Annually

Pfizer Logo Pfizer

Director, AI Platform Product Management

Artificial Intelligence • Healthtech • Machine Learning • Natural Language Processing • Biotech • Pharmaceutical
In-Office or Remote
30 Locations
121990 Employees
163K-272K Annually

Pfizer Logo Pfizer

Director, Build Engineer

Artificial Intelligence • Healthtech • Machine Learning • Natural Language Processing • Biotech • Pharmaceutical
In-Office or Remote
30 Locations
121990 Employees
177K-294K Annually

Similar Companies Hiring

Compa Thumbnail
Artificial Intelligence • HR Tech • Software • Business Intelligence
Irvine, California
75 Employees
NODA AI Thumbnail
Artificial Intelligence • Information Technology • Software • Cybersecurity
Sydney, AU
54 Employees
Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account