Sr Platform/ Infrastructure Engineer

Posted 9 Hours Ago
Be an Early Applicant
8 Locations
Remote
Senior level
Professional Services • Consulting
The Role
Design, deploy, and maintain production Kubernetes clusters and cloud-native infrastructure. Build Python automation, integrate Prometheus monitoring and Ceph storage, troubleshoot distributed systems, support platform modernization and migrations, document runbooks, and participate in on-call incident response with cross-team collaboration.
Summary Generated by Built In

We are seeking a senior Sr Platform/Infrastructure Engineer to strengthen our platform team and drive cloud-native infrastructure initiatives. This role focuses on deploying and maintaining Kubernetes services, integrating monitoring and storage platforms, and troubleshooting distributed systems to ensure resilient, scalable operations.

You will work with Python-driven tooling, Prometheus-based monitoring, Ceph-backed storage, and public cloud environments (AWS and Azure) to modernize and operate our platform. This is an opportunity to shape platform reliability and performance in a hands-on engineering role.

Responsibilities
  • Design, deploy, and maintain production Kubernetes clusters and related services.
  • Build and maintain automation and tooling using Python to support platform operations.
  • Integrate and operate Prometheus for monitoring, alerting, and observability.
  • Deploy and manage Ceph storage solutions for distributed workloads.
  • Support platform modernization initiatives and migrate services to cloud-native patterns.
  • Troubleshoot and resolve issues in distributed systems across compute, storage, and network layers.
  • Collaborate with development, SRE, and operations teams to define platform requirements and SLAs.
  • Document platform designs, runbooks, and operational procedures.
  • Participate in on-call rotations and incident response to maintain platform availability.
Requirements
  • 5+ years of experience in platform, infrastructure, or site reliability engineering roles.
  • Proven experience deploying and operating Kubernetes in production.
  • Strong Python skills for automation, tooling, and operational scripts.
  • Experience implementing and operating Prometheus-based monitoring and alerting.
  • Hands-on experience with Ceph or similar distributed storage systems.
  • Cloud experience with AWS and Azure (designing, deploying, and operating services).
  • Demonstrated ability to troubleshoot distributed systems and resolve production incidents.
  • Experience collaborating across teams to deliver platform improvements and migrations.
Nice to Have
  • Experience with OpenSearch.
  • Proficiency with Bash scripting.
  • Familiarity with Java-based services.
  • Experience with Fluent Bit for log collection.
  • Experience working with PostgreSQL.

Engagement & Logistics
  • Engagement Length: 12 months or more.
  • Time Zone: PST - 8:00 AM - 5:00 PM
  • Holiday Calendar: Client Holidays (USA – Mandatory)
  • Laptop: BYOD.
  • Overtime Required: No.

Selection process
  1. Meeting with Resilient Co. team.
  2. Technical interview
  3. Client (2 interviews - Manager + Technical panel)

Skills Required

  • 5+ years of experience in platform, infrastructure, or site reliability engineering roles.
  • Proven experience deploying and operating Kubernetes in production.
  • Strong Python skills for automation, tooling, and operational scripts.
  • Experience implementing and operating Prometheus-based monitoring and alerting.
  • Hands-on experience with Ceph or similar distributed storage systems.
  • Cloud experience with AWS and Azure (designing, deploying, and operating services).
  • Demonstrated ability to troubleshoot distributed systems and resolve production incidents.
  • Experience collaborating across teams to deliver platform improvements and migrations.
  • Experience with OpenSearch.
  • Proficiency with Bash scripting.
  • Familiarity with Java-based services.
  • Experience with Fluent Bit for log collection.
  • Experience working with PostgreSQL.
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
12 Employees
Year Founded: 2020

What We Do

ResilientCo is a professional consultancy firm specializing in all aspects of resilience, including community building, risk management, and emergency and disaster response. The company provides expert guidance in strategy development, organizational resilience, and the management of natural hazards and societal risks, aiming to enhance sustainability and the capacity of organizations and communities to withstand and recover from systemic challenges.

Similar Jobs

InterSystems Logo InterSystems

Technical Specialist

Artificial Intelligence • Big Data • Healthtech • Machine Learning • Software • Database • Analytics
Easy Apply
Remote
Chile
2100 Employees
Easy Apply
Remote
37 Locations
55 Employees
140K-178K Annually

InterSystems Logo InterSystems

Executive Assistant

Artificial Intelligence • Big Data • Healthtech • Machine Learning • Software • Database • Analytics
Easy Apply
Remote
Chile
2100 Employees

Cloudera Logo Cloudera

Staff System Engineer – Taikun APIs (C# & Go)

Artificial Intelligence • Cloud • Software • Big Data Analytics
In-Office or Remote
Santiago, Metropolitana de Santiago, CHL
3092 Employees

Similar Companies Hiring

Energy CX Thumbnail
Greentech • Professional Services • Business Intelligence • Consulting • Energy • Financial Services • Utilities
Chicago, IL
108 Employees
Northslope Thumbnail
Artificial Intelligence • Information Technology • Software • Analytics • Consulting • Generative AI
London, GB
100 Employees
Amplify Platform Thumbnail
Fintech • Financial Services • Consulting • Cloud • Business Intelligence • Big Data Analytics
Scottsdale, AZ
62 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account