Senior Ceph/Rook-Ceph + Kubernetes Storage Engineer

Posted 27 Days Ago
Be an Early Applicant
6 Locations
Remote
Senior level
Information Technology • Software
The Role
Senior storage engineer responsible for operating and troubleshooting production Ceph and Rook-Ceph clusters on Kubernetes, including capacity planning, performance tuning, upgrades, CSI integration, Linux-level troubleshooting, automation (Ansible), Helm, and monitoring with Prometheus/Grafana. Investigates storage performance, failures, replication/CRUSH topology, and ensures high availability at scale while communicating with stakeholders.
Summary Generated by Built In

We’re currently looking for a Senior Ceph/Rook-Ceph + Kubernetes Storage Engineer to join a long-term project. The role focuses on production Ceph environments, Kubernetes storage, and infrastructure troubleshooting.


Requirements
  • 5+ years in DevOps, SRE, infrastructure or storage engineering
  • Deep hands-on production Ceph experience: OSD, MON, MGR, PGs, recovery, backfill, capacity planning, performance tuning, scaling and high availability
  • Experience building, operating, upgrading and troubleshooting production Ceph clusters
  • Hands-on Rook-Ceph in Kubernetes, preferably business-critical production environments
  • Strong Kubernetes knowledge, including CSI/storage integration
  • Linux administration and troubleshooting at OS/hardware level
  • Automation experience with Ansible
  • Kubernetes tooling such as Helm
  • Monitoring/observability with Prometheus and Grafana
  • Experience investigating storage performance, latency, disk failures, network bottlenecks and recovery issues
  • Strong understanding of failure domains, CRUSH topology, replication and storage architecture
  • Good English communication skills

Highly valuable but not mandatory:

  • OpenStack experience, especially Ceph integration with Cinder, Glance, Nova and RBD
  • Petabyte-scale Ceph environments
  • Bare-metal infrastructure
  • Large production Rook-Ceph clusters
  • CephFS and RGW/S3 experience
  • Customer-facing troubleshooting/support experience

Skills Required

  • 5+ years in DevOps, SRE, infrastructure or storage engineering
  • Deep hands-on production Ceph experience (OSD, MON, MGR, PGs, recovery, backfill, capacity planning, performance tuning, scaling, high availability)
  • Experience building, operating, upgrading and troubleshooting production Ceph clusters
  • Hands-on Rook-Ceph in Kubernetes, preferably in production
  • Strong Kubernetes knowledge including CSI/storage integration
  • Linux administration and troubleshooting at OS/hardware level
  • Automation experience with Ansible
  • Kubernetes tooling such as Helm
  • Monitoring/observability with Prometheus and Grafana
  • Experience investigating storage performance, latency, disk failures, network bottlenecks and recovery issues
  • Strong understanding of failure domains, CRUSH topology, replication and storage architecture
  • Good English communication skills
  • OpenStack experience, especially Ceph integration with Cinder, Glance, Nova and RBD
  • Experience with petabyte-scale Ceph environments
  • Bare-metal infrastructure experience
  • Experience with large production Rook-Ceph clusters
  • CephFS and RGW/S3 experience
  • Customer-facing troubleshooting/support experience
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Kadima
294 Employees

What We Do

Globaldev Group provides end-to-end software development services and builds skillful teams of specialists to fuel your business’ growth. Leveraging our extensive 12-year background and deep industry knowledge, we have established ourselves as a reliable partner for numerous enterprises, SMEs, and startups. Globaldev Group is an ISO/IEC 27001 certified company. ▪️️ Global Teams ▪️ Globaldev offers R&D team extensions to both emerging technology companies and global corporations. Empowered by our global talent pool, we excel at connecting our clients with high-skilled specialists. Hiring, onboarding, and continuous management are all included in our comprehensive set of services, enabling our clients to quickly scale their teams and achieve their product development and implementation goals. ▪️ Global Engineering ▪️ - Obtain a custom-designed solution from scratch - Scale or enhance your current software infrastructure - Stay ahead with the latest cutting-edge technologies - Access a comprehensive range of services from a single vendor As a full-stack development company, we are equipped to fulfill all your business requirements. Whether it's product discovery, proof of concept, UX design, or development, Globaldev serves as your all-in-one solution provider.

Similar Jobs

Pfizer Logo Pfizer

Director R&D EHS Program Lead

Artificial Intelligence • Healthtech • Machine Learning • Natural Language Processing • Biotech • Pharmaceutical
In-Office or Remote
36 Locations
121990 Employees
177K-294K Annually

SailPoint Logo SailPoint

Consultant

Artificial Intelligence • Cloud • Sales • Security • Software • Cybersecurity • Data Privacy
Remote or Hybrid
2 Locations
2461 Employees

Mondelēz International Logo Mondelēz International

Buying Channel Enablement Lead

Big Data • Food • Hardware • Machine Learning • Retail • Automation • Manufacturing
Remote or Hybrid
5 Locations
90000 Employees
3K-3K Annually

Mondelēz International Logo Mondelēz International

o9 Change Readiness Lead

Big Data • Food • Hardware • Machine Learning • Retail • Automation • Manufacturing
Remote or Hybrid
11 Locations
90000 Employees

Similar Companies Hiring

Kepler  Thumbnail
Artificial Intelligence • Fintech • Software
New York, New York
9 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Revel.io Thumbnail
Aerospace • Hardware • Robotics • Software
US
50 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account