Mirantis, an IREN company, is the Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI, machine learning, and data-intensive applications. By combining open source innovation with deep expertise in Kubernetes orchestration, Mirantis empowers platform engineering teams to deliver composable, production-ready developer platforms across any environment—on-premises, in the cloud, at the edge, or in sovereign data centers. As enterprises navigate the growing complexity of AI-driven workloads, Mirantis delivers the automation, GPU orchestration, and policy-driven control needed to manage infrastructure with confidence and agility. Committed to open standards and freedom from lock-in, Mirantis ensures that customers retain full control of their infrastructure strategy. https://www.mirantis.com/
Job DescriptionWe are looking for a Senior Data Platform Engineer to run the PostgreSQL and Apache Kafka platform behind k0rdent-ai — our multi-tenant control plane for enterprise GPU infrastructure. Every cluster provisioned, every GPU-hour consumed, and every tenant action is recorded on this platform, so it must be reliable, secure, and recoverable.
This is a hands-on DevOps / SRE role for stateful systems. You will deploy and operate PostgreSQL (CloudNativePG) and Kafka (Strimzi) on Kubernetes using operators, Helm, and GitOps — across cloud and bare-metal clusters, in a global control plane and multiple regions. You will own automation, upgrades, observability, backups, and disaster recovery, and give product teams self-service access to databases and topics through code.
Main Responsibilities
- Run PostgreSQL and Kafka on Kubernetes: Deploy, upgrade, scale, and operate CloudNativePG and Strimzi clusters with Helm charts and operator custom resources, across cloud and bare-metal Kubernetes.
- GitOps & Infrastructure as Code: Manage everything declaratively with Argo CD (or Flux), Helm, and Terraform; build CI/CD pipelines for database and Kafka changes, schema migrations, and operator upgrades.
- High Availability & Disaster Recovery: Design and operate high availability within each region and cross-region disaster recovery — PostgreSQL replica clusters, backups and point-in-time recovery, Kafka MirrorMaker 2 — with defined RPO/RTO targets and regular failover drills.
- Self-Service Platform: Give product teams self-service provisioning of databases, users, topics, and ACLs through code, with secure defaults and guardrails.
- Observability & Operations: Build monitoring and alerting (Prometheus, Grafana, OpenTelemetry) for replication lag, backups, consumer lag, and capacity; participate in on-call and lead incident reviews.
- Security & Multi-Tenancy: Implement TLS/mTLS, secrets management (Vault/OpenBao), network policies, Kafka ACLs, and per-tenant isolation; integrate database and Kafka access with our identity platform (Keycloak / OIDC).
- Data Movement & CDC: Operate the outbox and change-data-capture pipelines (Debezium, Kafka Connect) that keep databases and event streams consistent between the global control plane and regions.
- Automation & Tooling: Write automation and small services in Go or Python to remove manual work, and partner with application teams on schema standards and performance troubleshooting.
Must Have
- Experience: 8+ years in DevOps, SRE, platform, or infrastructure engineering, including 3+ years running stateful systems (databases or message streaming) in production.
- Kubernetes: Strong hands-on Kubernetes operations — StatefulSets, persistent volumes and storage classes, pod disruption budgets, scheduling and affinity, network policies, and cluster upgrades — on cloud and/or bare-metal clusters.
- Operators & Helm: Production experience running PostgreSQL and/or Kafka through Kubernetes operators (CloudNativePG, Zalando, or Crunchy PGO; Strimzi, Confluent for Kubernetes, or similar), and authoring and maintaining Helm charts. Depth with one operator matters more than the specific product.
- GitOps & IaC: Day-to-day use of Argo CD or Flux, Terraform, and CI/CD pipelines (GitHub Actions, GitLab CI, or similar) in a declarative, review-driven workflow.
- PostgreSQL Operations: Practical PostgreSQL administration — replication and failover, backup and point-in-time recovery, connection pooling (PgBouncer), version upgrades, and performance troubleshooting.
- Kafka Operations: Running Kafka in production — brokers, topics and partitions, replication, consumer groups, monitoring, and capacity planning.
- Reliability & DR: Designing and testing high availability and disaster recovery for stateful systems, with defined RPO/RTO.
- Automation: Strong scripting and programming in Go (preferred) or Python, plus Bash.
Nice to Have
- Cross-region replication (CloudNativePG replica clusters, Kafka MirrorMaker 2) and global/regional multi-site architectures.
- Debezium, Kafka Connect, and transactional outbox / CDC patterns.
- Identity integration for data platforms (Keycloak, OIDC, LDAP/Active Directory).
- Observability stacks (Prometheus, Grafana, OpenTelemetry, VictoriaMetrics).
- Secrets and security tooling (Vault/OpenBao, cert-manager, mTLS).
- Bare-metal infrastructure, Ceph or object storage, or GPU/AI infrastructure.
- Contributions to CloudNativePG, Strimzi, or related open-source projects.
- Experience in SOC 2 or ISO 27001 environments.
Education and Experience
Bachelor's degree in Computer Science or a related field, or equivalent practical experience.
What does Mirantis offer you?
- Work with an established Silicon Valley leader in the cloud infrastructure industry;
- Work with exceptionally passionate, talented and engaging colleagues, helping Fortune 500 and Global 2000 customers implement next-generation cloud technologies;
- Be a part of cutting-edge, open-source innovation;
- Thrive in the high-energy environment of a young company where openness, collaboration, risk-taking, and continuous growth are valued;
- Professional development and training;
- Attend conferences and working groups;
- Company outings, happy hours, hackathons, and tech talks;
- Receive a competitive compensation package with a strong benefits plan.
We are a Leader for Container Management in G2 (#2 after AWS)!
Skills Required
- 8+ years of experience in DevOps, SRE, platform, or infrastructure engineering
- 3+ years running stateful systems such as databases or message streaming platforms in production
- Strong hands-on Kubernetes operations experience, including StatefulSets, persistent volumes, storage classes, pod disruption budgets, scheduling, affinity, network policies, and cluster upgrades
- Production experience operating PostgreSQL and/or Kafka through Kubernetes operators such as CloudNativePG, Zalando, Crunchy PGO, Strimzi, or Confluent for Kubernetes
- Experience authoring and maintaining Helm charts
- Day-to-day experience with Argo CD or Flux, Terraform, and CI/CD pipelines
- Practical PostgreSQL administration, including replication, failover, backups, point-in-time recovery, PgBouncer, version upgrades, and performance troubleshooting
- Production Kafka operations experience, including brokers, topics, partitions, replication, consumer groups, monitoring, and capacity planning
- Experience designing and testing high availability and disaster recovery for stateful systems with defined RPO and RTO
- Strong programming and scripting skills in Go or Python, plus Bash
- Bachelor's degree in Computer Science or a related field, or equivalent practical experience
- Experience with cross-region replication and multi-site architectures
- Experience with Debezium, Kafka Connect, and transactional outbox or CDC patterns
- Experience integrating data platforms with Keycloak, OIDC, LDAP, or Active Directory
- Experience with Prometheus, Grafana, OpenTelemetry, or VictoriaMetrics
- Experience with Vault, OpenBao, cert-manager, or mTLS
- Experience with bare-metal infrastructure, Ceph, object storage, or GPU/AI infrastructure
- Contributions to CloudNativePG, Strimzi, or related open-source projects
- Experience in SOC 2 or ISO 27001 environments
Mirantis Compensation & Benefits Highlights
The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Mirantis and has not been reviewed or approved by Mirantis.
-
Fair & Transparent Compensation — Pay is considered fair-to-good and competitive for certain technical and senior roles, and compensation is often viewed as a relative strength. Overall packages can feel solid in context, even if not universally top-of-market.
-
Healthcare Strength — Health coverage is described as comprehensive, including medical, dental, and vision, with employer-verified listings and mentions of strong coverage in practice. Core offerings align with what tech employees typically expect and are sometimes called out as a highlight.
-
Leave & Time Off Breadth — Policies include PTO, sick time, paid holidays, bereavement leave, and generous parental leave. These provisions are frequently characterized as generous and supportive of work-life balance.
Mirantis Insights
Similar Jobs
What We Do
We are dedicated to helping organizations increase developer productivity and ship code faster on public and private clouds. We provide a ZeroOps experience to remove the stress of managing cloud native infrastructure by combining software and automation tools with our cloud native expertise to deliver the industry's leading secure cloud platforms. Our capabilities allow us to provide a secure and reliable cloud native platform that includes validated FIPS-140-2 Encryption and DISA STIG ready capabilities. Who do we serve? We serve a wide range of industries, building on our extensive customer experience to provide distinct value in specific verticals including Financial Services, Government & Education, Healthcare, Manufacturing, and Telecommunications. Mirantis serves many of the world’s leading enterprises, including Adobe, DocuSign, Inmarsat, PayPal, Reliance Jio, Societe Generale, Splunk, and S&P Global. Learn more at www.mirantis.com.








