Senior Kubernetes Engineer (SME)

Posted 3 Days Ago
Be an Early Applicant
Mumbai, Maharashtra, IND
In-Office
Senior level
Artificial Intelligence • Cloud • Infrastructure as a Service (IaaS)
The Role
Own production Kubernetes infrastructure across bare-metal, cloud, and hybrid environments. Design and operate multi-cluster, multi-tenant platforms, managing networking, DNS, storage, service mesh, container runtimes, operating systems, security, GitOps, infrastructure as code, backups, and disaster recovery. Provide on-call incident response, root-cause resolution, documentation, and operational leadership for reliable platform delivery.
Summary Generated by Built In

About the Role:

We are building and running mission-critical production infrastructure on Kubernetes. As a Senior Kubernetes Engineer, you will own the full stack - from the underlying OS and container runtime through networking, storage, and the cluster control plane itself. You will engage across the full lifecycle: architecture, deployment, hardening, and day-2 operations for multi-cluster, multi-tenant environments. This is hands-on infrastructure work with direct ownership of production reliability.

What you will be doing:

      Design, deploy, and operate production-grade Kubernetes clusters across bare-metal, cloud, and hybrid environments.

      Manage cluster lifecycle end-to-end: provisioning, version upgrades, patching, scaling, and capacity planning

      Architect and run multi-cluster/multi-tenant setups using Kamaji, Rancher (hosted control planes), and vCluster (virtual clusters)

      Configure and troubleshoot CNI plugins (Calico, Cilium) — pod networking, network policies, BGP, and eBPF dataplanes.

      Own cluster DNS (CoreDNS) configuration, service discovery, and resolution troubleshooting.

      Manage container runtime (containerd, CRI-O) and underlying OS: Linux tuning, kernel/sysctl parameters, systemd, cgroups

      Deploy and operate service mesh (Istio, Cilium mesh) for traffic management

      Manage persistent storage: PV/PVC, StorageClasses, CSI drivers (Rook/Ceph, Longhorn, cloud-native CSI)

      Own the network stack: ingress controllers, load balancing, MetalLB/BGP, firewalling, and network troubleshooting

      Implement GitOps and Infrastructure-as-Code (ArgoCD/FluxCD, Terraform, Helm) for cluster and workload delivery

      Harden clusters: RBAC, Pod Security Standards, network policies, secrets management (Vault, Sealed Secrets), image scanning

      Own backup and disaster recovery (Velero, etcd snapshotting) and run DR drills

      Provide on-call production support: monitor cluster health, troubleshoot incidents, and drive root-cause resolution

      Create comprehensive documentation, runbooks, and knowledge bases for operational continuity and knowledge transfer

What we need to see:

          Core Kubernetes & Infrastructure (5+ years)

          Deep expertise in Kubernetes architecture: control plane, etcd, kube-apiserver, scheduler, controller-manager, kubelet

          Proven experience designing, deploying, and troubleshooting production clusters at scale

          Hands-on with multi-cluster/multi-tenant tooling — Rancher, Kamaji, and/or vCluster (strongly preferred)

          CNI expertise: Calico, Cilium — network policy design, BGP, VXLAN/IPIP encapsulation, eBPF

          Container runtime internals: containerd, CRI-O, runc — configuration and troubleshooting

          Strong Linux systems administration: kernel tuning, systemd, cgroups/namespaces, sysctl, package/OS lifecycle management

          Storage: PV/PVC, StorageClasses, CSI drivers, Ceph/Rook, Longhorn, NFS

          Service mesh experience: Istio, Linkerd, or Cilium service mesh

          CoreDNS configuration, custom resolvers, and DNS troubleshooting in cluster environments

          Networking depth: ingress controllers (NGINX, Envoy, Traefik), load balancing, MetalLB, and diagnostic tooling (tcpdump, iptables/nftables, conntrack)

Automation & Infrastructure-as-Code:

          Helm chart authoring and lifecycle management

          GitOps workflows: ArgoCD or FluxCD

          IaC and configuration management: Terraform, Ansible

          CI/CD pipeline integration for cluster and application delivery

          Scripting proficiency: Bash and Python (Go a plus)

Ways to stand out from the rest:

          Kubernetes certifications: CKA, CKAD, CKS

          Production experience with Rancher, Kamaji, and vCluster together (fleet/hosted-control-plane management)

          Multi-cloud Kubernetes: EKS, AKS, GKE, and bare-metal

          Experience with GPU-enabled clusters and AI/ML workload scheduling '

          Familiarity with AI/LLM serving stacks on Kubernetes — vLLM, KServe, Triton Inference Server, GPU operator/device plugin, MIG partitioning

          General AI infrastructure knowledge: model serving patterns, inference autoscaling, GPU scheduling constraints

          Contributions to CNCF projects or an active open-source/GitHub presence

Minimum Qualifications:

          Bachelor’s degree in computer science, Electrical/Computer Engineering, or related field (or equivalent industry experience)

          5+ years of hands-on production Kubernetes experience

          Demonstrated ownership of CNI, DNS, storage, and networking within Kubernetes environments

          Solid grounding in containerd/OS-level troubleshooting

          Production on-call and incident-response experience.

Soft Skills:

          Strong problem-solving and debugging abilities under production pressure

          Ownership mindset with accountability for platform reliability and stability

          Cross-functional collaboration with application, security, operations and platform teams

          Proactive approach to continuous learning and staying current with the Kubernetes/CNCF ecosystem

          Strong documentation and communication skills, with ability to defend design decisions in peer reviews

 

Skills Required

  • 5+ years of hands-on production Kubernetes experience
  • Bachelor's degree in computer science, Electrical/Computer Engineering, or related field, or equivalent industry experience
  • Deep expertise in Kubernetes architecture, including the control plane, etcd, kube-apiserver, scheduler, controller-manager, and kubelet
  • Experience designing, deploying, and troubleshooting production Kubernetes clusters at scale
  • Demonstrated ownership of CNI, DNS, storage, and networking within Kubernetes environments
  • Expertise with Calico or Cilium, including network policies, BGP, encapsulation, and eBPF
  • Container runtime and OS-level troubleshooting with containerd, CRI-O, or runc
  • Strong Linux systems administration, including kernel tuning, systemd, cgroups, namespaces, sysctl, and OS lifecycle management
  • Experience with Kubernetes storage, including PV/PVC, StorageClasses, CSI drivers, Ceph/Rook, Longhorn, or NFS
  • Production on-call and incident-response experience
  • Experience with Rancher, Kamaji, and/or vCluster
  • Service mesh experience with Istio, Linkerd, or Cilium service mesh
  • CoreDNS configuration and Kubernetes DNS troubleshooting
  • Networking experience with ingress controllers, load balancing, MetalLB, and diagnostic tools
  • Helm chart authoring and lifecycle management
  • GitOps workflows using ArgoCD or FluxCD
  • Infrastructure as code and configuration management using Terraform and Ansible
  • CI/CD pipeline integration for cluster and application delivery
  • Scripting proficiency in Bash and Python
  • Kubernetes certifications such as CKA, CKAD, or CKS
  • Multi-cloud Kubernetes experience with EKS, AKS, GKE, and bare-metal environments
  • Experience with GPU-enabled clusters and AI/ML workload scheduling
  • Familiarity with AI/LLM serving stacks on Kubernetes
  • Contributions to CNCF projects or an active open-source/GitHub presence
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
141 Employees
Year Founded: 2023

What We Do

Neysa is an India-based AI acceleration cloud provider focused on democratizing enterprise AI adoption. Co-founded by Sharad Sanghi and Anindya Das, it combines cloud infrastructure, AI systems, and cybersecurity expertise. Its flagship Neysa Velocis platform supports AI training, fine-tuning, inference, compute orchestration, security, and observability, helping organizations across high-growth markets such as India and beyond deploy and scale AI more quickly, safely, and cost-effectively.

Similar Jobs

Mastercard Logo Mastercard

Lead Product Manager

Blockchain • Fintech • Payments • Consulting • Cryptocurrency • Cybersecurity • Quantum Computing
Hybrid
Pune, Maharashtra, IND
38800 Employees

Coursera + Udemy  Logo Coursera + Udemy

Content Marketing Manager

Artificial Intelligence • Consumer Web • Edtech • Enterprise Web • HR Tech • Social Impact • Generative AI
Remote or Hybrid
India
1500 Employees
106K-143K Annually

Mastercard Logo Mastercard

Software Engineer

Blockchain • Fintech • Payments • Consulting • Cryptocurrency • Cybersecurity • Quantum Computing
Hybrid
Pune, Maharashtra, IND
38800 Employees

Mastercard Logo Mastercard

Senior Software Engineer

Blockchain • Fintech • Payments • Consulting • Cryptocurrency • Cybersecurity • Quantum Computing
Hybrid
Pune, Maharashtra, IND
38800 Employees

Similar Companies Hiring

Kepler  Thumbnail
Artificial Intelligence • Fintech • Software
New York, New York
9 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software • Productivity
US
15 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account