AI Infrastructure Engineer - GoLang, Kubernetes, Dockers, Linux - 8-12 years

Reposted 21 Days Ago
Be an Early Applicant
Bangalore, Bengaluru Urban, Karnataka, IND
In-Office
Senior level
Cloud • Information Technology • Internet of Things • Professional Services • Software
The Role
Design and implement scalable, highly available control plane components for AI workloads using Go and Python. Build Kubernetes operators/CRDs, GRPC/REST APIs, telemetry (eBPF), CI/CD, and live-upgrade mechanisms. Ensure operational simplicity, SLA/SLO monitoring, stateful application reliability, and GPU/CUDA compatibility across regions.
Summary Generated by Built In

Meet the Team
 

We are an innovation team on a mission to transform how enterprises harness AI. Operating with the agility of a startup and the focus of an incubator, we’re building a tight-knit group of AI and infrastructure experts driven by bold ideas and a shared goal: to rethink systems from the ground up and deliver breakthrough solutions that redefine what's possible — faster, leaner, and smarter.

We thrive in a fast-paced, experimentation-rich environment where new technologies aren’t just welcome — they’re expected. Here, you'll work side-by-side with seasoned engineers, architects, and thinkers to craft the kind of iconic products that can reshape industries and unlock entirely new models of operation for the enterprise.

If you're energized by the challenge of solving hard problems, love working at the edge of what's possible, and want to help shape the future of AI infrastructure — we'd love to meet you.

Impact

Cisco is seeking an experienced and innovative Control Plane Engineer to develop the control plane for the next-generation AI infrastructure. This role focuses on designing and implementing scalable, reliable, and efficient control plane components to manage AI workloads. The ideal candidate will have a strong background in microservices architecture, Kubernetes, and distributed systems, along with expertise in modern programming languages and cloud-native technologies.

As AI Control Plane Engineer your work will have a significant impact on:

Enabling seamless orchestration and management of AI workloads across distributed environments.

Improving the reliability, scalability, and performance of AI control plane services.

Ensuring operational simplicity and ease of debugging for control plane services.

Driving innovation in control plane architecture to improve efficiency and performance of AI infrastructure.

Your contributions will empower Cisco to deliver best-in-class AI infrastructure solutions and help customers scale AI workloads with confidence.

Key Responsibilities:

Design and implement control plane components using Golang AND Python,

Leverage Kubernetes (K8s) concepts, CRDs, and operator patterns (e.g., Kubebuilder) to build control plane services.

Develop scalable and highly available (HA) microservices that span across regions, ensuring reliability at scale.

Build and maintain GRPC, REST APIs, and CLI tools for seamless integration and control of AI infrastructure.

Address operational challenges of running applications as SaaS, focusing on ease of deployment and lifecycle management.

Establish and follow best practices for release management, including CI/CD pipelines and version control hygiene.

Design and implement strategies for live upgrades to minimize downtime and ensure service continuity.

Develop and implement telemetry collection mechanisms using eBPF and other tools to provide insights into system performance and health.

Define and monitor SLA/SLO metrics to ensure the reliability of control plane services.

Design and manage stateful applications, ensuring high performance and reliability of underlying databases.

Build systems with debuggability in mind, simplifying troubleshooting and remediation for operational teams.

Collaborate with other teams to ensure control plane compatibility with GPU/CUDA-related technologies.

Minimum Qualifications:

Proficiency in Golang and Python.

Strong expertise in Kubernetes (K8s), including CRDs, the operator pattern, and tools like Kubebuilder and Helm.

Experience with API design and implementation (GRPC, REST APIs, and CLI).

Proven track record in designing and scaling highly available microservices across regions. Strong understanding of distributed systems architecture and fundamentals.

Familiarity with telemetry tools and SLA/SLO monitoring for large-scale systems.

Strong debugging skills and experience building systems with easy remediation capabilities.

Passion for learning and staying updated on the latest trends in AI infrastructure and cloud-native technologies.

Bachelor’s degree+ and relevant 7+ years of Engineering work experience.

Preferred Qualifications:

Proficiency in programming languages such as C++, Golang.

Hands-on experience with eBPF for collecting insights and optimizing system performance.

Knowledge of GPU/CUDA technologies and their integration into infrastructure systems.

Knowledge of release management best practices, including versioning and rollback mechanisms.

Familiarity with SaaS operational models and challenges.

Why Cisco? 

At Cisco, we’re revolutionizing how data and infrastructure connect and protect organizations in the AI era – and beyond. We’ve been innovating fearlessly for 40 years to create solutions that power how humans and technology work together across the physical and digital worlds. These solutions provide customers with unparalleled security, visibility, and insights across the entire digital footprint.

Fueled by the depth and breadth of our technology, we experiment and create meaningful solutions. Add to that our worldwide network of doers and experts, and you’ll see that the opportunities to grow and build are limitless. We work as a team, collaborating with empathy to make really big things happen on a global scale. Because our solutions are everywhere, our impact is everywhere. 

We are Cisco, and our power starts with you. 

Skills Required

  • Proficiency in Golang and Python.
  • Strong expertise in Kubernetes, including CRDs, operator patterns, Kubebuilder, and Helm.
  • Experience with API design and implementation (gRPC, REST APIs, and CLI tools).
  • Proven track record in designing and scaling highly available microservices across regions; strong distributed systems fundamentals.
  • Familiarity with telemetry tools and SLA/SLO monitoring for large-scale systems.
  • Strong debugging skills and experience building systems with easy remediation capabilities.
  • Bachelor's degree (or higher) and relevant 7+ years of engineering work experience.
  • Hands-on experience with Docker and Linux environments.
  • Proficiency in C++.
  • Hands-on experience with eBPF for telemetry and performance optimization.
  • Knowledge of GPU/CUDA technologies and integration into infrastructure systems.
  • Knowledge of release management best practices, including versioning and rollback mechanisms.
  • Familiarity with SaaS operational models and challenges.

Cisco Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Cisco and has not been reviewed or approved by Cisco.

  • Healthcare Strength Health coverage is described as robust with multiple plan options and access to onsite/virtual LifeConnections Health Centers on major campuses. Company materials also highlight mental-health resources and comprehensive preventive care, supporting strong core medical benefits.
  • Leave & Time Off Breadth Time away includes company‑wide recharge days, a paid birthday, a year‑end shutdown, and paid Critical Time Off for emergencies. Paid volunteer days further expand opportunities to step away and recharge.
  • Parental & Family Support Policies include a global minimum for paid parental leave for primary caregivers, caregiving concierge services, and on‑site children’s learning centers in select locations. In the U.S., family‑building support is consolidated under Carrot with a defined lifetime maximum, indicating structured assistance across fertility, preservation, adoption, and surrogacy.

Cisco Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: San Jose, CA
77,500 Employees
Year Founded: 1984

What We Do

Cisco (NASDAQ: CSCO) enables people to make powerful connections--whether in business, education, philanthropy, or creativity. Cisco hardware, software, and service offerings are used to create the Internet solutions that make networks possible--providing easy access to information anywhere, at any time. Cisco was founded in 1984 by a small group of computer scientists from Stanford University. Since the company's inception, Cisco engineers have been leaders in the development of Internet Protocol (IP)-based networking technologies. Today, with more than 71,000 employees worldwide, this tradition of innovation continues with industry-leading products and solutions in the company's core development areas of routing and switching, as well as in advanced technologies such as home networking, IP telephony, optical networking, security, storage area networking, and wireless technology. In addition to its products, Cisco provides a broad range of service offerings, including technical support and advanced services. Cisco sells its products and services, both directly through its own sales force as well as through its channel partners, to large enterprises, commercial businesses, service providers, and consumers.

Similar Jobs

Shield AI Logo Shield AI

Simulation Specialist, India

Aerospace • Artificial Intelligence • Machine Learning • Robotics • Software
In-Office
Bangalore, Bengaluru Urban, Karnataka, IND

Cargill Logo Cargill

Senior Software Engineer

Food • Greentech • Logistics • Sharing Economy • Transportation • Agriculture • Industrial
In-Office
Bengaluru, Bengaluru Urban, Karnataka, IND
155000 Employees

Micron Technology Logo Micron Technology

Intern

Artificial Intelligence • Hardware • Information Technology • Machine Learning
In-Office
2 Locations
45000 Employees

Micron Technology Logo Micron Technology

Engineer

Artificial Intelligence • Hardware • Information Technology • Machine Learning
In-Office
2 Locations
45000 Employees

Similar Companies Hiring

Kepler  Thumbnail
Artificial Intelligence • Fintech • Software
New York, New York
9 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Revel.io Thumbnail
Aerospace • Hardware • Robotics • Software
US
50 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account