The Collaboration Technology Group is redefining the future of teamwork, building services that connect people effortlessly across devices, locations and time zones.
Our team builds, runs and continuously improves the platform services behind Cisco’s collaboration products, operating at global scale across numerous datacentres. We’re a passionate, collaborative team focused on reliability, innovation and engineering excellence.
Your impactAs a Technical Leader, you will drive the architectural vision and implementation of our next-generation AI-powered Production Intelligence platform. You will combine Site Reliability Engineering practices with modern agentic AI (reusable Skills, Model Context Protocol (MCP), and LLM tooling) to transform how engineering and leadership teams monitor, diagnose, and auto-remediate global SaaS infrastructure.
What you'll do- Technical Leadership & Architecture: Define the technical roadmap and architecture for AI-assisted observability, automated incident response, and self-healing cloud infrastructure.
- Agentic Workflows & Tooling: Design and build production-grade AI agents, MCP tool integrations, and deterministic evaluation pipelines for automated operational decision support.
- Telemetry & Insights: Architect ingestion and correlation pipelines across distributed logs, metrics, OpenTelemetry traces, change events, and runbooks to accelerate Mean Time to Detection (MTTD) and Resolution (MTTR).
- Safe Production Automation: Develop proactive anomaly detection and Human-in-the-Loop (HITL) remediation workflows with rigorous safety, security, and quality guardrails.
- Reliability & Scalability Engineering: Partner with application and infrastructure teams to define SLIs/SLOs, handle error budgets, and lead deep-dive post-incident reviews (PIRs).
- Mentorship & Collaboration: Mentor senior and mid-level engineers, establish engineering best practices, and drive alignment across global development and operations teams.
- You’ll manage priorities and deadlines, communicate progress clearly and work across teams to turn production needs into reliable software and AI-assisted capabilities.
- Bachelor’s degree + 8 years of related experience, Master’s + 6 years, or PhD + 3 years in Computer Science, Software Engineering, or a related technical field.
- Proven record as a Technical Lead or Lead SRE/Software Engineer delivering distributed, high-availability SaaS platforms at scale.
- Strong proficiency in Python, Go, Java, or C++ with experience designing microservices, APIs, and production automation.
- Deep experience with Kubernetes, Docker, and container orchestration in large-scale multi-cluster environments.
- Proven background in SRE practices: SLI/SLO design, observability platforms (metrics/logs/traces), incident management, and automated RCA.
- AI & Agentic Systems: Hands-on experience building LLM pipelines, AI Agents, Model Context Protocol (MCP) servers/clients, RAG architectures, and evaluation frameworks.
- Observability & Telemetry: Experience with OpenTelemetry (OTel), Prometheus, Grafana, Splunk, ThousandEyes, or distributed tracing systems.
- Cloud & Infrastructure: Expertise in public cloud providers (AWS, GCP, Azure), Terraform/IaC, and GitOps/CI/CD pipelines (Jenkins, GitHub Actions).
- Safe Automation & Guardrails: Experience implementing responsible AI guardrails, deterministic fallback logic, and policy-driven remediation engines.
- Data & Messaging: Experience with streaming and data platforms (Kafka, Redis, PostgreSQL, Elasticsearch/Vector DBs).
CollabHiring
Why Cisco?At Cisco, we’re revolutionizing how data and infrastructure connect and protect organizations in the AI era – and beyond. We’ve been innovating fearlessly for 40 years to create solutions that power how humans and technology work together across the physical and digital worlds. These solutions provide customers with unparalleled security, visibility, and insights across the entire digital footprint.
Fueled by the depth and breadth of our technology, we experiment and create meaningful solutions. Add to that our worldwide network of doers and experts, and you’ll see that the opportunities to grow and build are limitless. We work as a team, collaborating with empathy to make really big things happen on a global scale. Because our solutions are everywhere, our impact is everywhere.
We are Cisco, and our power starts with you.
Skills Required
- Bachelor's degree in Computer Science, Software Engineering, or a related technical field plus 8 years of related experience; or Master's degree plus 6 years; or PhD plus 3 years.
- Proven experience as a Technical Lead or Lead SRE/Software Engineer delivering distributed, highly available SaaS platforms at scale.
- Strong proficiency in Python, Go, Java, or C++.
- Experience designing microservices, APIs, and production automation.
- Deep experience with Kubernetes, Docker, and container orchestration in large-scale multi-cluster environments.
- Experience with SRE practices, including SLI/SLO design, observability, incident management, and automated root-cause analysis.
- Hands-on experience building LLM pipelines, AI agents, MCP servers or clients, RAG architectures, and evaluation frameworks.
- Experience with OpenTelemetry, Prometheus, Grafana, Splunk, ThousandEyes, or distributed tracing systems.
- Expertise with AWS, GCP, Azure, Terraform, and GitOps or CI/CD pipelines such as Jenkins or GitHub Actions.
- Experience implementing responsible AI guardrails, deterministic fallback logic, and policy-driven remediation engines.
- Experience with Kafka, Redis, PostgreSQL, Elasticsearch, or vector databases.
Cisco Compensation & Benefits Highlights
The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Cisco and has not been reviewed or approved by Cisco.
-
Healthcare Strength — Health coverage is described as robust with multiple plan options and access to onsite/virtual LifeConnections Health Centers on major campuses. Company materials also highlight mental-health resources and comprehensive preventive care, supporting strong core medical benefits.
-
Leave & Time Off Breadth — Time away includes company‑wide recharge days, a paid birthday, a year‑end shutdown, and paid Critical Time Off for emergencies. Paid volunteer days further expand opportunities to step away and recharge.
-
Parental & Family Support — Policies include a global minimum for paid parental leave for primary caregivers, caregiving concierge services, and on‑site children’s learning centers in select locations. In the U.S., family‑building support is consolidated under Carrot with a defined lifetime maximum, indicating structured assistance across fertility, preservation, adoption, and surrogacy.
Cisco Insights
What We Do
Cisco (NASDAQ: CSCO) enables people to make powerful connections--whether in business, education, philanthropy, or creativity. Cisco hardware, software, and service offerings are used to create the Internet solutions that make networks possible--providing easy access to information anywhere, at any time. Cisco was founded in 1984 by a small group of computer scientists from Stanford University. Since the company's inception, Cisco engineers have been leaders in the development of Internet Protocol (IP)-based networking technologies. Today, with more than 71,000 employees worldwide, this tradition of innovation continues with industry-leading products and solutions in the company's core development areas of routing and switching, as well as in advanced technologies such as home networking, IP telephony, optical networking, security, storage area networking, and wireless technology. In addition to its products, Cisco provides a broad range of service offerings, including technical support and advanced services. Cisco sells its products and services, both directly through its own sales force as well as through its channel partners, to large enterprises, commercial businesses, service providers, and consumers.







