The Network Assurance Data Platform team within Cisco ThousandEyes is responsible for building, operating, and scaling the core data infrastructure that powers large-scale network assurance, analytics, and intelligence capabilities.
The team manages critical cloud infrastructure, big data workflows, ML platform components, and reliability engineering practices across AWS environments. We focus on improving platform reliability, scalability, performance, automation, and cost efficiency while enabling engineering and data teams to deliver business-critical capabilities at scale.
As a Senior Site Reliability Engineer in the NADP team, you will work on highly scalable infrastructure supporting data pipelines, ML workloads, cloud-native services, and cost-optimized AWS operations.
Your Impact- Own and operate scalable, reliable, and cost-efficient infrastructure for the ThousandEyes Network Assurance Data Platform.
- Manage, optimize, and improve large-scale data workflows using Apache Airflow, AWS EMR, Spark, and Hadoop-based processing platforms.
- Operate and improve Amazon EKS environments supporting containerized services, ML workloads, and production-grade platform components.
- Build and maintain infrastructure automation using Terraform and other infrastructure-as-code practices.
- Develop Python-based automation, integrations, operational tooling, reporting, and reliability improvements.
- Drive FinOps practices across NADP infrastructure, including cost visibility, cost allocation, forecasting, anomaly detection, optimization, and governance.
- Partner with data engineering, ML, platform, finance, and product teams to improve reliability, performance, scalability, and cost efficiency.
- Identify infrastructure bottlenecks, performance issues, inefficient workloads, and cost optimization opportunities.
- Improve observability, alerting, incident response, and operational readiness across data and ML platforms.
- Support capacity planning, right-sizing, autoscaling, storage optimization, and workload efficiency across AWS services.
- Lead technical discussions, influence design decisions, and guide teams toward reliable and cost-conscious architecture.
- Provide senior-level technical leadership, mentorship, and operational guidance to engineers across the team.
- Drive continuous improvement in platform reliability, automation, deployment practices, and operational excellence.
- Bachelor’s degree or higher in Engineering, Computer Science, or equivalent practical experience.
- 8–10 years of relevant experience in Site Reliability Engineering, DevOps, Cloud Infrastructure, Platform Engineering, Data Infrastructure, or Production Engineering.
- Strong experience operating production infrastructure on AWS.
- Strong experience with Apache Airflow for workflow orchestration, pipeline operations, scheduling, monitoring, and troubleshooting.
- Experience with AWS EMR, Spark, Hadoop, or similar large-scale data processing platforms.
- Strong experience operating Amazon EKS or Kubernetes-based environments in production.
- Experience supporting containerized workloads, preferably including ML workloads or data platform services.
- Strong experience with Terraform and infrastructure-as-code practices.
- Strong Python programming or scripting experience for automation, integrations, operational tooling, and infrastructure workflows.
- Practical experience with cloud cost optimization, FinOps, AWS cost analysis, tagging, budgeting, forecasting, and cost governance.
- Strong understanding of Linux systems, networking, distributed systems, and production troubleshooting.
- Experience with observability tools such as CloudWatch, Prometheus, Grafana, Splunk, OpenSearch, Datadog, or similar platforms.
- Experience with incident management, production support, root cause analysis, reliability improvements, and operational excellence.
- Ability to analyze infrastructure, performance, and cost data and convert findings into clear technical recommendations.
- Strong communication skills with the ability to collaborate across data, ML, platform, finance, and engineering teams.
- Ability to operate independently, drive initiatives end to end, and provide technical leadership in a fast-paced environment.
- Experience working with large-scale SaaS platforms or high-volume data infrastructure.
- Experience with ML infrastructure, model execution platforms, batch processing, or data pipeline reliability.
- Experience optimizing EMR, Spark, Airflow, EKS, storage, and compute workloads for performance and cost.
- Experience with AWS services such as EC2, S3, RDS, IAM, VPC, CloudWatch, OpenSearch, Lambda, ElastiCache, and related cloud-native services.
- Experience with AWS Savings Plans, Reserved Instances, Spot adoption, Graviton migration, storage lifecycle management, and workload right-sizing.
- Experience with cloud cost management tools such as AWS Cost Explorer, AWS CUR, Cloudability, CloudHealth, Kubecost, or similar platforms.
- Experience driving FinOps programs, cost reviews, stakeholder reporting, OKR tracking, and executive-level updates.
- Experience with CI/CD systems, GitHub workflows, Atlantis, or similar deployment automation platforms.
- Experience with Puppet, Ansible, Helm, Argo CD, or other configuration and deployment management tools.
- FinOps certification or equivalent cloud financial management experience is a plus.
- Ability to influence engineering teams toward cost-aware, scalable, and reliable design patterns.
- NADP data and ML infrastructure is reliable, scalable, performant, and cost-efficient.
- Airflow, EMR, Spark, and EKS workloads are operated with strong observability, automation, and production maturity.
- AWS cost visibility, forecasting, and governance are improved across NADP-owned infrastructure.
- Cloud wastage is reduced through right-sizing, automation, storage optimization, and workload efficiency improvements.
- Infrastructure is managed through production-grade Terraform and automation practices.
- Engineering and data teams have better visibility into platform health, cost drivers, risks, and optimization opportunities.
- Cross-functional stakeholders receive clear updates on reliability, performance, cost trends, and execution progress.
- The team continuously improves operational excellence, incident response, and platform reliability.
At Cisco, we’re revolutionizing how data and infrastructure connect and protect organizations in the AI era – and beyond. We’ve been innovating fearlessly for 40 years to create solutions that power how humans and technology work together across the physical and digital worlds. These solutions provide customers with unparalleled security, visibility, and insights across the entire digital footprint.
Fueled by the depth and breadth of our technology, we experiment and create meaningful solutions. Add to that our worldwide network of doers and experts, and you’ll see that the opportunities to grow and build are limitless. We work as a team, collaborating with empathy to make really big things happen on a global scale. Because our solutions are everywhere, our impact is everywhere.
We are Cisco, and our power starts with you.
Skills Required
- Bachelor's degree in Engineering, Computer Science, or equivalent practical experience
- 8-10 years experience in Site Reliability Engineering, DevOps, Cloud Infrastructure, Platform Engineering, Data Infrastructure, or Production Engineering
- Strong experience operating production infrastructure on AWS
- Strong experience with Apache Airflow for workflow orchestration and pipeline operations
- Experience with AWS EMR, Spark, Hadoop, or similar large-scale data processing platforms
- Strong experience operating Amazon EKS or Kubernetes-based environments in production
- Experience supporting containerized workloads, including ML or data platform services
- Strong experience with Terraform and infrastructure-as-code practices
- Strong Python programming or scripting experience for automation and operational tooling
- Practical experience with cloud cost optimization, FinOps, cost analysis, tagging, budgeting, and governance
- Strong understanding of Linux systems, networking, distributed systems, and production troubleshooting
- Experience with observability tools such as CloudWatch, Prometheus, Grafana, Splunk, OpenSearch, or Datadog
- Experience with incident management, production support, root cause analysis, and reliability improvements
- Ability to analyze infrastructure, performance, and cost data and provide clear technical recommendations
- Strong communication skills and ability to collaborate across data, ML, platform, finance, and engineering teams
- Ability to operate independently, drive initiatives end to end, and provide technical leadership
- Experience working with large-scale SaaS platforms or high-volume data infrastructure
- Experience with ML infrastructure, model execution platforms, batch processing, or data pipeline reliability
- Experience optimizing EMR, Spark, Airflow, EKS, storage, and compute workloads for performance and cost
- Experience with AWS services such as EC2, S3, RDS, IAM, VPC, CloudWatch, OpenSearch, Lambda, ElastiCache
- Experience with AWS Savings Plans, Reserved Instances, Spot adoption, Graviton migration, storage lifecycle management, and workload right-sizing
- Experience with cloud cost management tools such as AWS Cost Explorer, AWS CUR, Cloudability, CloudHealth, Kubecost, or similar
- Experience driving FinOps programs, cost reviews, stakeholder reporting, OKR tracking, and executive-level updates
- Experience with CI/CD systems, GitHub workflows, Atlantis, or similar deployment automation platforms
- Experience with Puppet, Ansible, Helm, Argo CD, or other configuration and deployment management tools
- FinOps certification or equivalent cloud financial management experience
- Ability to influence engineering teams toward cost-aware, scalable, and reliable design patterns
Cisco Compensation & Benefits Highlights
The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Cisco and has not been reviewed or approved by Cisco.
-
Healthcare Strength — Health coverage is described as robust with multiple plan options and access to onsite/virtual LifeConnections Health Centers on major campuses. Company materials also highlight mental-health resources and comprehensive preventive care, supporting strong core medical benefits.
-
Leave & Time Off Breadth — Time away includes company‑wide recharge days, a paid birthday, a year‑end shutdown, and paid Critical Time Off for emergencies. Paid volunteer days further expand opportunities to step away and recharge.
-
Parental & Family Support — Policies include a global minimum for paid parental leave for primary caregivers, caregiving concierge services, and on‑site children’s learning centers in select locations. In the U.S., family‑building support is consolidated under Carrot with a defined lifetime maximum, indicating structured assistance across fertility, preservation, adoption, and surrogacy.
Cisco Insights
What We Do
Cisco (NASDAQ: CSCO) enables people to make powerful connections--whether in business, education, philanthropy, or creativity. Cisco hardware, software, and service offerings are used to create the Internet solutions that make networks possible--providing easy access to information anywhere, at any time. Cisco was founded in 1984 by a small group of computer scientists from Stanford University. Since the company's inception, Cisco engineers have been leaders in the development of Internet Protocol (IP)-based networking technologies. Today, with more than 71,000 employees worldwide, this tradition of innovation continues with industry-leading products and solutions in the company's core development areas of routing and switching, as well as in advanced technologies such as home networking, IP telephony, optical networking, security, storage area networking, and wireless technology. In addition to its products, Cisco provides a broad range of service offerings, including technical support and advanced services. Cisco sells its products and services, both directly through its own sales force as well as through its channel partners, to large enterprises, commercial businesses, service providers, and consumers.

.png)






