Senior Site Reliability Engineer

Posted 5 Days Ago
Be an Early Applicant
Bangalore, Bengaluru Urban, Karnataka, IND
In-Office
Senior level
Cloud • Information Technology • Internet of Things • Professional Services • Software
The Role
Lead AWS cost and performance optimization for ThousandEyes cloud platform: drive FinOps practices, identify savings, improve resource utilization and autoscaling, build automation and dashboards, support cost governance and stakeholder reporting, mentor engineers, and balance reliability, scalability, performance, and cost across large-scale cloud systems.
Summary Generated by Built In
Meet the Team

The ThousandEyes Efficiency & Performance team is responsible for improving the reliability, scalability, performance, and cost efficiency of the ThousandEyes cloud platform. The team partners closely with engineering, platform, finance, product, and leadership team members to drive infrastructure optimization, cloud cost governance, performance improvements, and operational excellence across AWS environments.

As a Senior Site Reliability Engineer in the Efficiency & Performance team, you will play a key role in managing and optimizing ThousandEyes cloud infrastructure cost, improving AWS efficiency, driving FinOps practices, and helping engineering teams make data-driven decisions around cost, performance, and scalability.

Your Impact
  • Own and drive AWS cost optimization and efficiency initiatives across ThousandEyes infrastructure.
  • Analyze cloud usage across compute, storage, databases, observability, networking, and data platform workloads to identify savings opportunities.
  • Partner with engineering teams to improve application and infrastructure performance while reducing cloud wastage.
  • Drive FinOps practices including cost visibility, tagging hygiene, showback/chargeback, budgeting, forecasting, anomaly detection, and cost allocation.
  • Improve infrastructure efficiency through better resource utilization, autoscaling, right-sizing, and capacity planning.
  • Build automation, dashboards, reports, and guardrails to improve cost governance and operational visibility.
  • Support ThousandEyes cost reviews, OKR tracking, leadership updates, and stakeholder communications.
  • Collaborate with product, finance, engineering, and platform teams to align cost optimization with business priorities.
  • Identify and reduce underutilized infrastructure, idle resources, over-provisioned workloads, and inefficient service usage.
  • Improve observability, alerting, and monitoring for cost, performance, and platform health indicators.
  • Influence engineering teams to adopt cost-aware architecture, reliable design patterns, and performance-efficient implementation practices.
  • Provide senior-level technical leadership, mentor engineers, and independently drive cross-functional initiatives to closure.
  • Balance reliability, scalability, performance, and cost efficiency in all platform decisions.
Minimum Qualifications
  • Bachelor’s degree or higher in Engineering, Computer Science, or equivalent practical experience.
  • 8–12 years of relevant experience in Site Reliability Engineering, Cloud Infrastructure, Platform Engineering, DevOps, Production Engineering, or Performance Engineering.
  • Strong experience with AWS services such as EC2, S3, RDS, EMR, Lambda, CloudWatch, OpenSearch, ElastiCache, IAM, VPC, and related cloud-native services.
  • Practical experience in cloud cost optimization, FinOps, AWS billing analysis, cost allocation, tagging, budget tracking, forecasting, and cost governance.
  • Experience with performance analysis, capacity planning, infrastructure optimization, and reliability improvements for large-scale cloud platforms.
  • Experience with infrastructure-as-code and automation tools such as Terraform, CloudFormation, Puppet, Ansible, or similar.
  • Strong scripting or programming skills in Python, Go, Shell, or similar languages for automation, reporting, and operational tooling.
  • Experience with observability and monitoring platforms such as ThousandEyes, CloudWatch, Prometheus, Grafana, Splunk, OpenSearch, Datadog, or similar.
  • Strong Linux systems knowledge and understanding of distributed systems.
  • Experience in incident management, production support, reliability engineering, and operational excellence.
  • Ability to analyze large-scale infrastructure, performance, and cost data and convert findings into actionable engineering recommendations.
  • Strong communication skills with the ability to present technical, performance, and cost insights to engineering teams, finance, leadership, and multi-functional stakeholders.
  • Experience working in Agile/Scrum environments and managing priorities across multiple teams.
Preferred Qualifications
  • Experience supporting or optimizing large-scale SaaS platforms in AWS.
  • FinOps certification or equivalent experience with cloud financial management practices.
  • Experience with AWS Savings Plans, Reserved Instances, Spot adoption, Graviton migration, storage lifecycle optimization, workload right-sizing, and commitment planning.
  • Experience with cloud cost management tools such as AWS Cost Explorer, AWS CUR, Cloud-ability, Cloud-Health, or similar platforms.
  • Experience with performance tuning of cloud infrastructure, distributed systems, databases, data pipelines, and high-scale services.
  • Experience with data platforms or large-scale infrastructure components such as EMR, Kafka, OpenSearch, RDS, Airflow, Spark, Redis, Cassandra, or similar.
  • Experience driving cross-team cost optimization programs, efficiency OKRs, governance reviews, executive reporting, and measurable savings outcomes.
  • Ability to influence engineering teams toward cost-conscious architecture, performance-aware design, and operational standard methodologies.
  • Exposure to AI/ML infrastructure cost optimization, capacity planning, GPU/accelerator cost governance, or AI-driven efficiency tooling is a plus.
Why Cisco? 

At Cisco, we’re revolutionizing how data and infrastructure connect and protect organizations in the AI era – and beyond. We’ve been innovating fearlessly for 40 years to create solutions that power how humans and technology work together across the physical and digital worlds. These solutions provide customers with unparalleled security, visibility, and insights across the entire digital footprint.

Fueled by the depth and breadth of our technology, we experiment and create meaningful solutions. Add to that our worldwide network of doers and experts, and you’ll see that the opportunities to grow and build are limitless. We work as a team, collaborating with empathy to make really big things happen on a global scale. Because our solutions are everywhere, our impact is everywhere. 

We are Cisco, and our power starts with you. 

Skills Required

  • Bachelor's degree in Engineering, Computer Science, or equivalent practical experience
  • 8-12 years experience in Site Reliability, Cloud Infrastructure, Platform, DevOps, Production, or Performance Engineering
  • Strong experience with AWS services (EC2, S3, RDS, EMR, Lambda, CloudWatch, OpenSearch, ElastiCache, IAM, VPC)
  • Practical experience in cloud cost optimization and FinOps: billing analysis, cost allocation, tagging, budgeting, forecasting
  • Experience with performance analysis, capacity planning, and infrastructure optimization for large-scale cloud platforms
  • Experience with infrastructure-as-code and automation (Terraform, CloudFormation, Puppet, Ansible)
  • Strong scripting or programming skills (Python, Go, Shell) for automation and tooling
  • Experience with observability and monitoring platforms (ThousandEyes, CloudWatch, Prometheus, Grafana, Splunk, OpenSearch, Datadog)
  • Strong Linux systems knowledge and distributed systems understanding
  • Experience in incident management, production support, and operational excellence
  • Ability to analyze large-scale infrastructure, performance, and cost data and convert to actionable recommendations
  • Strong communication skills for presenting technical and cost insights to engineering, finance, and leadership
  • Experience working in Agile/Scrum environments and managing cross-team priorities
  • Experience supporting or optimizing large-scale SaaS platforms in AWS
  • FinOps certification or equivalent cloud financial management experience
  • Experience with AWS Savings Plans, Reserved Instances, Spot adoption, Graviton migrations, and storage lifecycle optimization
  • Experience with cloud cost management tools (AWS Cost Explorer, AWS CUR, Cloudability, CloudHealth)
  • Performance tuning experience for distributed systems, databases, data pipelines, and high-scale services
  • Experience with data platform technologies (EMR, Kafka, RDS, Airflow, Spark, Redis, Cassandra, OpenSearch)
  • Experience driving cross-team cost optimization programs, OKRs, governance reviews, and measurable savings
  • Exposure to AI/ML infrastructure cost optimization and GPU/accelerator cost governance

Cisco Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Cisco and has not been reviewed or approved by Cisco.

  • Healthcare Strength Comprehensive medical, dental, and vision coverage, mental health support via an EAP, and access to on-site or virtual health centers indicate robust healthcare offerings. Wellness programs, fitness resources, and specialized services further reinforce coverage depth.
  • Leave & Time Off Breadth Generous PTO, a global minimum for paid parental leave, and unique programs like company-wide recharge days and paid volunteer time expand time-away options. Additional offerings such as Critical Time Off and adoption assistance add flexibility for life events.
  • Equity Value & Accessibility Restricted stock units and a discounted employee stock purchase plan are meaningful elements of total compensation. The prominence of equity can materially augment overall pay packages alongside salary and bonuses.

Cisco Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: San Jose, CA
77,500 Employees
Year Founded: 1984

What We Do

Cisco (NASDAQ: CSCO) enables people to make powerful connections--whether in business, education, philanthropy, or creativity. Cisco hardware, software, and service offerings are used to create the Internet solutions that make networks possible--providing easy access to information anywhere, at any time. Cisco was founded in 1984 by a small group of computer scientists from Stanford University. Since the company's inception, Cisco engineers have been leaders in the development of Internet Protocol (IP)-based networking technologies. Today, with more than 71,000 employees worldwide, this tradition of innovation continues with industry-leading products and solutions in the company's core development areas of routing and switching, as well as in advanced technologies such as home networking, IP telephony, optical networking, security, storage area networking, and wireless technology. In addition to its products, Cisco provides a broad range of service offerings, including technical support and advanced services. Cisco sells its products and services, both directly through its own sales force as well as through its channel partners, to large enterprises, commercial businesses, service providers, and consumers.

Similar Jobs

Nexthink Logo Nexthink

Senior Site Reliability Engineer

Artificial Intelligence • Big Data • Cloud • Information Technology • Machine Learning • Software
Hybrid
Bengaluru, Bengaluru Urban, Karnataka, IND
1200 Employees

Ping Identity Logo Ping Identity

Senior Site Reliability Engineer

Cloud • Security • Software
Hybrid
Bengaluru, Bengaluru Urban, Karnataka, IND
2300 Employees

Microsoft Logo Microsoft

Senior Site Reliability Engineer

Software • Quantum Computing • Metaverse • Infrastructure as a Service (IaaS)
In-Office
2 Locations
206870 Employees

Zeta Logo Zeta

Senior Site Reliability Engineer

Cloud • Fintech • Financial Services
In-Office
Bangalore, Bengaluru Urban, Karnataka, IND
1834 Employees

Similar Companies Hiring

Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees
Kepler  Thumbnail
Fintech • Software
New York, New York
6 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account