Staff Software Engineer

Reposted One Month Ago
Be an Early Applicant
Noida, Gautam Buddha Nagar, Uttar Pradesh, IND
In-Office
Senior level
Software
The Role
Lead product-area reliability efforts for planet-scale observability and security products. Define and execute reliability roadmaps, manage SLOs, reduce operational toil via automation, participate in on-call rotations and incident RCA, collaborate across infra and engineering teams, write production code, and hire/mentor SREs.
Summary Generated by Built In
Title: Staff Site Reliability Engineer, Cloud Cost & Engineering Efficiency
Location: Noida/ Bangalore (Hybrid)
Summary of role

Sumo Logic's microservices architecture, hosted on AWS, ingests petabytes of data daily across many geographic regions in support of our planet-scale observability and security products, serving hundreds of millions of queries a day against thousands of petabytes of data. At that scale, every inefficiency — in code, architecture, or infrastructure — compounds into real cost.

This role sits within the Product SRE organization, working alongside your global SRE team on your product area's reliability roadmap — with a mandate that leads with code as much as operations. You'll find where Sumo's systems are spending more compute, storage, or engineering time than necessary, and fix it in the code and architecture, not just the infra config, while also carrying the fuller SRE mandate — reliability, security posture, and improving the day-to-day experience of the engineers within your product area. You'll be part of a team that blends SRE and backend software engineering skillsets, partnering closely with product engineering teams across your product area.

This is an engineering role — the work is about shipping code, architecture, and system-level changes that improve unit economics, not managing cost dashboards, tagging, or reserved-instance/savings-plan purchasing.

What you’ll do

  • Continuously discover cost and efficiency opportunities through production profiling, telemetry, cost data, capacity trends, and system-level analysis — across algorithmic inefficiencies, resource-heavy code paths, and architectural decisions — and turn ambiguous problems into prioritized engineering initiatives.
  • Apply performance and capacity engineering techniques to understand CPU, memory, storage, network, and I/O behavior under real production workloads, and optimize the resulting resource footprint.
  • Write production-grade code to implement the optimizations you identify — JVM/GC tuning, algorithmic and resource-efficiency improvements, re-architecting inefficient services — in systems that process petabytes of data daily.
  • Define and track engineering efficiency metrics such as cost per GB ingested, cost per query, cost per event, resource utilization, or cost per customer workload, and translate the work into measurable impact for engineering and business stakeholders.
  • Lead complex, cross-team engineering initiatives from problem discovery through design, implementation, rollout, and measurement — influencing teams where you don't have direct ownership — and help establish engineering patterns and practices that make cost and efficiency a continuous part of the development lifecycle.
  • Partner with engineering teams in your product area to prioritize changes, and with developer infrastructure and Global SRE to align with the broader reliability roadmap.
  • Participate in the SRE responsibilities for different product areas — SLOs, on-call, incident response, and blameless RCA — using those experiences to identify systemic reliability, performance, and efficiency improvements.

What you’ll have

  • B.Tech, M.Tech, or equivalent degree in Computer Science or a related discipline.
  • 8+ years of industry experience with a demonstrated track record of ownership.
  • Strong CS fundamentals — comfortable with algorithmic complexity, data-structure performance characteristics, and system design at scale.
  • Ability to author production-ready code in at least one OO/systems language (Java, Scala, Go, C++, or similar) — depth of engineering ability matters more than which language.
  • Experience with distributed systems and microservice architectures in production.
  • Demonstrated track record of independently identifying ambiguous performance, scalability, or cost problems and driving engineering changes that produced measurable improvements in production.
  • Strong ability to reason quantitatively about system behavior, capacity, performance, and cost, and use production data to validate hypotheses and measure outcomes.
  • Working fluency with cloud infrastructure (AWS compute, storage, networking) — enough to reason about cost and architectural tradeoffs.
  • Comfort moving across the stack, from application code to the infrastructure it runs on, to find root causes of inefficiency.

Nice to have

  • JVM tuning and GC optimization experience at scale.
  • Exposure to cost-attribution/FinOps practices or tooling.
  • Experience with Kubernetes, Terraform, or modern CI/CD tooling.
  • Prior SRE experience — on-call, SLOs, incident response.
  • Experience with streaming technologies (Kafka, Kafka Streams) or observability/security platforms.

Why this role

  • Visibility — Cost/efficiency work maps directly to metrics the business already tracks.
  • Scope you define — you identify where the opportunities are rather than executing a fixed backlog.
  • Real scale — systems ingesting petabytes of data daily and serving hundreds of millions of queries.
About Us

Sumo Logic, Inc. helps make the digital world secure, fast, and reliable by unifying critical security and operational data through its Intelligent Operations Platform. Built to address the increasing complexity of modern cybersecurity and cloud operations challenges, we empower digital teams to move from reaction to readiness—combining agentic AI-powered SIEM and log analytics into a single platform to detect, investigate, and resolve modern challenges. Customers around the world rely on Sumo Logic for trusted insights to protect against security threats, ensure reliability, and gain powerful insights into their digital environments. For more information, visit www.sumologic.com.

Sumo Logic Privacy Policy. Employees will be responsible for complying with applicable federal privacy laws and regulations, as well as organizational policies related to data protection.


Skills Required

  • Cloud native application development experience leveraging best practices and design patterns
  • Strong debugging and troubleshooting skills across the entire technology stack
  • Deep understanding of AWS Networking, Compute, Storage, and managed services
  • Competency with modern CI/CD tooling like Kubernetes, Terraform, Ansible & Jenkins
  • Experience with full life cycle support of services, from creation to production support
  • Versed in Infrastructure as Code practices using technologies like Terraform or Cloud Formation
  • Ability to author production ready code in at least one of the following: Java, Scala or Go
  • Experience with Linux systems and at home on the command line
  • Understand and apply modern approaches to cloud-native software security
  • Experienced with agile frameworks, such as Scrum and Kanban
  • Flexible and willing to step into new roles and responsibilities
  • Willingness to learn and use Sumo Logic products for solving reliability and security issues
  • Bachelor's or Master's Degree in Computer Science, Electrical Engineering, or another scientific or technical discipline
  • 8+ years of professional experience in applied software security role
  • Experience using Sumo Logic or other observability products for reliability and security
  • Experience with planet-scale product development
  • Expert level experience running and operating SaaS products on AWS Cloud
  • Experience with streaming technologies like Kafka, Kafka Streams, or KSQL
  • Expert level experience in one or more of: Java, Go, Scala, or Python
  • Expert level experience in one or more of: Terraform, Jenkins, Kubernetes
  • Extensive experience running and tuning JVM workloads at scale
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Redwood, CA
913 Employees
Year Founded: 2010

What We Do

Sumo Logic is the pioneer in continuous intelligence, a new category of software, which enables organizations of all sizes to address the data challenges and opportunities presented by digital transformation, modern applications, and cloud computing. The Sumo Logic Continuous Intelligence Platform™ automates the collection, ingestion, and analysis of application, infrastructure, security, and IoT data to derive actionable insights within seconds. More than 2,100 customers around the world rely on Sumo Logic to build, run, and secure their modern applications and cloud infrastructures. Sumo Logic delivers its platform as a true, multi-tenant SaaS architecture, across multiple use-cases, enabling businesses to thrive in the Intelligence Economy.

Similar Jobs

Clearwater Analytics (CWAN) Logo Clearwater Analytics (CWAN)

Development Engineer

Fintech • Software • Financial Services
Hybrid
Block Noida Authority Office, Sector 6, Gautam Buddha Nagar, Uttar Pradesh, IND
1100 Employees
In-Office
Noida, Gautam Buddha Nagar, Uttar Pradesh, IND
687 Employees

AlphaSense Logo AlphaSense

Staff Software Engineer

Artificial Intelligence • Fintech • Machine Learning • Natural Language Processing • Business Intelligence
Remote or Hybrid
India
2000 Employees

Walmart Global Tech Logo Walmart Global Tech

Software Engineer

Big Data • Cloud • Logistics • Machine Learning • Retail
Remote or Hybrid
8 Locations
578950 Employees
110K-286K Annually

Similar Companies Hiring

Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Revel Thumbnail
Aerospace • Hardware • Robotics • Software
Marina Del Rey, California
60 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software
New York, New York
30 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account