Sr. Manager, SRE & Performance

Posted 5 Days Ago
Be an Early Applicant
Mountain View, CA, USA
In-Office
223K-310K Annually
Senior level
Artificial Intelligence • Information Technology • Software
The Role
Leads SRE and Performance Engineering teams responsible for the reliability, scalability, observability, performance, and operational excellence of a distributed SaaS platform. Defines SLOs, performance standards, capacity plans, incident management, and production-readiness practices; guides architecture reviews, resilience improvements, CI/CD validation, automation, and cloud cost optimization. Provides technical leadership during incidents and develops engineering teams while partnering across Engineering, Product, Security, Architecture, and SaaS Operations.
Summary Generated by Built In

Job Description:

We are Omnissa! 
 
Omnissa is the first AI-driven digital work platform, built to support flexible, secure, work-from anywhere experiences. We integrate industry-leading solutions—including Unified Endpoint Management, Virtual  Apps and Desktops, Digital Employee Experience, and Security & Compliance—into a seamless, autonomous workspace that adapts to how people work. Our platform boosts employee engagement while optimizing IT operations, security, and cost.  
 
Guided by our Core Values—Act in Alignment, Build Trust, Foster Inclusiveness, Drive Efficiency, and Maximize Customer Valuewe’re growing rapidly and committed to delivering meaningful impact. If you're passionate about shaping the future of work, we’d love to hear from you. 

At Omnissa, we are committed to maintaining a fair, consistent, and secure hiring process for all candidates. As part of this approach, we use standard interview and verification practices designed to ensure alignment and protect both candidates and the organization. These practices are applied thoughtfully and with respect for candidate privacy. 

What is the opportunity?

We are looking for a Senior Manager of Site Reliability Engineering (SRE) & Performance Engineering to lead teams responsible for the reliability, scalability, performance, and operational excellence of our distributed SaaS platform. This is a hands-on technical leadership role for someone who can operate at the intersection of software engineering, distributed systems, cloud infrastructure, SRE, and performance engineering. You will lead engineers while partnering closely with development, architecture, security, SaaS operations, and product teams to ensure our services are designed and operated for scale, resilience, efficiency, and predictable performance. You will help establish engineering standards around SLOs, observability, performance testing, capacity planning, incident management, production readiness, and continuous reliability improvement. Here's a breakdown:

  • Lead and grow SRE and Performance Engineering teams, providing technical direction, coaching, career development, and establishing a strong culture of ownership and engineering excellence.
  • Define and drive the organization's SRE and performance engineering strategy, including reliability goals, performance objectives, scalability standards, and operational readiness requirements.
  • Partner with engineering teams to ensure systems are designed for high availability, scalability, fault tolerance, performance, and operability from the beginning rather than addressing these concerns after deployment.
  • Establish and drive adoption of SLIs, SLOs, error budgets, service health indicators, and production readiness criteria for critical services.
  • Lead performance engineering initiatives across distributed systems, including workload modeling, benchmarking, profiling, scalability testing, capacity planning, and bottleneck analysis across CPU, memory, I/O, storage, database, and network layers.
  • Drive observability strategy across metrics, logs, traces, dashboards, and alerting, while improving signal quality and reducing noisy or non-actionable alerts.
  • Provide technical leadership during complex production incidents, helping teams diagnose distributed system failures, latency spikes, resource contention, database issues, network problems, and cascading failures.
  • Drive effective incident management, RCA/postmortem practices, corrective actions, and systemic reliability improvements, ensuring lessons from incidents translate into engineering changes.
  • Partner with architects and senior engineers on system design and architecture reviews, challenging designs around scalability, resilience, failure modes, performance, data architecture, and operational complexity.
  • Establish scalable approaches for performance regression detection and reliability validation within CI/CD pipelines, enabling issues to be identified earlier in the development lifecycle.
  • Drive capacity management and forecasting, using production telemetry, workload characteristics, and performance models to anticipate infrastructure requirements and scalability limits.
  • Improve platform efficiency through performance optimization, infrastructure right-sizing, and cost-aware engineering, balancing reliability, performance, and cloud cost.
  • Champion automation and engineering-driven operations, reducing manual operational work and toil through software, tooling, Infrastructure as Code, and automated remediation.
  • Establish measurable SRE and performance KPIs, such as availability/SLO attainment, MTTR, change failure rate, alert quality, performance regression rates, capacity headroom, operational toil, and recurring incident reduction.
  • Collaborate across Product, Engineering, Security, Cloud/SaaS Operations, and Architecture organizations to deliver secure, resilient, scalable, and production-ready services.
What will you bring to the company?Leadership & Engineering Management
  • 12+ years of software engineering experience, with significant experience building and operating large-scale backend or distributed systems.
  • 5+ years of engineering leadership/management experience, preferably leading SRE, Performance Engineering, Platform Engineering, Infrastructure, or backend engineering teams.
  • Proven ability to build, mentor, and develop high-performing engineering teams, including senior and staff-level engineers.
  • Strong ability to balance people leadership, technical strategy, operational priorities, and business objectives.
  • Experience influencing engineering practices across teams without relying solely on organizational authority.
  • Demonstrated ability to work effectively with senior engineers, architects, engineering managers, product leaders, security teams, and executive stakeholders.
Technical Depth
  • Strong understanding of distributed systems and microservices architectures, including scalability, availability, consistency, fault tolerance, failure modes, and architectural trade-offs.
  • Strong software engineering background, preferably with experience in C#/.NET, Go or Java, and the ability to participate meaningfully in architecture, design, and code-level technical discussions.
  • Deep understanding of Linux systems, including processes, memory, CPU scheduling, networking, file systems, containers, and system-level performance diagnostics.
  • Strong experience with Docker and container orchestration using HashiStack technologies such as Nomad, Consul, and Vault.
  • Experience operating high-scale data platforms using technologies such as Kafka, PostgreSQL, OpenSearch, Redis/Valkey, or equivalent technologies.
  • Strong cloud experience, preferably with AWS, including services such as EC2, EKS, MSK, Aurora/RDS, OpenSearch, networking, storage, and cloud observability services.
  • Strong understanding of CI/CD, Infrastructure as Code, deployment automation, release strategies, and modern DevOps practices.
SRE & Performance Engineering
  • Deep understanding of SRE principles, including SLIs/SLOs, error budgets, availability engineering, incident management, production readiness, operational toil, and reliability automation.
  • Proven experience designing and implementing observability strategies using metrics, logging, tracing, dashboards, and actionable alerting.
  • Strong understanding of performance engineering methodologies, including workload modeling, benchmarking, profiling, stress/load testing, scalability analysis, and performance regression detection.
  • Experience diagnosing performance issues across applications, databases, operating systems, containers, infrastructure, and network layers.
  • Experience with capacity planning and forecasting, including translating workload growth into infrastructure and service capacity requirements.
  • Strong understanding of resilience engineering, including graceful degradation, retries, timeouts, circuit breakers, backpressure, rate limiting, disaster recovery, and failure testing.
  • Demonstrated ability to turn production incidents and performance findings into systemic engineering improvements rather than tactical fixes.

Location:   Mountain View, CA
Location Type: hybrid
Education: Bachelor's Degree preferred, or equivalent combination of education and relevant professional experience. 

Compensation: The typical base salary for this role is between USD $223,000 – $310,000 per year and it may be eligible for participation in a corporate bonus program. Actual compensation offer may vary from posted hiring range based upon geographic location, work experience, education, skill level, or other relevant factors. In addition to competitive compensation, Omnissa offers a variety of benefits such as employee ownership, health insurance, 401k with matching contributions, disability insurance, paid-time off, growth opportunities, and more. 
 
Omnissa is an Equal Employment Opportunity company and Prohibits Discrimination and Harassment of Any Kind:  
Omnissa is committed to the principle of equal employment opportunity and to providing a work environment free of discrimination and harassment. All employment decisions at Omnissa are based on business needs, job requirements and individual qualifications, without regard to race, color, religion, ancestry, ethnicity, national, social or ethnic origin, sex (including pregnancy), age, physical, mental or sensory disability, HIV status, sexual orientation, gender identity and/or expression, marital, civil union or domestic partnership status, past, present, or prospective service in the uniformed services, family medical history or genetic information, family or parental status, veteran status, or any other status protected by applicable laws or regulations in the locations where we operate. Omnissa will not tolerate discrimination or harassment based on any of these characteristics. Omnissa welcomes applicants of all ages. Omnissa will provide reasonable accommodations to applicants and employees who have protected disabilities consistent with applicable federal, state and local law. 

Skills Required

  • 12+ years of software engineering experience, including significant experience building and operating large-scale backend or distributed systems
  • 5+ years of engineering leadership or management experience, preferably leading SRE, Performance Engineering, Platform Engineering, Infrastructure, or backend engineering teams
  • Experience building, mentoring, and developing high-performing engineering teams, including senior and staff-level engineers
  • Ability to balance people leadership, technical strategy, operational priorities, and business objectives
  • Experience influencing engineering practices across teams without relying solely on organizational authority
  • Ability to work effectively with senior engineers, architects, engineering managers, product leaders, security teams, and executive stakeholders
  • Strong understanding of distributed systems and microservices architectures, including scalability, availability, consistency, fault tolerance, failure modes, and architectural trade-offs
  • Strong software engineering background, preferably with experience in C#/.NET, Go, or Java
  • Deep understanding of Linux systems, including processes, memory, CPU scheduling, networking, file systems, containers, and system-level performance diagnostics
  • Strong experience with Docker and container orchestration using Nomad, Consul, and Vault
  • Experience operating high-scale data platforms using Kafka, PostgreSQL, OpenSearch, Redis/Valkey, or equivalent technologies
  • Strong cloud experience, preferably with AWS and services such as EC2, EKS, MSK, Aurora/RDS, OpenSearch, networking, storage, and cloud observability services
  • Strong understanding of CI/CD, Infrastructure as Code, deployment automation, release strategies, and modern DevOps practices
  • Deep understanding of SRE principles, including SLIs, SLOs, error budgets, availability engineering, incident management, production readiness, operational toil, and reliability automation
  • Experience designing and implementing observability strategies using metrics, logging, tracing, dashboards, and actionable alerting
  • Strong understanding of performance engineering methodologies, including workload modeling, benchmarking, profiling, stress/load testing, scalability analysis, and performance regression detection
  • Experience diagnosing performance issues across applications, databases, operating systems, containers, infrastructure, and network layers
  • Experience with capacity planning and forecasting, including translating workload growth into infrastructure and service capacity requirements
  • Strong understanding of resilience engineering, including graceful degradation, retries, timeouts, circuit breakers, backpressure, rate limiting, disaster recovery, and failure testing
  • Ability to turn production incidents and performance findings into systemic engineering improvements
  • Bachelor's degree or equivalent combination of education and relevant professional experience

Omnissa Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Omnissa and has not been reviewed or approved by Omnissa.

  • Healthcare Strength Healthcare offerings include comprehensive medical, dental, and vision coverage, with wellness options referenced across materials. Health plans are characterized as decent to strong within a standard tech package.
  • Retirement Support A 401(k) with company match is part of the core package and is specifically highlighted as a valued benefit in U.S. materials. Retirement support is presented as a stable element of total rewards.
  • Leave & Time Off Breadth Vacation and PTO are highlighted positively, with generous paid time off and holidays noted in public benefits descriptions. Time-off programs are portrayed as supportive of work-life balance.

Omnissa Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Mountain View, California
2,430 Employees

What We Do

Omnissa is the digital work platform leader, trusted by thousands of organizations worldwide as the former VMware End-User Computing business. We make digital work, work – for businesses and their people. No painful IT processes or productivity trade-offs. Instead, a seamlessly delivered digital employee experience that simplifies work. Our comprehensive digital work platform enables IT teams to provide secure, personalized experiences for every employee, on any device. Omnissa unifies, automates, and efficiently scales the digital workspace. By empowering employees to do their best work, anywhere, we help workforces everywhere unlock exponential business value. All is made possible with the Omnissa™ Platform, the first AI-driven digital work platform for smart, seamless, and secure work experiences from anywhere. It integrates multiple industry-leading solutions across Unified Endpoint Management, Virtual Desktops and Apps, Digital Employee Experience, and Security and Compliance. By continuously adapting to users’ work styles, Omnissa optimizes user experience, security, IT operations and costs.

Similar Jobs

Liberty Mutual Insurance Logo Liberty Mutual Insurance

Executive Underwriter, Liberty Mutual Mobility Solutions

Artificial Intelligence • Fintech • Insurance • Marketing Tech • Software • Analytics
Hybrid
6 Locations
40000 Employees
68K-278K Annually

Airwallex Logo Airwallex

Senior Product Manager

Artificial Intelligence • Fintech • Payments • Business Intelligence • Financial Services • Generative AI
Hybrid
San Francisco, CA, USA
2300 Employees
160K-230K Annually

Superhuman Logo Superhuman

Data Engineer

Artificial Intelligence • Information Technology • Machine Learning • Natural Language Processing • Productivity • Software • Generative AI
Hybrid
2 Locations
1500 Employees
190K-240K Annually

PwC Logo PwC

Tax - Japanese Business Network - Intern - Winter 2028 - Destination CPA

Artificial Intelligence • Professional Services • Business Intelligence • Consulting • Cybersecurity • Generative AI
Hybrid
7 Locations
370000 Employees
29K-48K Hourly

Similar Companies Hiring

Kepler  Thumbnail
Artificial Intelligence • Fintech • Software
New York, New York
9 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Revel.io Thumbnail
Aerospace • Hardware • Robotics • Software
US
50 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account