Site Reliability Engineer (SRE) – Observability & Platform Operations

Reposted Yesterday
Be an Early Applicant
Sofia, Sofia-grad, BGR
In-Office
Mid level
Artificial Intelligence • Information Technology • Software
The Role
As a Site Reliability Engineer, you will ensure system reliability and performance, manage incident response, design observability solutions, and enhance automation workflows.
Summary Generated by Built In

Job Description:

We Are Omnissa!

Omnissa is the first AI-driven digital work platform, built to support flexible, secure, work-from-anywhere experiences. We integrate industry-leading solutions—including Unified Endpoint Management, Virtual Apps and Desktops, Digital Employee Experience, and Security & Compliance—into a seamless, autonomous workspace that adapts to how people work. Our platform boosts employee engagement while optimizing IT operations, security, and cost.

Guided by our Core Values—Act in Alignment, Build Trust, Foster Inclusiveness, Drive Efficiency, and Maximize Customer Value—we’re growing rapidly and committed to delivering meaningful impact. If you’re passionate about shaping the future of work, we’d love to hear from you.

The Team

Our internal Platform Engineering team architects and operates Omnissa's enterprise-grade infrastructure. Our environment includes:

  • Core platforms: VMware Cloud Foundation, Apache CloudStack, Proxmox, Kubernetes, and S3-compatible object storage

  • Observability: Prometheus, Grafana, Loki, and Ansible

  • AI-driven automation: An internal incident diagnosis platform built on Ollama, n8n, and MCP servers to reduce MTTD and MTTR

The Role

We're seeking an SRE with deep observability expertise (Grafana, Loki, Prometheus, automation, and scripting) to maintain the reliability, performance, and operational integrity of our platforms. You'll work across planned and unplanned workstreams with engineering, incident management, and service owners. The role includes an on-call rotation covering nights and weekends.
 

Key Responsibilities

  • Design, deploy, and maintain Loki, Grafana, Prometheus, and observability pipelines; expand logging, metrics, and tracing coverage

  • Build and refine automation and AI workflows for incident analysis and auto-remediation

  • Drive reliability through capacity planning, performance optimization, SLIs/SLOs, and root cause analysis

  • Participate in the global on-call rotation; manage incidents and outages and lead post-mortem reviews

  • Use Atlassian tools (Jira, Confluence, Opsgenie) for task, change, and incident management

  • Operate and improve internal clouds (vCF, CloudStack, Proxmox), Kubernetes clusters, and S3-compatible storage

Required Skills

  • Hands-on expertise with Grafana, Loki, Tempo (or similar tracing), and Prometheus

  • At least one scripting/programming language

  • Configuration management tools (Ansible, SaltStack)

  • Strong Linux skills and experience operating large-scale, highly available distributed systems

  • Familiarity with Kubernetes, CI/CD, and Infrastructure as Code

  • Comfortable with on-call participation and incident leadership

  • Experience with Atlassian tools; proficiency in Linux and Windows

Nice to Have

  • Exposure to Ollama, n8n, or similar AI orchestration tooling

  • Experience with S3/open-source object stores (SeaweedFS, Ceph)

  • Knowledge of virtualization stacks (Proxmox, vSphere/vCF, CloudStack)

  • Background in SRE culture, including SLIs/SLOs and error budgeting

Omnissa is committed to building a workforce that reflects the communities we serve across the globe. We believe this brings unique perspectives, experiences, and ideas, which are essential for driving innovation and achieving business success. We hire based on merit and provide equal opportunity for all.

Skills Required

  • Hands-on expertise with Grafana, Loki, Tempo, and Prometheus
  • At least one scripting/programming language
  • Configuration management tools (Ansible, SaltStack)
  • Strong Linux skills and experience operating large-scale, highly available distributed systems
  • Familiarity with Kubernetes, CI/CD, and Infrastructure as Code
  • Comfortable with on-call participation and incident leadership
  • Experience with Atlassian tools; proficiency in Linux and Windows

Omnissa Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Omnissa and has not been reviewed or approved by Omnissa.

  • Healthcare Strength Healthcare offerings include comprehensive medical, dental, and vision coverage, with wellness options referenced across materials. Health plans are characterized as decent to strong within a standard tech package.
  • Retirement Support A 401(k) with company match is part of the core package and is specifically highlighted as a valued benefit in U.S. materials. Retirement support is presented as a stable element of total rewards.
  • Leave & Time Off Breadth Vacation and PTO are highlighted positively, with generous paid time off and holidays noted in public benefits descriptions. Time-off programs are portrayed as supportive of work-life balance.

Omnissa Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Mountain View, California
2,430 Employees

What We Do

Omnissa is the digital work platform leader, trusted by thousands of organizations worldwide as the former VMware End-User Computing business. We make digital work, work – for businesses and their people. No painful IT processes or productivity trade-offs. Instead, a seamlessly delivered digital employee experience that simplifies work. Our comprehensive digital work platform enables IT teams to provide secure, personalized experiences for every employee, on any device. Omnissa unifies, automates, and efficiently scales the digital workspace. By empowering employees to do their best work, anywhere, we help workforces everywhere unlock exponential business value. All is made possible with the Omnissa™ Platform, the first AI-driven digital work platform for smart, seamless, and secure work experiences from anywhere. It integrates multiple industry-leading solutions across Unified Endpoint Management, Virtual Desktops and Apps, Digital Employee Experience, and Security and Compliance. By continuously adapting to users’ work styles, Omnissa optimizes user experience, security, IT operations and costs.

Similar Jobs

Tufin Logo Tufin

Consultant

Security • Cybersecurity
Remote or Hybrid
27 Locations
500 Employees

Pfizer Logo Pfizer

Vice President, Strategy, Value & Innovation

Artificial Intelligence • Healthtech • Machine Learning • Natural Language Processing • Biotech • Pharmaceutical
In-Office or Remote
43 Locations
121990 Employees

Pfizer Logo Pfizer

Vice President, Build - Data, Engineering & AI

Artificial Intelligence • Healthtech • Machine Learning • Natural Language Processing • Biotech • Pharmaceutical
In-Office or Remote
43 Locations
121990 Employees

Pfizer Logo Pfizer

Quality Assurance Manager

Artificial Intelligence • Healthtech • Machine Learning • Natural Language Processing • Biotech • Pharmaceutical
Hybrid
2 Locations
121990 Employees

Similar Companies Hiring

Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees
Kepler  Thumbnail
Fintech • Software
New York, New York
6 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account