Principal Site Reliability Engineer

Posted 3 Days Ago
Be an Early Applicant
Hiring Remotely in Sofia, Sofia-grad, BGR
Remote
Senior level
Natural Language Processing • Software • Conversational AI
It’s not about the AI, it’s about the conversation. We put conversations at the center of business for our customers.
The Role
Provides technical leadership for site reliability and platform engineering across multiple teams. Designs scalable, secure GCP systems; leads Kubernetes, networking, observability, security, CI/CD, and infrastructure-as-code initiatives; establishes SRE practices; responds to incidents; develops automation and reusable platforms; and mentors senior technical staff.
Summary Generated by Built In

LivePerson (NASDAQ: LPSN) is the global leader in enterprise conversations. Hundreds of the world’s leading brands — including HSBC, Chipotle, and Virgin Media — use our award-winning Conversational Cloud platform to connect with millions of consumers. We power nearly a billion conversational interactions every month, providing a uniquely rich data set and safety tools to unlock the power of Conversational AI for better customer experiences.  

At LivePerson, we foster an inclusive workplace culture that encourages meaningful connection, collaboration, and innovation. Everyone is invited to ask questions, actively seek new ways to achieve success, nd reach their full potential. We are continually looking for ways to improve our products and make things better. This means spotting opportunities, solving ambiguities, and seeking effective solutions to the problems our customers care about. 


Overview:

LivePerson transforms customer care from voice calls to mobile messaging. Our cloud-based software platform, LiveEngage, allows brands with millions of customers and tens of thousands of care agents to deliver digital experiences at scale. As the market leader in real-time intelligent customer engagement, we are a B2B SaaS company with 20 years of experience and the heart of a startup.

The Cloud DevOps team at LivePerson is looking for a Principal Site Reliability Engineer (Principal SRE) to provide technical leadership across the organization and help shape the reliability, scalability, security, and operational excellence of our cloud platforms and services.

The ideal candidate is a highly experienced engineer who can solve complex technical problems, influence engineering teams without direct authority, and drive large-scale initiatives from strategy through implementation. This is a highly technical role focused on architecture, reliability engineering, automation, and improving the overall engineering maturity of the organization.


You will: 

  • Provide technical leadership and direction for reliability and platform engineering across multiple teams.
  • Design and evolve highly available, scalable, secure, and resilient systems, with a strong focus on Google Cloud Platform (GCP).
  • Lead complex, cross-team initiatives across cloud infrastructure, Kubernetes, networking, observability, security, and software delivery.
  • Define and drive SRE practices including SLOs, SLIs, error budgets, reliability reviews, capacity planning, and operational readiness.
  • Lead technical response to complex production incidents and drive long-term corrective and preventative actions.
  • Develop automation and infrastructure-as-code solutions using Python, Terraform, Ansible, Bash, and other modern engineering tools.
  • Provide technical leadership for Kubernetes platforms and containerized workloads, including architecture, scalability, performance, and reliability.
  • Establish and evolve GitOps deployment practices using Kubernetes, Helm, and FluxCD.Define and improve CI/CD practices using GitLab CI/CD, focusing on reliability, security, scalability, and developer experience.
  • Drive observability improvements using metrics, logs, traces, dashboards, and actionable alerting.Identify systemic reliability risks, technical debt, and architectural weaknesses and drive sustainable solutions.
  • Create reusable platforms, tooling, and engineering patterns that enable teams to operate reliable services independently.Mentor Senior SREs, Team Leads, and other technical leaders while raising engineering standards across the organization.Influence architecture and technical decisions across teams and communicate complex technical concepts and trade-offs to engineering leadership.

You have:

  • Bachelor's degree in Computer Science, Engineering, or a related field, or equivalent practical experience.
  • 7+ years of experience in SRE, DevOps, Cloud Infrastructure, Systems Engineering, or a related field.
  • Experience operating at a Principal, Staff+, Architect, or equivalent senior technical leadership level.
  • Extensive experience designing and operating highly available, scalable, secure, and distributed production systems.
  • Strong hands-on experience with Google Cloud Platform (GCP) and cloud architecture, including networking, IAM, compute, storage, and managed services.
  • Extensive experience with Kubernetes and containerization technologies such as Docker.
  • Hands-on experience with service meshes such as Istio.
  • Strong programming and scripting experience with Python, Bash, or similar languages, focused on automation and platform engineering.
  • Extensive experience with Terraform and automation/configuration management tools such as Ansible.
  • Strong experience with GitOps, Helm, and FluxCD.
  • Strong experience designing and maintaining CI/CD pipelines using GitLab CI/CD.
  • Deep understanding of Linux, networking, DNS, load balancing, TLS/SSL, authentication, and security fundamentals.
  • Strong experience with observability platforms such as Prometheus, Grafana, and Alertmanager.
  • Deep understanding of SRE principles, distributed systems, scalability, fault tolerance, and performance engineering.
  • Proven ability to influence architecture and technical decisions across multiple teams and organizational boundaries.
  • Excellent communication skills and the ability to mentor senior engineers and technical leaders.

Nice to Have:

  • Experience operating large-scale B2B SaaS or globally distributed systems.
  • Experience with secrets management and security platforms such as HashiCorp Vault.
  • Experience with PostgreSQL and other distributed data services.
  • Experience operating and modernizing hybrid cloud and on-premises infrastructure.
  • Experience leading large-scale cloud migrations or infrastructure modernization initiatives.
  • Experience designing internal developer platforms and self-service infrastructure capabilities.
  • Experience with disaster recovery, business continuity, and resilience engineering.
  • Experience defining organization-wide engineering standards and reliability frameworks.

Benefits:  

  • Health: medical, dental, and vision
  • Time away: 28 vacation days
  • Development: Generous tuition reimbursement and access to internal professional development resources. 
  • Additional: Food Vouchers.
  • #LI-Remote

Why you’ll love working here:

As leaders in enterprise customer conversations, we celebrate diversity, empowering our team to forge impactful conversations globally. LivePerson is a place where uniqueness is embraced, growth is constant, and everyone is empowered to create their own success. And, we're very proud to have earned recognition from Fast Company, Newsweek, and BuiltIn for being a top innovative, beloved, and remote-friendly workplace. 

Belonging at LivePerson:

We are proud to be an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to age, ancestry, color, family or medical care leave, gender identity or expression, genetic information, marital status, medical condition, national origin, physical or mental disability, protected veteran status, race, religion, sex (including pregnancy), sexual orientation, or any other characteristic protected by applicable laws, regulations and ordinances. We also consider qualified applicants with criminal histories, consistent with applicable federal, state, and local law.

We are committed to the accessibility needs of applicants and employees. We provide reasonable accommodations to job applicants with physical or mental disabilities. Applicants with a disability who require reasonable accommodation for any part of the application or hiring process should inform their recruiting contact upon initial connection.



The talent acquisition team at LivePerson has recently been notified of a phishing scam targeting candidates applying for our open roles. Scammers have been posing as hiring managers and recruiters in an effort to access candidates' personal and financial information.  This phishing scam is not isolated to only LivePerson and has been documented in news articles and media outlets.Please note that any communication from our hiring teams at LivePerson regarding a job opportunity will only be made by a LivePerson employee with an @liveperson.com email address.

LivePerson does not ask for personal or financial information as part of our interview process, including but not limited to your social security number, online account passwords, credit card numbers, passport information and other related banking information. If you have any questions and or concerns, please feel free to contact [email protected]



Skills Required

  • Bachelor's degree in Computer Science, Engineering, or a related field, or equivalent practical experience
  • 7+ years of experience in SRE, DevOps, cloud infrastructure, systems engineering, or a related field
  • Experience at a Principal, Staff+, Architect, or equivalent senior technical leadership level
  • Experience designing and operating highly available, scalable, secure, and distributed production systems
  • Strong hands-on experience with Google Cloud Platform and cloud architecture, including networking, IAM, compute, storage, and managed services
  • Extensive experience with Kubernetes and containerization technologies such as Docker
  • Hands-on experience with service meshes such as Istio
  • Strong programming and scripting experience with Python, Bash, or similar languages
  • Extensive experience with Terraform and automation or configuration management tools such as Ansible
  • Strong experience with GitOps, Helm, and FluxCD
  • Strong experience designing and maintaining CI/CD pipelines using GitLab CI/CD
  • Deep understanding of Linux, networking, DNS, load balancing, TLS/SSL, authentication, and security fundamentals
  • Strong experience with Prometheus, Grafana, Alertmanager, or comparable observability platforms
  • Deep understanding of SRE principles, distributed systems, scalability, fault tolerance, and performance engineering
  • Ability to influence architecture and technical decisions across multiple teams and organizational boundaries
  • Excellent communication skills and ability to mentor senior engineers and technical leaders
  • Experience operating large-scale B2B SaaS or globally distributed systems
  • Experience with secrets management and security platforms such as HashiCorp Vault
  • Experience with PostgreSQL and other distributed data services
  • Experience operating and modernizing hybrid cloud and on-premises infrastructure
  • Experience leading large-scale cloud migrations or infrastructure modernization initiatives
  • Experience designing internal developer platforms and self-service infrastructure capabilities
  • Experience with disaster recovery, business continuity, and resilience engineering
  • Experience defining organization-wide engineering standards and reliability frameworks
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: New York, New York
1,000 Employees
Year Founded: 1995

What We Do

Forget the AI hype. We’re living in the age of conversation. Authentic, ongoing conversations are what fuel relationships, earn loyalty, and ultimately, drive growth. From the dawn of chat and messaging to the conversational AI era, LivePerson has been connecting businesses and customers through conversation for nearly three decades. Our award-winning Conversational AI platform, Conversational Cloud®, is built using large language models fine-tuned by billions of real customer conversations. With safety and security guardrails designed for the world’s largest enterprises, you remain firmly in control of the conversation.

Why Work With Us

We build warmth into our work and our workplace by making all people feel seen, valued, and heard. Dream Big, Help Others, Pursue Expertise and Own It. These four company values guide our continued, holistic growth as individuals, as teams, and as a global organization with over 1,000 employees.

Gallery

Gallery

Similar Jobs

Menlo Security Inc. Logo Menlo Security Inc.

Infrastructure Engineer

Cloud • Security • Cybersecurity
Remote
27 Locations
312 Employees

Mondelēz International Logo Mondelēz International

Senior Director, Global Supply Chain Excellence Program & Focused Improvement Lead

Big Data • Food • Hardware • Machine Learning • Retail • Automation • Manufacturing
Remote or Hybrid
29 Locations
90000 Employees
174K-287K Annually

Drata Logo Drata

Enterprise Account Executive

Security • Software • Cybersecurity • Automation
Remote
26 Locations
600 Employees
194K-273K Annually

Pfizer Logo Pfizer

Machine Learning Engineer

Artificial Intelligence • Healthtech • Machine Learning • Natural Language Processing • Biotech • Pharmaceutical
In-Office or Remote
32 Locations
121990 Employees
163K-272K Annually

Similar Companies Hiring

Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Revel.io Thumbnail
Aerospace • Hardware • Robotics • Software
US
50 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software • Productivity
US
15 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account