Site Reliability Engineer, K8s (Remote International)

Posted Yesterday
Be an Early Applicant
Hiring Remotely in United Kingdom
Remote
Mid level
Healthtech
The Role
Design, build, and operate a large-scale Kubernetes platform supporting distributed systems, data infrastructure, and business-critical services. Responsibilities include Kubernetes lifecycle management, reliability and observability, incident response, infrastructure automation, GitOps workflows, networking, platform security, developer self-service, and bare-metal operations. The role requires hands-on architecture, operational excellence, and technical leadership across infrastructure platforms.
Summary Generated by Built In
WebMD and its affiliates is an Equal Opportunity/Affirmative Action employer and does not discriminate on the basis of race, ancestry, color, religion, sex, gender, age, marital status, sexual orientation, gender identity, national origin, medical condition, disability, veterans status, or any other basis protected by law. 
Build the platform behind PulsePoint

PulsePoint operates large-scale data and advertising platforms that power business-critical services used every day across the company.

Our Platform Engineering team owns the foundation that enables engineering teams to move quickly and safely. We build and operate the infrastructure that supports Kubernetes workloads, data platforms, developer tooling, observability and production operations at scale.

Unlike many cloud-only environments, we own the full lifecycle of the infrastructure - from bare-metal hardware and networking to Kubernetes, observability and developer experience.

The environment supports large-scale Kubernetes workloads, multi-petabyte data systems and business-critical services used across multiple engineering organizations.

We're looking for an experienced engineer to help shape its next stage of growth.

This is a hands-on role focused on architecture, reliability, automation and operational excellence. You will work on complex distributed systems, drive platform improvements and help shape the technical direction of infrastructure across the company.

What you'll work on

You'll help design, build and operate the Kubernetes platform used across PulsePoint.

Examples of challenges you may work on include:

  • Platform architecture and Kubernetes lifecycle management

  • Reliability, observability and incident response

  • Infrastructure automation and GitOps workflows

  • Networking, service connectivity and platform security

  • Developer experience and self-service platform capabilities

  • Large-scale distributed systems running on bare-metal infrastructure

Technology

You'll work in an environment that includes:

  • Kubernetes and platform services

  • Multi-petabyte data infrastructure

  • Bare-metal and cloud environments

  • GitOps and infrastructure automation

  • Modern observability and reliability engineering practices

Technologies commonly used across the environment include Kubernetes, ArgoCD, Puppet, Terraform, OpenTelemetry, Prometheus, Alertmanager, Kafka, Redis and Ceph.

Experience with every technology is not required.

Who we’re looking for

Success in this role is not measured by the number of tickets closed or clusters operated. Success means building platform capabilities that make engineering teams more reliable, productive and autonomous.

We're especially interested in engineers who:

  • Have operated production infrastructure at meaningful scale

  • Understand how distributed systems fail and recover

  • Prefer automation over repetitive operational work

  • Enjoy simplifying systems rather than adding complexity

  • Take ownership beyond the boundaries of a single component

  • Willing and able to work 9am-6pm ET U.S. hours. You can work fully remotely
Why this role

This is not a ticket-driven operational role.

You'll help define platform architecture, influence engineering standards and work on infrastructure that supports multiple engineering organizations.

The engineer joining this role is expected to become a key technical contributor shaping the future of the platform.

We try to keep the process focused and practical.

  1. Introductory conversation (~60m)
    Learn about your background and discuss the role.

  2. Technical discussion (~60m)
    Deep dive into systems engineering, Kubernetes and operational experience.

  3. Architecture discussion (~60m)
    Explore platform design, distributed systems and technical decision making.

  4. Leadership conversation (~30m)
    Meet engineering leadership and discuss team, strategy and long-term direction.

Skills Required

  • Experience operating production infrastructure at meaningful scale
  • Understanding of distributed systems failure and recovery
  • Experience with automation and infrastructure operations
  • Willingness and ability to work 9:00 a.m. to 6:00 p.m. U.S. Eastern Time
  • Experience with every listed technology is not required
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: New York, NY
2,062 Employees
Year Founded: 1995

What We Do

Our Company: WebMD is a leading provider of health information services to consumers, physicians, healthcare professionals, employers and health plans. Our Business: WebMD is the leading provider of health information and services to consumers and healthcare professionals. The online healthcare information, decision-support applications and communications services that we provide: help consumers take an active role in managing their health by providing objective healthcare information and lifestyle information. make it easier for physicians and healthcare professionals to access clinical reference sources, stay abreast of the latest clinical information, learn about new treatment options, earn continuing medical education credits and communicate with peers. enable employers and health plans to provide their employees and plan members with access to personalized heath and benefit information and decision support technology that helps them make informed benefit, provider and treatment choices.

Similar Jobs

Remote
United Kingdom
25 Employees

Rubrik Logo Rubrik

Senior Talent Partner EMEA

Artificial Intelligence • Big Data • Cloud • Information Technology • Software • Cybersecurity • Data Privacy
In-Office or Remote
London, Greater London, England, GBR
3000 Employees

Samsara Logo Samsara

Senior Software Engineer

Artificial Intelligence • Cloud • Computer Vision • Hardware • Internet of Things • Software
Easy Apply
Remote or Hybrid
UK
4000 Employees
Remote or Hybrid
2 Locations
289097 Employees

Similar Companies Hiring

Sailor Health Thumbnail
Healthtech • Social Impact • Telehealth
New York City, NY
20 Employees
Granted Thumbnail
Artificial Intelligence • Healthtech • Insurance • Mobile • Financial Services
New York, New York
23 Employees
OneImaging Thumbnail
Healthtech
Miami, FL
62 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account