Site Reliability Engineer- Spacetime UK

Posted 19 Days Ago
Be an Early Applicant
London, Greater London, England, GBR
In-Office
Mid level
Aerospace • Manufacturing
Connectivity Everywhere
The Role
Design and build a centralized observability platform for satellite and deep-space networks. Implement metrics, logging, and tracing (Prometheus, Loki, Tempo/OpenTelemetry), define SLOs/SLIs/error budgets, automate deployments with IaC and GitOps, integrate observability into Kubernetes and cloud environments, support engineers with instrumentation and tooling, lead monitoring/alerting/incident response, and participate in on-call rotations.
Summary Generated by Built In
About Aalyria:

Aalyria is a leading technology company that supplies laser communications technology and temporospatial software-defined networking platforms to the aerospace industry. With technology acquired from Google, Aalyria is at the forefront of innovation in satellite and airborne mesh networks, as well as cislunar and deep-space communications. We are revolutionizing the orchestration and management of planetary mesh networks using any radio or optical spectrum, any orbit, and any hardware across land, sea, air, and space.

Role Overview:

This isn't a "keep the lights on" SRE role. This is a strategic, high-impact opportunity to build the nervous system for a platform that transforms how networks of satellites, ground stations, and fleets are interconnected and orchestrated. You will be building the core observability stack that ensures the reliability of systems critical to the operation of satellite megaconstellations and missions to deep space.

This is a greenfield/brownfield opportunity. You will be a trusted expert, helping to define and implement the strategy and building the tools that empower our engineers. You will support the roadmap to mature our observability stack, moving from cloud-native tools to a robust, scalable, and insightful platform built on best-in-class technologies (Prometheus, OpenTelemetry, etc.). If you are an SRE who thrives on platform-building challenges and wants to be relied upon to build a production-grade observability stack from the ground up, this role is for you.

Note: this role includes on-call responsibilities.

Key Responsibilities:
  • Help design and build Aalyria's centralized observability platform, integrating and scaling tools for metrics (e.g. Prometheus), logging (e.g. Loki), and distributed tracing (e.g. Tempo/OpenTelemetry).
  • Define, implement, and manage a robust framework of Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets for our core products, ensuring we are launch-ready.
  • Partner with SWEs to implement observability best practices, develop standard templates and documentation, and configure tooling (e.g., OpenTelemetry libraries).
  • Automate the deployment, scaling, and management of the entire observability stack using Infrastructure as Code (e.g. Terraform) and GitOps principles (e.g. ArgoCD).
  • Partner closely with the core infrastructure team to ensure deep visibility into our Kubernetes clusters and underlying GCP and AWS environments.
  • Develop and lead the company's monitoring, alerting, and incident response strategy, driving a culture of proactive reliability and blameless post-mortems.
Required Qualifications:
  • 4+ years of experience in an SRE or platform engineering role, with a focus on observability for large-scale, distributed compute or network systems.
  • Deep, hands-on expertise building, scaling, and managing observability platforms (e.g., Prometheus, Grafana, Loki/ELK, OpenTelemetry, Tempo/Jaeger, Honeycomb, etc.). You have proven experience using these tools to support performance analysis and debugging of complex distributed systems.
  • Strong production-level experience with Google Cloud Platform (GCP) and Kubernetes.
  • Experience using Infrastructure as Code (IaC) and GitOps principles (e.g., ArgoCD).
  • Proficiency in a systems programming language, with a strong preference for Go and Python for debugging and writing tooling.
  • Demonstrable experience defining, implementing, and managing SLOs, SLIs, and error budgets for production services for high availability distributed systems.
Preferred Qualifications:
  • Experience operating a multi-cloud environment, specifically GCP and AWS.
  • Hands-on experience with GitLab CI for CI/CD pipelines.
  • Working knowledge of service mesh technologies such as Istio or Linkerd.
  • Familiarity with instrumenting applications written in Go and C++.
  • An active Secret clearance, or higher, is preferred for this position.
  • Experience with JVM observability (tuning, monitoring) for Java-based applications.
What We Offer:
  • Competitive salary benchmarked to UK aerospace and defence technology market rates
  • Equity participation — share in Aalyria's growth at an early stage
  • Comprehensive benefits including pension, private health insurance, and generous annual leave
  • Flexible and hybrid working arrangements
  • The opportunity to work on genuinely novel technology with real-world operational impact across national security, commercial satellite, and deep-space programmes
  • A collaborative, low-hierarchy team environment with direct exposure to technical leadership and customers
Equal Opportunity Employer Statement:

Aalyria Technologies is an equal opportunities employer. We are committed to building an inclusive workplace and welcome applicants from all backgrounds. We do not discriminate on the basis of race, religion, gender, sexual orientation, age, disability, or any other protected characteristic under UK law.

Aalyria Technologies operates in sectors subject to UK export control regulations. Candidates may be asked to confirm their eligibility to access export-controlled technology as part of the hiring process.


#LI-Remote

Skills Required

  • 4+ years in an SRE or platform engineering role focused on observability for large-scale distributed systems
  • Hands-on expertise with observability platforms (Prometheus, Grafana, Loki/ELK, OpenTelemetry, Tempo/Jaeger, Honeycomb)
  • Production-level experience with Google Cloud Platform (GCP) and Kubernetes
  • Experience using Infrastructure as Code (e.g., Terraform) and GitOps principles (e.g., ArgoCD)
  • Proficiency in a systems programming language; strong preference for Go and Python
  • Demonstrable experience defining, implementing, and managing SLOs, SLIs, and error budgets
  • On-call responsibilities
  • Experience operating a multi-cloud environment, specifically GCP and AWS
  • Hands-on experience with GitLab CI for CI/CD pipelines
  • Working knowledge of service mesh technologies such as Istio or Linkerd
  • Familiarity with instrumenting applications written in Go and C++
  • An active Secret clearance, or higher
  • Experience with JVM observability (tuning, monitoring) for Java-based applications
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Livermore, CA
83 Employees

What We Do

Aalyria is creating, organizing and managing the world’s most advanced networks to enable connectivity everywhere at the speed of discovery. Aalyria brings together two technologies originally developed at Alphabet as part of its wireless connectivity efforts: atmospheric laser communications technology and a software platform for orchestrating networks across land, sea, air, space and beyond. It is backed by leading Silicon Valley investors including the founders of Accel, J2 Ventures and Housatonic.

Similar Jobs

iManage Logo iManage

Senior Site Reliability Engineer

Artificial Intelligence • Cloud • Information Technology • Legal Tech • Productivity • Software
Hybrid
London, Greater London, England, GBR
1100 Employees
In-Office
Nottingham, Nottinghamshire, England, GBR
15967 Employees
In-Office
London, Greater London, England, GBR
22000 Employees

FNZ Group Logo FNZ Group

Site Reliability Engineer

Fintech • Payments • Financial Services
In-Office
2 Locations
4252 Employees

Similar Companies Hiring

Fortune Brands Innovations Thumbnail
Manufacturing
Deerfield, IL
10000 Employees
Amalgamated Sugar Thumbnail
Food • Greentech • Agriculture • Industrial • Manufacturing
Boise, Idaho
768 Employees
Outpost Space Thumbnail
Aerospace • Defense
US
24 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account