Site Reliability Engineer

Posted 19 Days Ago
Be an Early Applicant
Trondheim, Trøndelag, NOR
In-Office
Mid level
Artificial Intelligence • Software
The Role
Operate and improve globally distributed Vespa Cloud production systems. Responsibilities include reliability engineering, incident response, blameless postmortems, automation, observability, SLO/SLI tracking, proactive alerting, capacity planning, remediation, and collaboration on scalable, operable features. The role participates in a 24x7 on-call rotation and focuses on solving operational problems through code across a multi-cloud environment.
Summary Generated by Built In

Does it sound interesting to work on an open source platform managing the data and real-time search and inference for some of the largest companies in the world? Would you thrive on keeping large, globally distributed systems reliable, fast, and observable — and on writing the code that lets a small team operate at massive scale? If so, we want you to join our team at Vespa.ai as a Site Reliability Engineer!

About Vespa.ai:

Vespa.ai is a team of passionate builders. We maintain and develop the Apache 2.0 licensed open-source AI search platform Vespa.

Vespa is a fully featured search engine and vector database. It supports vector search (ANN), lexical search, and structured data search, all in a single query. Integrated machine-learning model inference enables the application of AI to make sense of data in real time. Together with Vespa's proven scalability and high availability, this empowers to create production-ready search applications at any scale and with any combination of features. Our users and customers are #1 in e-commerce, content, and financial services globally, and are powering companies such as Perplexity, Spotify, Yahoo, Wix, and many more.

In addition to our open-source platform, Vespa.ai develops and runs Vespa Cloud, a robust SaaS offering that allows businesses to harness the power of our technology with ease.

At Vespa.ai, we are extremely focused on automating everything we do to grow fast and maintain high quality. In all roles, we scale through technology, not simply by adding larger teams. We take pride in being small, nimble, and the most productive.

Position overview

At Vespa.ai, we embrace DevOps as a company culture, seeking to solve technical problems with automation and code rather than repetitive manual effort. For our Vespa Cloud production systems, we have had this mindset from day one.

We are seeking a Site Reliability Engineer to join the engineers who operate and improve Vespa Cloud. This is a development role first: we are looking for a strong engineer who, when faced with an operational problem, chooses to fix it in code so it never comes back. You will work across a large multi-cloud fleet on observability, automation and incident response. You will also participate in our 24x7 on-call rotation, approximately every third to fourth week.

At our Trondheim office, we work office-first: you will be based on-site most of the time, with the flexibility to work from home/remotely when needed, as agreed with your manager.



Responsibilities

  • Help ensure the reliability, availability, and performance of Vespa Cloud production systems running globally at scale.
  • Participate in a 24x7 on-call rotation (approximately every 3rd–4th week), respond to incidents, and drive blameless postmortems through to durable fixes.
  • Eliminate operational toil by solving problems with automation and code rather than manual effort.
  • Build and improve observability — metrics, logging, and tracing — across a large fleet.
  • Help define and track SLOs/SLIs, and build proactive alerting, capacity planning, and remediation.
  • Work with the rest of the Vespa.ai development team on reliability, scalability, and operability of new features.

Qualifications

  • Solid programming skills in Java, Python, Go, or similar languages, and a strong preference for solving problems in code.
  • Experience with cloud platforms (AWS, Azure, or GCP).
  • Solid understanding of networking, operating systems, distributed systems, and security principles.
  • Incident management and on-call experience.
  • Good understanding of sound software engineering principles and practices.
  • Excellent problem-solving and analytical skills.
  • Ability to work independently and as part of a team.
  • Fluent written and spoken English. Norwegian is not required.

Desired Skills

  • 4+ years building and operating large-scale production systems.
  • Experience with Infrastructure as Code tools such as Terraform, Tofu, Spacelift, etc.
  • Familiarity with observability stacks (Prometheus, Grafana, OpenTelemetry, ELK).
  • Experience with CI/CD tooling such as GitHub Actions, Buildkite, etc.
  • Experience operating data-intensive or stateful systems at scale.

Some of Our Tools and Services

  • JumpCloud, Google Workspace, and Slack
  • GitHub Enterprise Cloud (including GitHub Actions)
  • Jira Cloud and Jira Service Desk
  • StrongDM, Grafana, Spacelift, and Buildkite
  • AWS, GCP, and Azure

Why Join Us:

  • Opportunities for professional growth and development as part of one of Europe's most exciting start-ups!
  • Be part of a cutting-edge team working on innovative search and recommendation technology.
  • Work on a team where we don't believe in silos between engineers; there aren't "developers", "ops people", and "sysadmins". We're all engineers solving problems the smart way together!
  • Competitive salary and benefits.
  • Relocating to Norway? We'll help you settle in, and offer voluntary Norwegian language training.

Skills Required

  • Solid programming skills in Java, Python, Go, or similar languages
  • Experience with cloud platforms such as AWS, Azure, or GCP
  • Understanding of networking, operating systems, distributed systems, and security principles
  • Incident management and on-call experience
  • Understanding of software engineering principles and practices
  • Excellent problem-solving and analytical skills
  • Ability to work independently and as part of a team
  • Fluent written and spoken English
  • 4+ years building and operating large-scale production systems
  • Experience with Infrastructure as Code tools such as Terraform, Tofu, or Spacelift
  • Familiarity with observability stacks such as Prometheus, Grafana, OpenTelemetry, or ELK
  • Experience with CI/CD tooling such as GitHub Actions or Buildkite
  • Experience operating data-intensive or stateful systems at scale
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Trondheim
Year Founded: 2023

What We Do

Vespa.ai operates Vespa Cloud - used by companies to run Big Data serving with AI, online. We maintain the Vespa open-source project, continuously released and used by organizations with high performance, availability, and functional requirements. We are hiring! See the Jobs page, or visit our website.

Similar Jobs

ScorePlay Logo ScorePlay

Senior Platform Engineer

Artificial Intelligence • Digital Media • Software • Sports
In-Office or Remote
27 Locations
70 Employees

Nebius Logo Nebius

Site Reliability Engineer

Artificial Intelligence • Information Technology • Consulting
In-Office or Remote
27 Locations
473 Employees

Nebius Logo Nebius

Senior Site Reliability Engineer

Artificial Intelligence • Information Technology • Consulting
In-Office or Remote
27 Locations
473 Employees

Vespa.ai Logo Vespa.ai

Site Reliability Engineer

Artificial Intelligence • Software
In-Office
Trondheim, Trøndelag, NOR

Similar Companies Hiring

Revel Thumbnail
Aerospace • Hardware • Robotics • Software
Marina Del Rey, California
60 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software
New York, New York
30 Employees
Vega Thumbnail
Artificial Intelligence • Automotive • Insurance • Transportation
US
43 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account