Site Reliability Engineer

vespaai· Engineering
Apply Now ↗
📍 Trondheim, NorwayFull Time

About this role

Does it sound interesting to work on an open source platform managing the data and real-time search and inference for some of the largest companies in the world? Would you thrive on keeping large, globally distributed systems reliable, fast, and observable — and on writing the code that lets a small team operate at massive scale? If so, we want you to join our team at Vespa.ai as a Site Reliability Engineer!


About Vespa.ai:

Vespa.ai is a team of passionate builders. We maintain and develop the Apache 2.0 licensed open-source AI search platform Vespa.

Vespa is a fully featured search engine and vector database. It supports vector search (ANN), lexical search, and structured data search, all in a single query. Integrated machine-learning model inference enables the application of AI to make sense of data in real time. Together with Vespa's proven scalability and high availability, this empowers to create production-ready search applications at any scale and with any combination of features. Our users and customers are #1 in e-commerce, content, and financial services globally, and are powering companies such as Perplexity, Spotify, Yahoo, Wix, and many more.

In addition to our open-source platform, Vespa.ai develops and runs Vespa Cloud, a robust SaaS offering that allows businesses to harness the power of our technology with ease.

At Vespa.ai, we are extremely focused on automating everything we do to grow fast and maintain high quality. In all roles, we scale through technology, not simply by adding larger teams. We take pride in being small, nimble, and the most productive.

Position overview

At Vespa.ai, we embrace DevOps as a company culture, seeking to solve technical problems with automation and code rather than repetitive manual effort. For our Vespa Cloud production systems, we have had this mindset from day one.

We are seeking a Site Reliability Engineer to join the engineers who operate and improve Vespa Cloud. This is a development role first: we are looking for a strong engineer who, when faced with an operational problem, chooses to fix it in code so it never comes back. You will work across a large multi-cloud fleet on observability, automation and incident response. You will also participate in our 24x7 on-call rotation, approximately every third to fourth week.

At our Trondheim office, we work office-first: you will be based on-site most of the time, with the flexibility to work from home/remotely when needed, as agreed with your manager.



Responsibilities

  • Help ensure the reliability, availability, and performance of Vespa Cloud production systems running globally at scale.
  • Participate in a 24x7 on-call rotation (approximately every 3rd–4th week), respond to incidents, and drive blameless postmortems through to durable fixes.
  • Eliminate operational toil by solving problems with automation and code rather than manual effort.
  • Build and improve observability — metrics, logging, and tracing — across a large fleet.
  • Help define and track SLOs/SLIs, and build proactive alerting, capacity planning, and remediation.
  • Work with the rest of the Vespa.ai development team on reliability, scalability, and operability of new features.

Qualifications

  • Solid programming skills in Java, Python, Go, or similar languages, and a strong preference for solving problems in code.
  • Experience with cloud platforms (AWS, Azure, or GCP).
  • Solid understanding of networking, operating systems, distributed systems, and security principles.
  • Incident management and on-call experience.
  • Good understanding of sound software engineering principles and practices.
  • Excellent problem-solving and analytical skills.
  • Ability to work independently and as part of a team.
  • Fluent written and spoken English. Norwegian is not required.

Desired Skills

  • 4+ years building and operating large-scale production systems.
  • Experience with Infrastructure as Code tools such as Terraform, Tofu, Spacelift, etc.
  • Familiarity with observability stacks (Prometheus, Grafana, OpenTelemetry, ELK).
  • Experience with CI/CD tooling such as GitHub Actions, Buildkite, etc.
  • Experience operating data-intensive or stateful systems at scale.

Some of Our Tools and Services

  • JumpCloud, Google Workspace, and Slack
  • GitHub Enterprise Cloud (including GitHub Actions)
  • Jira Cloud and Jira Service Desk
  • StrongDM, Grafana, Spacelift, and Buildkite
  • AWS, GCP, and Azure

Why Join Us:

  • Opportunities for professional growth and development as part of one of Europe's most exciting start-ups!
  • Be part of a cutting-edge team working on innovative search and recommendation technology.
  • Work on a team where we don't believe in silos between engineers; there aren't "developers", "ops people", and "sysadmins". We're all engineers solving problems the smart way together!
  • Competitive salary and benefits.
  • Relocating to Norway? We'll help you settle in, and offer voluntary Norwegian language training.

Frequently Asked Questions

Is the salary disclosed for the Site Reliability Engineer position at vespaai?
The salary for this Site Reliability Engineer role at vespaai is not publicly listed. Click "Apply Now" to learn more about the compensation package on their official careers page.
Where is the Site Reliability Engineer position at vespaai located?
This Site Reliability Engineer role at vespaai is based in Trondheim, Norway. The position is listed as on-site or hybrid. Check the full job description or apply directly to confirm the work arrangement.
Is the Site Reliability Engineer role at vespaai full-time or part-time?
This is listed as a Full Time position. It is posted as a Site Reliability Engineer role in the Engineering department at vespaai.
Which team or department does the Site Reliability Engineer at vespaai belong to?
This Site Reliability Engineer position is part of the Engineering department at vespaai. See the full job description for more information about the team structure and responsibilities.
How do I apply for the Site Reliability Engineer position at vespaai?
Click the "Apply Now" button on this page. You will be redirected to vespaai's official application portal hosted on bamboohr where you can submit your application directly.
When was the Site Reliability Engineer job at vespaai posted?
This Site Reliability Engineer position at vespaai was posted on Sep 9, 2026. Apply as soon as possible — early applications are often reviewed first.
Site Reliability Engineer
vespaai
Apply for this role ↗

You'll be redirected to vespaai's official application page on bamboohr.