Vespa.ai’s Personalized Search: Advanced Ranking & Tensor framework

📺 Click to view the youtube embed player and accept their cookies.
Session Abstract

Modern search demands scalable personalization. Discover Vespa’s multi-stage ranking and tensor framework for hybrid queries, multimodal retrieval and real-time ML Learn how to deploy low-latency, high-relevance search systems at petabyte scale.

Session Description

Today’s applications require search engines to unify text, vectors, and business logic with millisecond latency at petabyte scale. It’s not easy to balance speed, relevance, and personalization for a large user population and a billion scale item base. Vespa.ai, the open-source engine powering Yahoo, Perplexity, Qwant, Vinted, Spotify addresses this through multi-stage ranking with close to data tensor operations and easy to understand custom functions.
Vespa’s phased architecture enables high performance due to the ability to filter candidates via hybrid retrieval (text + multi vector + filters) before applying ML models for precision or logic for personalisation. Its tensor framework enables multimodal (text/image/video) and multivector queries with real-time individual personalization, scaling beyond 100k QPS with milliseconds latency.
You will learn Vespa.ai configuration concepts and ideas how all the building blocks (LLMs, VLMs, embedding models, sparse and dense representations for items and users) can be connected together.


This session is sponsored by Vespa.ai.

Palais Atelier
17.Jun 2025
11:10am - 11:50am
Talk