Using KV cache as embeddings BreadBowl-Embed, released under Apache-2.0 with open weights, introduces a late-interaction embedding architecture that stores every passage as 16 routing-value slots rather than one vector per document or one per token. Each slot holds a 256-dimensional routing vector that is indexed and searched plus a 256-dimensional value vector that a query reads after routing, letting retrieval and reranking share one representation so documents are encoded once and reranking never re-reads their text. The author estimates roughly 80% of work in their own coding-agent sessions goes to finding the right context, the cost BreadBowl-Embed targets by avoiding the per-query full transformer pass a cross-encoder reranker requires. Introducing BreadBowl-Embed · Open weights · Apache-2.0 One vector is too few . One per token is too many . BreadBowl-Embed is a new late-interaction architecture for embeddings. Instead of one vector per document, or one per token, it stores every passage as 16 routing–value slots. The routing vectors find candidates in an index; then the same vectors decide which stored values each query reads. Retrieval and reranking share one representation: documents are encoded once, and reranking never re-reads their text. 01 — The problem High-precision retrieval reads your documents twice. Once to find them, and again to decide which ones actually matter. Modern search and RAG pipelines run in two stages. First, a bi-encoder compresses every document into a single vector, so a nearest-neighbor index can search millions of them in milliseconds. Then, because one vector throws away the details that decide relevance the date, the exception, who did what to whom , a cross-encoder re-reads the top candidates next to the query and scores them again. That second stage is where the precision comes from, and it is expensive in a very specific way: the reranker runs a full transformer pass over every query, candidate pair, every time a query arrives. Most of that work can't be cached ahead of time, because it depends on the query. Agents make this worse. A coding or research agent doesn't issue one query per task; it issues dozens, and each one pays the reranking bill again. In my own coding-agent sessions, I'd estimate roughly 80% of the work is finding the right context: searching, reading and deciding what matters. Finding context is becoming the inner loop of AI work. So we asked a simple question: how much of the reranker's job can the retrieval representation do by itself , using only what was computed and stored before the query arrived? 02 — The representation spectrum One vector is too few. One per token is too many. There have been two classic answers to the question how should a document be stored for search? One vector per document. Bi-encoders such as Qwen3-Embedding, E5 or EmbeddingGemma pool the whole text into a single point. It's cheap to store and fast to search. But every query is compared against the same summary: a question about the date and a question about the architect hit the same point in space. The score is one dot product, so there is nothing left to refine. That's why single vectors so often get paired with a reranker. One vector per token. Late-interaction models like ColBERT keep a contextual vector for every token and let each query token find its best match MaxSim . That preserves detail and ranks well. The costs come with it: the index grows with every token stored, scoring cost grows with both query and document length, and all the scorer can do with a stored token vector is measure its similarity to a query token. Each query token takes its single best match; there is no separate content to read . BreadBowl-Embed sits in between on purpose. Every passage gets a fixed budget of 16 slots, however long it is. Each slot carries two vectors with two different jobs: - a routing vector 256-d that says where to look . It is what gets indexed and searched. - a value vector 256-d that holds what to read once a query has decided where to look. What do you actually need from a representation? Here's Table 1 from our paper, made interactive. Switch on the requirements one at a time and watch which designs survive. | Architecture | Retriever | Reranker | Extra backbone pass | Separate value readout | Stored doc. vectors | Scoring cost | |---|---|---|---|---|---|---| | Bi-encoder | ✓ | ✓