Skip to content

Archive

Retrieval

19 articles
Artificial Intelligence 22 Sep 2026 6 min read

Cosine Similarity Discards Embedding Magnitude

Two embedding vectors can point in the same direction while having very different norms. Cosine similarity gives those vectors the same directional score. A raw dot product does not. That distinction becomes an implementation boundary when a retrieval system changes index metrics, normalizes vectors at ingestion, or mixes embeddings produced by different pipelines. The issue is not that one metric is universally preferable. The relevant question is whether vector magnitude carries information that the scoring contract intends to preserve. Once vectors are normalized to unit length, that information is removed from the similarity calculation.

Artificial Intelligence 16 Sep 2026 5 min read

Version Embedding Spaces as Incompatible Interfaces

Two embedding models can emit vectors with the same number of dimensions and still produce similarity scores that have no useful cross-version meaning. A vector database accepts the shapes, the distance function runs normally, and retrieval returns ranked results. Nothing in that execution path proves that query and document vectors occupy a compatible representation space. This makes embedding model identity part of the retrieval interface. Replacing an encoder is not equivalent to swapping a serialization routine. Unless compatibility is explicitly established, vectors produced by separate model versions should be treated as belonging to separate spaces.

Artificial Intelligence 16 Sep 2026 5 min read

Quantize Embeddings Without Hiding Retrieval Error

Embedding quantization replaces higher-precision vector values with a smaller representation. The storage reduction is easy to measure. The retrieval effect is less direct: a small numeric error can be harmless for one query and change the candidate order for another when several similarity scores are close. That makes quantization a ranking concern, not only a storage format choice. The relevant question is how the compressed representation changes the comparisons used to select neighbors.

Artificial Intelligence 16 Sep 2026 5 min read

Measure Embedding Anisotropy Before Vector Search

Cosine similarity assumes that vector direction carries useful discrimination. That assumption becomes less informative when many embeddings occupy a narrow region of the space. In that case, unrelated items can share a substantial directional component, cosine scores can cluster into a compressed range, and small residual differences can decide the ranking. This geometric pattern is often described as embedding anisotropy. It exists before a vector index chooses candidates, so index tuning alone cannot establish whether the representation has enough angular separation for the retrieval task.

Artificial Intelligence 15 Sep 2026 6 min read

Treat Embedding Model Changes as Index Migrations

An embedding model change does not merely replace a function that emits arrays of the same length. It can change the coordinate system in which stored items and incoming queries are represented. Even when two models produce vectors with identical dimensions, their coordinates and similarity-score distributions are not interchangeable by default. That makes an embedding model version part of the index schema. A retrieval system that changes the query encoder while retaining vectors produced by an older encoder can still return numeric scores, but those scores no longer have a justified geometric interpretation unless cross-version compatibility is an explicit property of the models.

Artificial Intelligence 15 Sep 2026 5 min read

Account for Hubness in Embedding Retrieval

An embedding index can return the same few items for many unrelated queries. Their similarity scores may look ordinary, and the nearest-neighbor algorithm may be operating correctly. The distortion can come from the geometry of the representation itself: some vectors become neighbors of unusually many other vectors. This effect is commonly called hubness. Hubness matters because nearest-neighbor retrieval is usually interpreted locally. A query asks which stored vectors are closest to it, but the index does not normally expose how often each candidate also appears near other queries. A candidate that repeatedly occupies neighbor lists can receive more retrieval opportunities than its semantic relevance warrants.

Artificial Intelligence 13 Sep 2026 7 min read

Merge Retrieval Rankings with Reciprocal Rank Fusion

A lexical retriever and an embedding retriever can return useful results for the same query while assigning scores that have no common numerical meaning. Adding those raw scores treats incomparable scales as if they were calibrated measurements. Reciprocal rank fusion avoids that assumption by combining positions rather than score magnitudes. This makes RRF useful in retrieval-augmented generation systems that mix distinct retrieval signals. Each retriever keeps its own scoring model. The fusion layer only needs ordered result lists and stable document identities.

Artificial Intelligence 13 Sep 2026 7 min read

Diversify Retrieval Results with Maximum Marginal Relevance

A retriever can fill its top positions with passages that are individually relevant but nearly interchangeable. Several chunks from one document may repeat the same fact, leaving little room for other evidence in a fixed context budget. Maximum marginal relevance, commonly abbreviated MMR, addresses this at the selection stage by considering both query relevance and redundancy with items already chosen. MMR does not change the embedding model or recover candidates that retrieval missed. It reranks a candidate pool. That boundary matters: the method can improve variety among available candidates, but it cannot compensate for poor candidate recall.

Artificial Intelligence 13 Sep 2026 7 min read

Diagnose Embedding Anisotropy in Vector Retrieval

A vector retriever can return stable nearest neighbors even when its embedding space uses only a narrow set of directions. In that case, high cosine similarity may reflect shared global structure as well as query-specific semantic alignment. This geometric concentration is commonly described as embedding anisotropy. Anisotropy matters at the retrieval boundary because nearest-neighbor search operates on the geometry it receives. An index can reproduce cosine or inner-product rankings correctly while those rankings still have weak separation between relevant and irrelevant candidates. Treating every retrieval issue as an indexing problem can therefore hide a representation problem upstream.

Artificial Intelligence 12 Sep 2026 7 min read

Hubness Can Distort Nearest-Neighbor Embedding Retrieval

Embedding retrieval usually treats each query independently: encode the query, compare it with stored vectors, then return the closest items. That local view can miss a collection-level pattern. Some stored vectors may appear in the nearest-neighbor lists of many unrelated queries far more often than other vectors. This pattern is called hubness. A hub is not necessarily a broadly relevant item. It is a vector that becomes a neighbor unusually often under the representation and distance geometry in use. For developers, the distinction matters because a retrieval pipeline can compute cosine similarity or Euclidean distance exactly as specified and still produce systematically repetitive candidates.

Artificial Intelligence 12 Sep 2026 7 min read

Embedding Anisotropy Can Compress Cosine Score Separation

Two embedding vectors can have a high cosine similarity even when the items they represent are not close in the task-specific sense a retrieval system needs. One source of this mismatch is embedding anisotropy: vectors are distributed unevenly across representation space, often with substantial mass concentrated around shared directions. Cosine similarity removes vector magnitude from the comparison, but it does not remove a common directional component. If many vectors point partly in the same direction, unrelated pairs can start from an elevated cosine baseline. The useful distinction between relevant and irrelevant items then has to appear within a narrower score range.

Artificial Intelligence 10 Sep 2026 10 min read

Use Late Interaction for Fine-Grained Neural Retrieval

Use Late Interaction for Fine-Grained Neural Retrieval A single text embedding is convenient: encode a query into one vector, encode each document into one vector, then rank documents by vector similarity. That design scales well, but compression happens early. A paragraph containing several distinct ideas must squeeze all of them into one fixed-size representation before the query arrives. Late interaction keeps more of that detail. Instead of representing each text with only one vector, it retains multiple contextual token vectors and compares them at retrieval time. The document can still be encoded ahead of time, but the final relevance score is computed from fine-grained query-to-document matches.

Artificial Intelligence 08 Sep 2026 13 min read

Rewrite RAG Queries Without Losing User Intent

A retrieval-augmented generation system often searches with the user’s latest message. That works for self-contained questions, but conversational questions frequently depend on earlier turns. Consider a support assistant. The user first asks about a failed database migration, discusses PostgreSQL for several turns, and then asks: Does the rollback command work on version 16 too? Searching that sentence literally may retrieve pages about unrelated rollback commands because the query does not say what is being rolled back. A query rewriter can turn the conversational message into a self-contained retrieval query such as:

Artificial Intelligence 07 Sep 2026 11 min read

Diversify RAG Retrieval with Maximal Marginal Relevance

A retrieval-augmented generation (RAG) system can retrieve highly relevant passages and still build a poor context. The problem appears when the top results repeat the same fact in slightly different wording. Suppose five retrieved chunks all explain how to reset an API token, while a lower-ranked chunk explains the permission change that must happen afterward. Filling the context window with the five near-duplicates gives the language model less useful evidence than selecting a smaller set that covers both parts of the task.

Artificial Intelligence 07 Sep 2026 11 min read

Diagnose Hubness in Embedding Retrieval

Embedding retrieval is usually explained one query at a time: encode the query, compare it with stored vectors, and return the nearest items. That view can hide a collection-level failure mode. A document may look reasonably similar to many unrelated queries and therefore appear in far more result lists than it should. Such an item is called a hub. The broader phenomenon, hubness, is a tendency for some points in a vector space to become nearest neighbors of unusually many other points. It matters because a retriever can have healthy-looking similarity scores while repeatedly wasting top positions on generic or geometrically favored items.

Artificial Intelligence 07 Sep 2026 11 min read

Chunking Documents for RAG Without Losing Context

A retrieval-augmented generation (RAG) system can have a strong embedding model and still retrieve poor evidence. One common reason is document chunking: the text was divided into units that are awkward to search or incomplete when read on their own. If chunks are too large, one embedding must represent several unrelated ideas and retrieval becomes less precise. If chunks are too small, the retrieved text may omit the definitions, qualifiers, or surrounding steps needed to answer correctly. The problem is therefore not to find one universally correct chunk size. It is to create retrieval units that are focused enough to match a query and complete enough to be useful after retrieval.

Artificial Intelligence 05 Sep 2026 9 min read

Diversify RAG Context with Maximum Marginal Relevance

A retrieval-augmented generation (RAG) system can retrieve highly relevant passages and still build poor context. The problem appears when several top results say almost the same thing. Sending all of them to the language model consumes context without adding much evidence, while a slightly lower-ranked passage containing a different useful fact may be excluded. Maximum marginal relevance (MMR) is a selection strategy for this situation. Instead of choosing passages only by their relevance to the query, MMR repeatedly chooses a passage that is both relevant and sufficiently different from passages already selected.

Artificial Intelligence 05 Sep 2026 10 min read

Combine Retrieval Rankings with Reciprocal Rank Fusion

A retrieval-augmented generation (RAG) system often has more than one useful way to find evidence. Keyword retrieval is good at exact names, identifiers, and rare terms. Embedding retrieval can find passages that express the same idea with different wording. Using both can improve candidate coverage, but it creates a practical problem: their scores usually do not mean the same thing. A keyword score of 12.4 and a cosine similarity of 0.81 cannot be safely averaged just because both are numbers. Their scales, distributions, and even direction conventions depend on the retrieval methods and implementations.

Artificial Intelligence 03 Sep 2026 7 min read

Improve RAG Retrieval with Reranking

Retrieval-augmented generation (RAG) depends on finding useful evidence before asking a language model to answer. A vector search can retrieve candidates quickly, but the nearest vectors are not always the passages that best answer the user’s question. Reranking adds a second relevance step. The system first retrieves a reasonably broad candidate set with a fast method, then applies a more precise model to reorder those candidates before selecting context for the LLM.