Skip to content

Archive

Embeddings

40 articles
Artificial Intelligence 07 Sep 2026 11 min read

Diversify RAG Retrieval with Maximal Marginal Relevance

A retrieval-augmented generation (RAG) system can retrieve highly relevant passages and still build a poor context. The problem appears when the top results repeat the same fact in slightly different wording. Suppose five retrieved chunks all explain how to reset an API token, while a lower-ranked chunk explains the permission change that must happen afterward. Filling the context window with the five near-duplicates gives the language model less useful evidence than selecting a smaller set that covers both parts of the task.

Artificial Intelligence 07 Sep 2026 11 min read

Diagnose Hubness in Embedding Retrieval

Embedding retrieval is usually explained one query at a time: encode the query, compare it with stored vectors, and return the nearest items. That view can hide a collection-level failure mode. A document may look reasonably similar to many unrelated queries and therefore appear in far more result lists than it should. Such an item is called a hub. The broader phenomenon, hubness, is a tendency for some points in a vector space to become nearest neighbors of unusually many other points. It matters because a retriever can have healthy-looking similarity scores while repeatedly wasting top positions on generic or geometrically favored items.

Artificial Intelligence 07 Sep 2026 11 min read

Chunking Documents for RAG Without Losing Context

A retrieval-augmented generation (RAG) system can have a strong embedding model and still retrieve poor evidence. One common reason is document chunking: the text was divided into units that are awkward to search or incomplete when read on their own. If chunks are too large, one embedding must represent several unrelated ideas and retrieval becomes less precise. If chunks are too small, the retrieved text may omit the definitions, qualifiers, or surrounding steps needed to answer correctly. The problem is therefore not to find one universally correct chunk size. It is to create retrieval units that are focused enough to match a query and complete enough to be useful after retrieval.

Artificial Intelligence 06 Sep 2026 11 min read

Contrastive Learning for Text Embeddings

A text embedding model turns text into a vector so that software can compare meaning numerically. The difficult part is not producing vectors. A neural network can produce vectors for almost any input. The difficult part is teaching the geometry of those vectors so that distances correspond to the relationships your application cares about. Contrastive learning provides a practical way to do that. Instead of asking a model to predict a class label, you show it examples that should be close together and examples that should be farther apart. Training adjusts the encoder so that those relationships become easier to recover from the resulting vectors.

Artificial Intelligence 06 Sep 2026 11 min read

Compress Embeddings with Scalar Quantization

Embedding systems can become expensive for a reason that has little to do with the embedding model itself: storing and scanning the vectors. A collection of millions of dense vectors can consume gigabytes even before an index adds its own data structures. Moving those vectors through memory can also become part of query latency. Scalar quantization reduces that cost by representing each embedding coordinate with fewer bits. Instead of storing every coordinate as a 32-bit floating-point value, a system might map it to an 8-bit integer and keep enough information to approximately reconstruct or compare the original value.

Artificial Intelligence 06 Sep 2026 9 min read

Choose Pooling Strategies for Text Embeddings

A transformer usually produces one contextual representation for every input token. Many applications, however, need one vector for an entire sentence, query, or document. Semantic search, clustering, and similarity systems commonly compare these fixed-size vectors rather than every token representation separately. The operation that turns a variable number of token vectors into one vector is called pooling. It can look like a minor implementation detail, but changing it changes the representation being compared. Averaging every meaningful token, selecting a designated token, or emphasizing particular positions encodes different assumptions about where useful information lives.

Artificial Intelligence 05 Sep 2026 9 min read

Tie Input and Output Embeddings in Language Models

A language model needs token representations in two places. At the input, it converts token IDs into vectors. At the output, it converts a hidden vector into one score for every token in the vocabulary. A straightforward design gives these two operations separate parameter matrices, even though both matrices associate vocabulary items with vectors. Weight tying removes that duplication by using the same matrix for both roles. The model looks up input embeddings from the matrix and later uses its transpose to produce output logits. This can remove a large block of parameters, but it also couples two parts of the model that would otherwise learn independently.

Artificial Intelligence 05 Sep 2026 8 min read

Pooling Token Embeddings into Sequence Representations

Transformer encoders usually produce one vector for every input token. Many downstream tasks, however, need one vector for the whole input: a classifier may need a single representation of a support ticket, and a retrieval system may need one vector for an entire passage. The step that converts a variable number of token vectors into one fixed-size vector is pooling. It looks simple, but the choice of pooling rule changes what information survives, how padding must be handled, and whether the resulting vector matches the way a model was trained.

Artificial Intelligence 05 Sep 2026 8 min read

Pool Token Embeddings into Text Representations

A Transformer usually produces one contextual vector for every input token. Many downstream tasks, however, need one vector for the whole text. Semantic search may need one vector per document, clustering needs one vector per item, and similarity scoring often expects two fixed-size vectors to compare. Pooling is the step that turns a variable number of token vectors into one fixed-size representation. The operation looks simple, but small implementation choices can change the resulting geometry. Averaging padding tokens, assuming the first token is meaningful for every model, or changing pooling at deployment time can make an otherwise correct embedding pipeline behave poorly.

Artificial Intelligence 05 Sep 2026 9 min read

Normalize Embeddings Before Dot-Product Similarity

Embedding systems often compare vectors with cosine similarity or a dot product. The formulas look similar enough that it is easy to treat the two metrics as interchangeable. They are not interchangeable for arbitrary vectors. A dot product depends on both the angle between two vectors and their magnitudes. Cosine similarity removes magnitude and compares direction only. That difference can change nearest-neighbor rankings, retrieval results, and similarity thresholds. This article builds a practical mental model for deciding whether to normalize embeddings. You will see why L2 normalization makes dot product equivalent to cosine similarity, how inconsistent normalization breaks comparisons, and when preserving vector magnitude may be intentional.

Artificial Intelligence 05 Sep 2026 10 min read

Combine Retrieval Rankings with Reciprocal Rank Fusion

A retrieval-augmented generation (RAG) system often has more than one useful way to find evidence. Keyword retrieval is good at exact names, identifiers, and rare terms. Embedding retrieval can find passages that express the same idea with different wording. Using both can improve candidate coverage, but it creates a practical problem: their scores usually do not mean the same thing. A keyword score of 12.4 and a cosine similarity of 0.81 cannot be safely averaged just because both are numbers. Their scales, distributions, and even direction conventions depend on the retrieval methods and implementations.

Artificial Intelligence 04 Sep 2026 10 min read

Negative Sampling for Representation Learning

Some representation-learning problems have an awkward shape: each training example has one observed target, but the model could choose from thousands or millions of alternatives. Computing a score and normalization term for every alternative on every update can become a major training cost. Negative sampling changes the training problem. Instead of comparing the observed target with every possible alternative, the model learns from the observed positive pair and a small set of deliberately sampled negative pairs. The update becomes much cheaper, but it also optimizes a sampled discrimination objective rather than the exact full-class objective.

Artificial Intelligence 04 Sep 2026 9 min read

Contrastive Learning with Positive and Negative Pairs

An embedding model turns an input into a vector so that useful relationships can be measured numerically. The difficult part is not producing vectors. It is teaching the geometry of the vector space: which inputs should be close, which should be far apart, and what “similar” should mean for the application. Contrastive learning provides a practical answer. Instead of training only from a class label such as billing or technical support, it trains from relationships between examples. A positive pair contains examples that should have similar representations. A negative pair contains examples that should not.

Artificial Intelligence 02 Sep 2026 5 min read

Semantic Caching for LLM Applications Without Serving Stale Answers

Large language model requests are expensive compared with ordinary cache lookups. When users repeatedly ask questions with slightly different wording, an exact string cache misses even though the intended answer may be identical. A semantic cache uses vector similarity to decide whether a new request is close enough to a previous request that its answer can be reused. The idea is attractive, but the difficult part is not storing embeddings. It is deciding when reuse is actually safe.

Artificial Intelligence 02 Sep 2026 7 min read

Embeddings and Similarity Search

Many AI applications need to find items by meaning rather than by exact words. A user may search for “reset my password” while the relevant document says “recover account access.” Traditional keyword matching can miss that relationship because the phrases share few terms. Embeddings provide another representation. An embedding model converts an input such as text into a numeric vector. Inputs with related meaning are often placed near one another in that vector space, making it possible to retrieve semantically similar items with mathematical distance or similarity measures.

Artificial Intelligence 01 Sep 2026 4 min read

Version Embeddings for Safe Semantic Search Migrations

Semantic search systems often look simple from the outside: encode a document, store its vector, encode a query, and compare the vectors. The operational difficulty appears later, when the embedding model changes. Two models can produce vectors with the same dimension and still define completely different coordinate spaces. Mixing vectors from model A with query vectors from model B can silently destroy ranking quality without producing an obvious error. The safe approach is to treat an embedding model as a versioned data dependency, not a drop-in function.