Skip to content

Archive

Representation

4 articles
Artificial Intelligence 24 Sep 2026 5 min read

Embedding Anisotropy Compresses Cosine Similarity Ranges

Two embedding candidates can differ in semantic fit yet receive cosine scores packed into a narrow interval. The similarity function may be implemented correctly. The compression can instead come from the geometry of the embedding space: vectors may occupy preferred directions rather than spreading evenly across the available dimensions. This directional concentration is commonly described as anisotropy. Anisotropy matters to retrieval because cosine similarity measures angular alignment. When many vectors share a substantial common component, unrelated pairs can start from an elevated baseline alignment. Relevant pairs may still score higher, but the usable gap between relevant and irrelevant candidates can shrink.

Artificial Intelligence 15 Sep 2026 6 min read

Measure Anisotropy in Embedding Spaces

Embedding vectors can occupy a narrow cone instead of spreading evenly across their available dimensions. In that geometry, unrelated items may still have noticeably positive cosine similarity because many vectors share a common directional component. This concentration is called anisotropy. For developers, anisotropy matters at the point where vector geometry becomes an application signal. A similarity threshold, nearest-neighbor ranking, clustering rule, or novelty detector inherits the distribution produced by the embedding model. The same cosine value can carry different meaning across representation spaces with different directional concentration.

Artificial Intelligence 14 Sep 2026 6 min read

Measure Embedding Anisotropy Before Trusting Cosine Similarity

Cosine similarity is often treated as if a score has the same meaning across any embedding space. That assumption breaks when vectors occupy a narrow region of the available geometry. If many embeddings share a strong common direction, unrelated items can receive positive cosine scores simply because both align with that direction. This behavior is usually described as embedding anisotropy. It is not a defect in cosine similarity itself. The issue is that cosine measures angles in the representation it receives, including global structure that may have little value for the downstream comparison.

Artificial Intelligence 14 Sep 2026 5 min read

Control Classifier Logits with Cosine Normalization

A linear classification head mixes two signals in each logit: the angle between a feature vector and a class weight vector, and the magnitudes of both vectors. Cosine normalization removes the magnitude terms, so class scores depend on directional alignment instead. That change is small in code but substantial in interpretation. Feature norm no longer increases every class comparison merely by growing, class-weight norm no longer acts as an implicit class-specific scale, and the overall sharpness of the softmax must be supplied separately.