Two embedding candidates can differ in semantic fit yet receive cosine scores packed into a narrow interval. The similarity function may be implemented correctly. The compression can instead come from the geometry of the embedding space: vectors may occupy preferred directions rather than spreading evenly across the available dimensions. This directional concentration is commonly described as anisotropy.

Anisotropy matters to retrieval because cosine similarity measures angular alignment. When many vectors share a substantial common component, unrelated pairs can start from an elevated baseline alignment. Relevant pairs may still score higher, but the usable gap between relevant and irrelevant candidates can shrink.

A shared direction shifts the score distribution

Consider embeddings represented schematically as a shared component plus item-specific variation:

x = c + r_x
y = c + r_y

Here c is a direction present across many vectors, while r_x and r_y contain variation specific to each item. This decomposition is descriptive rather than a guarantee about a particular embedding model.

If the shared component contributes strongly to vector direction, the dot product contains a term from c · c in addition to the item-specific terms. After L2 normalization, cosine similarity still reflects the resulting direction. Normalization fixes vector length; it does not remove a common directional component.

That distinction is operationally useful. A system can normalize every vector correctly and still observe cosine similarities concentrated in a relatively narrow positive range. Treating normalization as a complete correction for representation geometry can therefore hide the actual source of score compression.

Narrow scores change threshold behavior

A cosine threshold has meaning only relative to the score distribution produced by a particular embedding model, corpus, preprocessing pipeline, and query population. A threshold such as 0.8 is not a universal semantic boundary.

Suppose relevant results tend to occupy one score band and irrelevant results another. If anisotropy raises the common baseline and compresses both bands, their numerical values can move closer together even when their ranking remains partly useful. A threshold copied from another model or corpus can then accept too many candidates or reject valid ones.

The diagnostic target is not merely the average cosine score. Compare score distributions for known relevant and irrelevant pairs, and inspect their overlap. Also compare random or mismatched pairs against query-candidate pairs. If unrelated pairs already have substantial positive similarity, absolute score magnitude carries less information than it would in a space with a lower baseline alignment.

Ranking metrics and threshold metrics expose different consequences. A retrieval system can preserve useful ordering while offering poor separation for a fixed acceptance threshold. Conversely, a transformation that spreads scores numerically is not automatically beneficial if it damages relevance ordering.

Mean centering changes the coordinate system

One possible geometric transformation subtracts an estimated mean vector:

x_centered = x - μ

where μ is estimated from a reference set. This removes the first moment of that reference distribution. It does not guarantee an isotropic space, because covariance can still be strongly directional.

The reference set also matters. A mean estimated from one corpus or traffic distribution may not represent another. If production data drifts, the transformation can become misaligned with the vectors it processes. Query and document embeddings must also remain compatible under any transformation used for similarity computation.

More aggressive transformations can rescale directions according to covariance structure or suppress dominant components. Those operations alter the geometry that the embedding model produced. They can reduce directional concentration under a chosen measurement, but they can also remove dimensions that carry task-relevant signal. Geometric uniformity is not itself the retrieval objective.

Anisotropy describes directional concentration in a vector distribution. Hubness describes candidates that occur unusually often in nearest-neighbor lists. The two can coexist, and representation geometry can contribute to both, but one does not establish the other.

This separation prevents a misleading diagnosis. A narrow cosine distribution does not prove that a small set of candidates dominates retrieval. Likewise, repeated nearest neighbors do not by themselves identify a single global direction as the cause. Score-distribution analysis and neighbor-frequency analysis measure different properties and should remain separate during evaluation.

Approximate nearest-neighbor indexing is another distinct layer. An index can introduce recall error relative to exact search, but it does not create the underlying embedding distribution. Comparing a manageable exact-search sample with production retrieval helps isolate index approximation from representation geometry.

Geometry diagnostics need task labels

Anisotropy can be measured without relevance labels, but deciding whether it is harmful requires a task-level criterion. Pairwise cosine histograms, mean vectors, principal directions, and covariance summaries can reveal concentration. None of those measurements alone establishes that retrieval quality is poor.

For a retrieval application, the useful test connects geometry to decisions. Compare relevant and irrelevant score distributions, ranking quality, threshold error patterns, and candidate recurrence before and after any transformation. Keep the embedding model, corpus slice, and evaluation queries controlled where possible so that a geometry change is not confused with a data change.

The implementation boundary is straightforward: anisotropy is a property of the representation distribution, while retrieval quality is a property of the representation combined with data, similarity rules, indexing, and the application objective. A narrower or more uniform geometric distribution is useful only when it improves the decisions the retrieval system is expected to make.