Skip to content

Archive

Model Interpretability

2 articles
Artificial Intelligence 08 Sep 2026 10 min read

Transformer Attention Weights: What They Show and What They Do Not

Transformer attention maps are visually compelling. A token appears to assign most of its attention to another token, so it is tempting to conclude that the second token caused the model’s prediction. That conclusion is stronger than the data supports. An attention weight has a precise local meaning: inside one attention operation, it controls how strongly a query mixes information from available value vectors. A complete transformer prediction, however, also depends on value vectors, residual connections, feed-forward layers, normalization, later layers, and often many attention heads. A large weight is therefore evidence about one routing operation, not a complete causal explanation.

Artificial Intelligence 06 Sep 2026 10 min read

Steer Language Models by Editing Hidden Activations

Prompting changes what a language model reads. Fine-tuning changes its parameters. There is another, more experimental way to influence generation: change the model’s internal activations while it runs. This technique is commonly called activation steering or activation engineering. A simple version measures how hidden representations differ between examples that express opposite properties, turns that difference into a steering vector, and adds a scaled version of the vector during inference. The model weights stay unchanged.