Skip to content

Archive

LLM

76 articles
Artificial Intelligence 12 Sep 2026 7 min read

Inspect Transformer Layer Predictions with the Logit Lens

A decoder-only transformer normally exposes token logits only after its final block and output normalization. The residual stream inside earlier blocks has the same model-width shape, which makes another operation possible: take an intermediate state, apply the model’s output-side normalization when required, and project that state through the output matrix. The resulting vocabulary scores form the logit lens. They provide a token-space view of an internal representation before the remaining transformer blocks have processed it. That view is useful for inspecting how candidate tokens change across depth, but it is not a record of tokens that the model has secretly selected in advance.

Artificial Intelligence 12 Sep 2026 10 min read

Improve LLM Generation with Contrastive Decoding

Improve LLM Generation with Contrastive Decoding A language model can assign high probability to text that is fluent but bland, repetitive, or overly driven by common patterns. Sampling adds variety, but increasing randomness can also admit weak continuations. Contrastive decoding takes a different route: compare a stronger model with a weaker reference model at each generation step, then favor tokens that the stronger model supports more distinctly. The method changes decoding rather than model parameters. It can therefore be useful when you control inference for compatible models and want to experiment with generation quality without another training run. The extra model pass is not free, and the method needs a plausibility guard to avoid promoting strange tokens.

Artificial Intelligence 12 Sep 2026 7 min read

Extend RoPE Context with Position Interpolation

A transformer that uses rotary position embeddings can accept a larger token buffer at the serving layer and still behave poorly at positions far beyond the range used during model training. The tensor shapes may be valid while the positional phases presented to attention are outside the regime the model adapted to. Position interpolation addresses that mismatch by compressing a longer sequence’s position indices into the original position interval before applying RoPE. It does not add memory to the architecture, and it does not make long-context behavior equivalent to native training at the extended length. It changes the positional coordinates supplied to attention.

Artificial Intelligence 12 Sep 2026 9 min read

Extend RoPE Context Windows with Position Interpolation

Extend RoPE Context Windows with Position Interpolation A RoPE-based language model trained on sequences up to a fixed length can behave poorly when inference suddenly asks it to process much larger position indices. The tokens are valid, but the positional pattern can move far outside the range used during training. Position interpolation changes that geometry. Instead of sending larger position indices directly into rotary position embeddings, it compresses a longer sequence into the positional range the model already uses. With suitable adaptation, this can extend the usable context window without changing the transformer architecture.

Artificial Intelligence 12 Sep 2026 9 min read

Control LLM Behavior with Activation Steering

Control LLM Behavior with Activation Steering Prompting controls a language model through its input tokens. Fine-tuning changes model parameters. Activation steering offers a third option: change selected internal activations while the model runs, without rewriting its weights. That makes activation steering useful for experiments where you want to test whether an internal direction is connected to a behavior, or apply a lightweight behavior shift during generation. It also creates new engineering questions. A steering vector can help at one layer and damage output at another. A strength that works on short prompts can become excessive on different inputs. A behavioral shift can also come with losses in fluency or task accuracy.

Artificial Intelligence 12 Sep 2026 7 min read

Bound KV Cache Growth in Streaming LLM Inference

Autoregressive transformer inference normally keeps key and value states from earlier tokens so each new token can attend to prior context without recomputing those states. The cache grows with sequence length. For a long-running stream, that growth eventually becomes a memory constraint even when generation itself continues one token at a time. KV cache eviction puts a bound on that state by discarding selected cached positions. The memory effect is straightforward: fewer retained key-value pairs occupy less cache space. The model effect is more subtle. Once a position is removed, later attention layers cannot use its cached key and value in the ordinary attention calculation. An eviction policy therefore changes both resource use and the effective attention history.

Artificial Intelligence 12 Sep 2026 10 min read

Align LLMs with Direct Preference Optimization

Align LLMs with Direct Preference Optimization Supervised fine-tuning works well when you can provide a target response for each prompt. It becomes less natural when the signal is comparative: one answer is preferred over another, but neither is a perfect target to copy. Direct Preference Optimization (DPO) turns those preference pairs into a training objective for a language model without requiring a separately trained reward model or an online reinforcement step.

Artificial Intelligence 11 Sep 2026 9 min read

Adapt Language Models with Prefix Tuning

Adapt Language Models with Prefix Tuning Full fine-tuning changes a model’s weights for each task. That can be effective, but storing and serving a separate full checkpoint for every task becomes expensive as model size and task count grow. Prefix tuning offers a different arrangement: keep the pretrained model frozen and train a small set of task-specific states that participate in attention. This article builds a practical mental model for prefix tuning, shows how it differs from text prompts and low-rank weight adapters, and explains the trade-offs that matter when training or serving several task variants.

Artificial Intelligence 10 Sep 2026 10 min read

Temperature in LLM Sampling

A language model can produce very different continuations from the same prompt even when its weights and context haven’t changed. One of the controls behind that variation is temperature. Temperature is often described as a creativity knob. That description is convenient but incomplete. Temperature doesn’t add ideas to a model, improve its knowledge, or directly control factual accuracy. It changes the probability distribution used to choose the next token. The practical effect depends on what the model already considers plausible at that step.

Artificial Intelligence 10 Sep 2026 9 min read

Reduce Repetitive Generation with Unlikelihood Training

Reduce Repetitive Generation with Unlikelihood Training A language model can learn to predict ordinary text well and still assign too much probability to behavior you don’t want at generation time. Repetition is a common example: once a phrase appears, the model may keep making recently used tokens plausible enough that a decoding loop becomes hard to escape. Changing the decoder can hide some of this behavior, but it doesn’t change the probabilities learned by the model. Unlikelihood training takes a different approach. During training, it identifies undesirable candidates and explicitly pushes their probabilities down while the usual likelihood objective pushes desired tokens up.

Artificial Intelligence 10 Sep 2026 10 min read

Pack Training Sequences Without Leaking Between Examples

Pack Training Sequences Without Leaking Between Examples Language-model training often wastes computation on padding. If a batch contains examples with very different lengths, shorter examples are extended with padding so tensors have compatible shapes. The model still has to move those tensor positions through parts of the training pipeline even though they contain no training content. Sequence packing reduces that waste by placing multiple shorter examples into one fixed-length training sequence. The idea is simple; the boundary handling is not. If attention or loss masks are wrong, one example can accidentally use another example as context, or the model can be trained to predict tokens that should not count as targets.

Artificial Intelligence 10 Sep 2026 8 min read

Mask Prompt Tokens During Instruction Fine-Tuning

A supervised language-model example often contains more than the text you want the model to produce. It may include a system message, a user request, separators, and an assistant answer. If you compute next-token loss over the entire sequence, the model is trained to predict all of those tokens, not just the assistant response. That may be intentional for some training objectives. For instruction fine-tuning, though, developers often want the prompt to provide context while only selected response tokens contribute to the supervised loss. A loss mask makes that distinction explicit.

Artificial Intelligence 10 Sep 2026 10 min read

Evaluate Knowledge Edits Before Trusting Model Updates

Changing one fact in a language model sounds simpler than retraining it. If a product name changes, an organization moves offices, or a fictional knowledge base is updated, knowledge editing aims to change a model’s behavior for that information without running broad training again. The hard part isn’t making one prompt produce the new answer. The hard part is knowing what else changed. A useful evaluation therefore asks more than “did the edit work?” It checks whether the new fact survives reasonable paraphrases, whether unrelated behavior stays stable, and whether the model can use the edited information when another answer depends on it. This article builds that evaluation model and shows how to turn it into a practical test suite.

Artificial Intelligence 09 Sep 2026 9 min read

LLM Sampling: Temperature, Top-K, and Top-P

A language model does not normally produce a single inevitable next token. Given a prefix, it assigns scores to many possible tokens. A decoding algorithm then decides how to turn those scores into the next output. That last step matters. If you sample too freely, a model can drift into unlikely continuations. If you restrict sampling too aggressively, outputs can become repetitive or lose useful variation. Parameters such as temperature, top-k, and top-p control different parts of this trade-off, so treating them as interchangeable “creativity settings” leads to confusing results.

Artificial Intelligence 09 Sep 2026 8 min read

Contextual Calibration for Few-Shot Classifiers

A language model can behave like a classifier without any parameter updates: give it a few labeled examples, present a new input, and ask it to choose a label. The surprising problem is that the answer can depend not only on the new input, but also on details such as the prompt wording, demonstration order, and label tokens. That creates a practical debugging trap. A prompt may appear to teach the task while also giving the model a baseline preference for one answer before meaningful input is considered.

Artificial Intelligence 09 Sep 2026 10 min read

Choose Consensus Outputs with Minimum Bayes Risk Decoding

A generative model can assign high probability to an output that is not the most useful answer for your application. This is especially visible when several different outputs are plausible: a translation can have multiple valid phrasings, a summary can emphasize different details, and a structured generator can produce several semantically similar candidates. Greedy decoding chooses locally likely tokens. Beam search searches for a high-probability sequence. Sampling gives you diverse candidates. None of those methods, by itself, asks a different question that is often closer to the application goal: which candidate agrees best with the distribution of plausible outputs?

Artificial Intelligence 08 Sep 2026 11 min read

Retrieve Multi-Hop Evidence with Graph RAG

Retrieval-augmented generation (RAG) usually starts with a simple idea: find text chunks similar to a question, place the most relevant chunks in the model’s context, and ask the model to answer from that evidence. This works well when the answer is stated in one passage or in several passages that are independently easy to retrieve. Some questions are harder because the useful evidence is connected by relationships, not just by similar wording. A developer may need to answer, “Which service depends on the library maintained by the team that owns the payment API?” No single chunk has to contain all of those words. The answer may require following several links across services, libraries, teams, and APIs.

Artificial Intelligence 08 Sep 2026 8 min read

Estimate LLM Uncertainty with Semantic Entropy

A language model can produce a fluent answer even when it is uncertain. Token probabilities help describe uncertainty during generation, but they can be misleading at the answer level because many different strings can express the same meaning. Consider a question whose correct answer is Paris. A model might generate Paris, The answer is Paris, and France's capital is Paris. These strings differ, yet they represent the same answer. Treating them as three unrelated outcomes exaggerates the apparent uncertainty.

Artificial Intelligence 08 Sep 2026 10 min read

Avoid Tokenization Boundary Failures in LLM Generation

A language model application usually treats a prompt as text: provide a prefix, then ask the model to continue it. The model sees something more specific. Its tokenizer first converts that text into tokens, and the end of the prompt forces the last token to end at exactly that position. That detail can matter when the prompt ends at a character position that would normally fall inside a larger token if the prompt and its continuation were tokenized together. The resulting tokenization boundary problem, also called the partial token problem, can make an otherwise natural continuation unexpectedly unlikely.

Artificial Intelligence 07 Sep 2026 10 min read

Repetition Controls for Language Model Generation

A language model can produce fluent text and still get stuck repeating a phrase, sentence pattern, or idea. A support reply may restate the same apology three times. A summarizer may loop over one point. A long generation can begin copying a short phrase again and again. The tempting fix is to turn up a “repetition” setting until the duplicate text disappears. That can solve one symptom while creating another: names become awkward, code becomes invalid, required terms disappear, or a model avoids repeating words that the task genuinely needs.

Artificial Intelligence 07 Sep 2026 12 min read

Measure Tokenization Efficiency Across Languages

Two prompts can communicate roughly the same amount of information and still consume very different numbers of model tokens. The difference can appear between languages, writing systems, domains, or even formatting styles. That matters because language-model systems usually operate on tokens rather than characters or words. A context window is measured in tokens. Many hosted APIs account for usage in tokens. Longer token sequences can also increase inference work, although the exact latency and compute effect depends on the model, serving stack, batching, caching, and whether the tokens belong to the input or generated output.

Artificial Intelligence 07 Sep 2026 9 min read

Improve Reasoning Reliability with Self-Consistency

A language model can reach different answers to the same reasoning problem depending on how generation unfolds. One sampled path may make an arithmetic mistake, another may misread a condition, and a third may reach the correct result. If an application trusts only one path, its answer depends heavily on that single generation. Self-consistency uses this variability instead of trying to eliminate it. It samples several reasoning paths for the same problem, extracts their final answers, and chooses the answer supported by the largest share of the samples. The technique is an inference-time strategy: it does not require changing model weights.

Artificial Intelligence 06 Sep 2026 10 min read

Use Best-of-N Sampling to Spend More Compute at Inference

A language model can produce several plausible answers to the same prompt. That variability is often treated as noise, but it can also be used deliberately: generate multiple candidates, evaluate them, and return the strongest one. This pattern is called best-of-N sampling. Instead of trusting one generation, the system samples N responses and uses a scoring rule to select one. The extra samples spend more compute at inference time in exchange for more opportunities to find a good response.

Artificial Intelligence 06 Sep 2026 9 min read

Token and Sequence Biases for LLM Decoding

Sometimes an LLM produces generally good text but makes one narrow decoding choice too often. Perhaps a domain-specific abbreviation should be preferred, a deprecated product name should be discouraged, or a particular token must not appear in generated text. Changing temperature is a poor fit for this problem because temperature affects the whole next-token distribution. Retraining a model is usually excessive when the desired change is local. Token and sequence biases provide a narrower tool: modify selected prediction scores during decoding while leaving the model parameters unchanged.