Skip to content

Archive

LLM

76 articles
Artificial Intelligence 24 Sep 2026 5 min read

Iteration-Level Scheduling Rebuilds LLM Batches Between Decode Steps

Autoregressive serving does not give every request the same completion point. One sequence may emit an end token after a few decode steps while another remains active for hundreds. If the serving engine keeps the original request batch fixed until every member finishes, completed sequences leave execution capacity stranded behind longer sequences. Iteration-level scheduling moves the scheduling boundary inward. Instead of treating an entire request as the indivisible scheduling unit, the engine returns to the scheduler after a model iteration. Finished sequences can leave, waiting work can enter, and the next iteration can run with a different set of active sequences.

Artificial Intelligence 16 Sep 2026 6 min read

Verify Speculative Decoding Without Changing Model Output

Autoregressive generation normally asks the target model to produce one next-token distribution at a time. Speculative decoding changes that execution pattern. A cheaper draft model proposes several tokens, then the target model evaluates those candidates in a batch and decides how much of the proposal can be accepted. The useful property is not merely that two models participate. The verification rule determines whether the optimization preserves the target model’s intended decoding distribution or silently changes it.

Artificial Intelligence 16 Sep 2026 6 min read

Test Transformer Mechanisms with Activation Patching

A transformer can produce two different outputs from prompts that differ in one relevant detail, yet inspection of attention weights or hidden-state similarity does not establish which internal states actually matter for that difference. Activation patching addresses a narrower question by intervening on a forward pass: replace a selected activation with the corresponding activation from another run, then measure how the output changes. The result is causal with respect to that intervention. It does not automatically identify a complete circuit, a unique mechanism, or a human-readable feature. That boundary is central to using patching results correctly.

Artificial Intelligence 16 Sep 2026 6 min read

Preserve Attention Sinks in Streaming KV Caches

A bounded KV cache seems to invite a simple eviction rule: keep the newest tokens and discard the oldest ones. For some transformer language models, that rule can degrade generation even when the discarded prefix carries little obvious semantic value. A small set of early positions may attract substantial attention across later decoding steps. These positions are commonly called attention sinks. This behavior matters for streaming inference because cache eviction changes the attention computation itself. A fixed-size cache that preserves a few sink positions plus a recent window can behave differently from a cache containing only the same number of recent positions.

Artificial Intelligence 16 Sep 2026 5 min read

Control Beam Search Length Bias with Sequence Scoring

Beam search can prefer a short completed sequence even when a longer continuation looks locally plausible at every token. The effect follows from the score being optimized. If a decoder ranks complete hypotheses by the sum of token log probabilities, every additional token contributes a value that is at most zero. Extending a sequence therefore cannot increase its raw accumulated log probability. This property is not a defect in probability theory. A sequence probability is a product of conditional probabilities, and its logarithm is their sum. The implementation concern appears when raw sequence probability is also used as the ranking objective for outputs whose lengths vary.

Artificial Intelligence 16 Sep 2026 6 min read

Account for Exposure Bias in Autoregressive Decoding

An autoregressive model can receive cleaner context during training than it receives during generation. Under teacher forcing, the next-token prediction is conditioned on a reference prefix from the training sequence. During free-running decoding, the model instead conditions on tokens it generated itself. Once a generated token differs from the intended continuation, later predictions operate on a prefix that training may have represented less often. This mismatch is commonly called exposure bias. It is not simply a claim that autoregressive models make errors. The specific issue is that the distribution of prefixes presented to the model can change between optimization and generation, and an early deviation can change every subsequent conditional prediction.

Artificial Intelligence 15 Sep 2026 6 min read

Control Length Bias in Beam Search Scoring

Beam search compares multiple partial outputs while autoregressive generation advances token by token. A common scoring rule adds token log probabilities along each candidate sequence. That rule is mathematically consistent with sequence probability, but it also creates a structural preference that developers can miss: extending a sequence normally makes its accumulated log score smaller. This matters whenever candidates of different lengths compete. A decoder can rank a short completed sequence above a longer candidate even when the longer candidate is more useful for the application. Length normalization and length penalties modify that ranking, but they also change the objective being optimized.

Artificial Intelligence 14 Sep 2026 6 min read

Speculative Decoding Trades Draft Accuracy for Target Model Work

Autoregressive generation normally commits tokens one position at a time. Even when a large model has ample parallel compute available, each next-token decision depends on the prefix produced so far. That serial dependency makes decoding latency sensitive to the number of target-model passes. Speculative decoding changes the unit of work. A cheaper draft model proposes several candidate tokens, then the target model evaluates the proposed continuation in one pass. Accepted candidates advance generation by multiple positions without requiring one separate target pass per accepted token.

Artificial Intelligence 14 Sep 2026 7 min read

Reduce Repetition with Unlikelihood Training

An autoregressive language model is usually trained to increase the probability of the observed next token. That positive objective does not directly state which plausible but unwanted alternatives should receive less probability. When repetitive tokens or phrases remain locally probable, ordinary next-token training can leave generation with a strong route back into content that has already appeared. Unlikelihood training adds a negative signal for selected candidates. Instead of only rewarding the target token, the objective can also penalize tokens chosen because they represent an unwanted behavior, such as repetition within the generated prefix.

Artificial Intelligence 14 Sep 2026 6 min read

Reduce KV Cache Memory with Grouped-Query Attention

Autoregressive transformer inference stores key and value vectors from earlier tokens so each new token does not have to recompute them. With standard multi-head attention, every attention head has its own key and value projections, so the KV cache grows with the number of key-value heads. Grouped-query attention changes that head layout. It keeps multiple query heads but lets several query heads share one key head and one value head. The result reduces cached key-value state without collapsing all query heads into a single shared projection.

Artificial Intelligence 14 Sep 2026 5 min read

Prevent Cross-Example Attention in Packed Sequences

Short training examples can waste much of a fixed-length transformer batch on padding. Sequence packing reduces that waste by placing several examples into one token buffer, but concatenation alone changes the computation. A causal mask prevents a token from attending to future positions; it does not prevent that token from attending to an earlier, unrelated example. The distinction matters whenever packed examples are intended to remain independent. The token buffer may be contiguous for storage and compute while attention, position handling, and loss accounting still need explicit example boundaries.

Artificial Intelligence 14 Sep 2026 6 min read

Bound Extreme Logits with Soft Capping

A transformer can produce logits whose magnitudes grow far beyond the range needed to express a strong preference. Large attention scores can make a softmax distribution extremely concentrated, while large output logits can make token probabilities nearly one-hot. A hard clamp can bound those values, but it introduces a flat region with an abrupt derivative change at the threshold. Logit soft capping uses a smooth saturating function instead. One common form is:

Artificial Intelligence 14 Sep 2026 5 min read

Account for Exposure Bias in Autoregressive Generation

An autoregressive model can receive a clean prefix at every training position and still face a different input distribution during generation. Training commonly scores the next reference token while conditioning on earlier reference tokens. At inference time, the prefix contains the model’s own outputs instead. This mismatch is called exposure bias. It matters because an early generation error does more than make one token incorrect. That token becomes part of the context for later predictions, placing the model in a prefix state that may have been rare or absent during training.

Artificial Intelligence 13 Sep 2026 7 min read

Steer LLM Behavior with Activation Vectors

A transformer can produce different continuations without changing its prompt, weights, or decoding settings if an internal activation is modified during the forward pass. Activation steering uses this property as an inference-time control mechanism. A vector representing a target attribute is added to, or subtracted from, a hidden representation at selected model locations. The operation is simple, but its effect depends on where the vector came from, where it is injected, and how strongly it is scaled. A direction that separates two sets of prompts in one layer is not automatically a portable semantic control across layers, model revisions, or prompt distributions.

Artificial Intelligence 13 Sep 2026 6 min read

Reuse Shared Prefix State in LLM Inference

Autoregressive LLM serving often repeats the same initial tokens across many requests. A fixed system prompt, tool schema, or document prefix can occupy thousands of tokens before request-specific text begins. Computing attention state for that identical prefix on every request repeats prefill work that has already produced the same cached keys and values under compatible execution conditions. Prefix caching stores reusable attention state for such shared token prefixes. It changes the amount of prefill computation required for a cache hit, but it does not make arbitrary similar prompts interchangeable. The reusable unit is tied to exact model input state, not semantic resemblance.

Artificial Intelligence 13 Sep 2026 7 min read

Quantize KV Caches with Separate Key and Value Error Budgets

Autoregressive transformer inference keeps past key and value tensors so each new token can attend to prior positions without recomputing the full prefix. That KV cache grows with sequence length, layer count, batch size, and the number of stored key-value heads. At long contexts, its memory footprint can become a direct limit on concurrent requests or usable context length. Quantizing the KV cache reduces bytes per stored element. The resulting approximation is not equivalent to quantizing a passive data structure, however. Cached keys participate in attention score computation, while cached values are mixed according to the resulting attention weights. Error in those two tensors therefore enters the attention operation at different points.

Artificial Intelligence 13 Sep 2026 8 min read

Preserve Attention Sinks in Sliding-Window LLM Inference

A sliding-window KV cache seems mechanically simple: keep the most recent tokens, evict older key-value entries, and continue decoding within a fixed memory budget. The complication is that some transformer models place substantial attention mass on a small set of early positions even when those positions carry little direct semantic relevance to the current token. Those positions are often called attention sinks. If a cache policy removes them while preserving only the newest tokens, the attention distribution seen during decoding can change abruptly. A bounded cache can therefore behave differently from full-context inference even when the evicted text appears unrelated to the current request.

Artificial Intelligence 13 Sep 2026 6 min read

Preserve Attention Sinks in Bounded KV Caches

A decoder that keeps only the newest key-value states can degrade even when its cache still contains enough recent text for the immediate task. In some transformer models, early token positions attract substantial attention across later decoding steps. Evicting those states changes the attention distribution, not just the amount of accessible history. Attention sink retention addresses that specific failure mode. A bounded cache preserves a small prefix of initial key-value states together with a moving window of recent states. Tokens between those regions can be discarded, keeping cache size bounded as generation continues.

Artificial Intelligence 13 Sep 2026 6 min read

Control Token Repetition with Logit Penalties

A language model can assign high probability to a token that has already appeared several times in the generated text. If the decoder keeps selecting that token or a short pattern containing it, the output may settle into repetition even though each individual choice is plausible under the model. A repetition penalty changes this behavior at decoding time. It modifies candidate scores according to token history before the next token is selected. The model parameters stay fixed, but the effective distribution used by the decoder no longer matches the model’s unmodified next-token distribution.

Artificial Intelligence 13 Sep 2026 7 min read

Control Sequence Length Bias in Beam Search

Beam search can return a shorter sequence even when a longer candidate contains locally plausible tokens at every position. The behavior follows directly from sequence scoring: autoregressive models multiply conditional token probabilities, or equivalently add their log probabilities. Since token probabilities are at most one, each additional token contributes a non-positive log term. That arithmetic makes sequence length part of decoding. Beam width changes which candidates survive, but it does not remove the scoring effect. A decoder therefore needs a deliberate policy for comparing hypotheses of different lengths and for deciding when a completed hypothesis is good enough to stop the search.

Artificial Intelligence 13 Sep 2026 6 min read

Contrastive Decoding with Expert and Amateur Models

A language model can give high next-token probability to text that is fluent but generic. Contrastive decoding changes the ranking by asking for a second signal: does a weaker model also find the same candidate easy to predict? A candidate favored by the expert but not by the amateur receives stronger relative support than one both models score highly. This is an inference-time mechanism. It does not alter either model’s parameters, and it does not convert the amateur model into a verifier. The decoder combines two token distributions and then selects from the resulting scores.

Artificial Intelligence 12 Sep 2026 8 min read

Test Transformer Circuits with Activation Patching

A transformer can expose a clear internal pattern without that pattern being responsible for the output under inspection. Activation patching addresses this gap by changing an internal state and measuring the downstream effect. Instead of asking whether a feature is visible at a layer, it asks whether replacing a selected state changes a defined model behavior. The method is simple in form but sensitive to experimental design. A patch has meaning only relative to the paired inputs, the patched location, the replacement value, and the output metric. Changing any of those can change the causal question being tested.

Artificial Intelligence 12 Sep 2026 9 min read

Stream Long LLM Sessions with Attention Sinks

Stream Long LLM Sessions with Attention Sinks Long-running LLM sessions create a simple resource problem: every generated token can add keys and values to the attention cache. Keep the entire history and memory use keeps growing. Keep only the newest tokens and some transformer models degrade sharply once older cache entries disappear. Attention sinks provide a useful middle ground for compatible models. Instead of retaining the full KV cache, keep a small group of initial tokens plus a moving window of recent tokens. The cache stays bounded, yet the model can remain much more stable than with a recent-token window alone.

Artificial Intelligence 12 Sep 2026 7 min read

Speculative Decoding Depends on Draft Acceptance

Autoregressive generation normally commits one token after each model pass, creating a serial dependency across the output sequence. Speculative decoding changes that execution pattern. A cheaper draft process proposes several future tokens, then the target model evaluates those proposals together and determines which tokens can be committed. The attraction is fewer serial target-model iterations per generated token. That does not make speculative decoding an automatic latency reduction. Its useful operating point depends on how cheaply candidates are produced, how many survive verification, and how much extra work the target model performs while checking them.