Skip to content

Archive

Inference

87 articles
Artificial Intelligence 22 Sep 2026 6 min read

Padding Masks Do Not Remove Padding Compute in Dense Attention

A batch can contain two prompts with very different token counts yet represent both with the same rectangular tensor. The shorter prompt is extended with padding so its tensor shape matches the longest sequence in the batch. An attention mask can stop those padded positions from contributing to attention probabilities, but that semantic exclusion does not imply that dense kernels skip every operation associated with the padded rows and columns.

Artificial Intelligence 19 Sep 2026 8 min read

Where PEGASUS-XSum Inference Time Goes

google/pegasus-xsum is a summarization checkpoint, not a compact text utility. Its latency follows directly from the work performed by a large Transformer encoder-decoder: first encode the source document, then run the decoder repeatedly until the summary is complete. On a CPU, the second phase is usually the part that makes a short output feel disproportionately expensive. The original PEGASUS work describes a Transformer encoder-decoder pretrained with Gap Sentences Generation and reports a 568M-parameter best model. The XSum checkpoint is fine-tuned for highly abstractive single-document summarization. That combination is useful when summary quality matters, but it also means inference has substantially more machinery than extracting a few source sentences or running a small classifier.

Artificial Intelligence 19 Sep 2026 6 min read

Prefix Caching Reuses KV State Only Across Identical Prompt Prefixes

Autoregressive transformer serving often repeats the same prompt prefix across requests: a system message, a long document header, or a fixed tool schema may precede user-specific text. Prefix caching stores the key-value state produced by that shared prefix so a later request can resume computation from the cached boundary instead of recomputing the entire prefix. The useful boundary is narrower than “similar prompts.” Reuse depends on the exact token sequence and on model state that affects the cached activations. A one-character text edit may preserve most tokens, shift tokenization near the edit, or change every token after a formatting boundary. The cache can only reuse the portion whose effective input is still identical.

Artificial Intelligence 19 Sep 2026 6 min read

Ollama Gemma 3 270M Memory Is More Than the Model File

Ollama lists gemma3:270m at about 292 MB. That number is useful for storage planning, but it is not a RAM requirement. It describes the packaged model data for the default Ollama variant, which uses Q8_0 quantization. Once inference starts, the runtime also needs memory for model metadata, execution buffers, token state, and the key-value cache used by attention. That distinction matters on small machines. A device with 512 MB of RAM may appear large enough when compared only with a 292 MB model file, yet the remaining memory must also accommodate Ollama and the operating system. The context configuration can move the total substantially.

Artificial Intelligence 19 Sep 2026 8 min read

Chunk Long Prefills to Limit Decode Stalls in LLM Serving

A long prompt can occupy an accelerator for a much larger scheduling interval than a single decode iteration. When a serving engine mixes new prefills with requests that are already generating tokens, that difference can show up as irregular time between output tokens. The model has not changed; the interference comes from how two distinct inference phases share execution time. Prefill processes a prompt and builds the key-value state required by later causal attention. Decode then extends the sequence autoregressively, usually one new token per active request per iteration. Those phases place different pressure on hardware, so treating them as interchangeable scheduling units can produce avoidable stalls.

Artificial Intelligence 16 Sep 2026 6 min read

Verify Speculative Decoding Without Changing Model Output

Autoregressive generation normally asks the target model to produce one next-token distribution at a time. Speculative decoding changes that execution pattern. A cheaper draft model proposes several tokens, then the target model evaluates those candidates in a batch and decides how much of the proposal can be accepted. The useful property is not merely that two models participate. The verification rule determines whether the optimization preserves the target model’s intended decoding distribution or silently changes it.

Artificial Intelligence 16 Sep 2026 6 min read

Preserve Attention Sinks in Streaming KV Caches

A bounded KV cache seems to invite a simple eviction rule: keep the newest tokens and discard the oldest ones. For some transformer language models, that rule can degrade generation even when the discarded prefix carries little obvious semantic value. A small set of early positions may attract substantial attention across later decoding steps. These positions are commonly called attention sinks. This behavior matters for streaming inference because cache eviction changes the attention computation itself. A fixed-size cache that preserves a few sink positions plus a recent window can behave differently from a cache containing only the same number of recent positions.

Artificial Intelligence 16 Sep 2026 5 min read

Preserve Attention Sinks in Sliding KV Caches

A sliding key-value cache seems to offer a simple bound on transformer inference memory: retain the most recent tokens and evict the oldest entries as generation continues. That policy preserves local context, but it can change attention behavior more sharply than token age alone suggests. Some early positions can attract substantial attention even when their lexical content is not directly relevant to the current token. Removing those positions can disturb the distribution that later layers receive.

Artificial Intelligence 16 Sep 2026 5 min read

Control Beam Search Length Bias with Sequence Scoring

Beam search can prefer a short completed sequence even when a longer continuation looks locally plausible at every token. The effect follows from the score being optimized. If a decoder ranks complete hypotheses by the sum of token log probabilities, every additional token contributes a value that is at most zero. Extending a sequence therefore cannot increase its raw accumulated log probability. This property is not a defect in probability theory. A sequence probability is a product of conditional probabilities, and its logarithm is their sum. The implementation concern appears when raw sequence probability is also used as the ranking objective for outputs whose lengths vary.

Artificial Intelligence 16 Sep 2026 6 min read

Account for Exposure Bias in Autoregressive Decoding

An autoregressive model can receive cleaner context during training than it receives during generation. Under teacher forcing, the next-token prediction is conditioned on a reference prefix from the training sequence. During free-running decoding, the model instead conditions on tokens it generated itself. Once a generated token differs from the intended continuation, later predictions operate on a prefix that training may have represented less often. This mismatch is commonly called exposure bias. It is not simply a claim that autoregressive models make errors. The specific issue is that the distribution of prefixes presented to the model can change between optimization and generation, and an early deviation can change every subsequent conditional prediction.

Artificial Intelligence 15 Sep 2026 4 min read

Preserve Attention Sinks in Streaming Transformers

A bounded attention cache creates a specific failure mode in autoregressive transformers: removing every old token can disturb attention even when those tokens no longer carry useful task content. Some early positions can absorb attention mass across many later queries. If cache eviction removes them, generation quality can degrade more than their semantic value would suggest. These positions are often called attention sinks. The practical implication is narrow but useful: a streaming cache can keep a small prefix of sink positions while rotating the rest of its capacity through recent tokens.

Artificial Intelligence 14 Sep 2026 6 min read

Speculative Decoding Trades Draft Accuracy for Target Model Work

Autoregressive generation normally commits tokens one position at a time. Even when a large model has ample parallel compute available, each next-token decision depends on the prefix produced so far. That serial dependency makes decoding latency sensitive to the number of target-model passes. Speculative decoding changes the unit of work. A cheaper draft model proposes several candidate tokens, then the target model evaluates the proposed continuation in one pass. Accepted candidates advance generation by multiple positions without requiring one separate target pass per accepted token.

Artificial Intelligence 14 Sep 2026 6 min read

Reduce KV Cache Memory with Grouped-Query Attention

Autoregressive transformer inference stores key and value vectors from earlier tokens so each new token does not have to recompute them. With standard multi-head attention, every attention head has its own key and value projections, so the KV cache grows with the number of key-value heads. Grouped-query attention changes that head layout. It keeps multiple query heads but lets several query heads share one key head and one value head. The result reduces cached key-value state without collapsing all query heads into a single shared projection.

Artificial Intelligence 14 Sep 2026 5 min read

Contrastive Decoding with an Amateur Model

A language model can assign high probability to a token for several reasons. Some reflect context-specific structure; others reflect generic tendencies that also appear in a weaker model. Contrastive decoding separates those signals by scoring candidate tokens with two models rather than one. The larger model acts as an expert. A smaller or otherwise weaker model acts as an amateur. Generation favors tokens that the expert supports more strongly relative to the amateur, subject to a plausibility constraint from the expert distribution.

Artificial Intelligence 14 Sep 2026 5 min read

Account for Exposure Bias in Autoregressive Generation

An autoregressive model can receive a clean prefix at every training position and still face a different input distribution during generation. Training commonly scores the next reference token while conditioning on earlier reference tokens. At inference time, the prefix contains the model’s own outputs instead. This mismatch is called exposure bias. It matters because an early generation error does more than make one token incorrect. That token becomes part of the context for later predictions, placing the model in a prefix state that may have been rare or absent during training.

Artificial Intelligence 13 Sep 2026 7 min read

Steer LLM Behavior with Activation Vectors

A transformer can produce different continuations without changing its prompt, weights, or decoding settings if an internal activation is modified during the forward pass. Activation steering uses this property as an inference-time control mechanism. A vector representing a target attribute is added to, or subtracted from, a hidden representation at selected model locations. The operation is simple, but its effect depends on where the vector came from, where it is injected, and how strongly it is scaled. A direction that separates two sets of prompts in one layer is not automatically a portable semantic control across layers, model revisions, or prompt distributions.

Artificial Intelligence 13 Sep 2026 6 min read

Reuse Shared Prefix State in LLM Inference

Autoregressive LLM serving often repeats the same initial tokens across many requests. A fixed system prompt, tool schema, or document prefix can occupy thousands of tokens before request-specific text begins. Computing attention state for that identical prefix on every request repeats prefill work that has already produced the same cached keys and values under compatible execution conditions. Prefix caching stores reusable attention state for such shared token prefixes. It changes the amount of prefill computation required for a cache hit, but it does not make arbitrary similar prompts interchangeable. The reusable unit is tied to exact model input state, not semantic resemblance.

Artificial Intelligence 13 Sep 2026 7 min read

Quantize KV Caches with Separate Key and Value Error Budgets

Autoregressive transformer inference keeps past key and value tensors so each new token can attend to prior positions without recomputing the full prefix. That KV cache grows with sequence length, layer count, batch size, and the number of stored key-value heads. At long contexts, its memory footprint can become a direct limit on concurrent requests or usable context length. Quantizing the KV cache reduces bytes per stored element. The resulting approximation is not equivalent to quantizing a passive data structure, however. Cached keys participate in attention score computation, while cached values are mixed according to the resulting attention weights. Error in those two tensors therefore enters the attention operation at different points.

Artificial Intelligence 13 Sep 2026 8 min read

Preserve Attention Sinks in Sliding-Window LLM Inference

A sliding-window KV cache seems mechanically simple: keep the most recent tokens, evict older key-value entries, and continue decoding within a fixed memory budget. The complication is that some transformer models place substantial attention mass on a small set of early positions even when those positions carry little direct semantic relevance to the current token. Those positions are often called attention sinks. If a cache policy removes them while preserving only the newest tokens, the attention distribution seen during decoding can change abruptly. A bounded cache can therefore behave differently from full-context inference even when the evicted text appears unrelated to the current request.

Artificial Intelligence 13 Sep 2026 6 min read

Control Token Repetition with Logit Penalties

A language model can assign high probability to a token that has already appeared several times in the generated text. If the decoder keeps selecting that token or a short pattern containing it, the output may settle into repetition even though each individual choice is plausible under the model. A repetition penalty changes this behavior at decoding time. It modifies candidate scores according to token history before the next token is selected. The model parameters stay fixed, but the effective distribution used by the decoder no longer matches the model’s unmodified next-token distribution.

Artificial Intelligence 13 Sep 2026 7 min read

Control Sequence Length Bias in Beam Search

Beam search can return a shorter sequence even when a longer candidate contains locally plausible tokens at every position. The behavior follows directly from sequence scoring: autoregressive models multiply conditional token probabilities, or equivalently add their log probabilities. Since token probabilities are at most one, each additional token contributes a non-positive log term. That arithmetic makes sequence length part of decoding. Beam width changes which candidates survive, but it does not remove the scoring effect. A decoder therefore needs a deliberate policy for comparing hypotheses of different lengths and for deciding when a completed hypothesis is good enough to stop the search.

Artificial Intelligence 13 Sep 2026 6 min read

Contrastive Decoding with Expert and Amateur Models

A language model can give high next-token probability to text that is fluent but generic. Contrastive decoding changes the ranking by asking for a second signal: does a weaker model also find the same candidate easy to predict? A candidate favored by the expert but not by the amateur receives stronger relative support than one both models score highly. This is an inference-time mechanism. It does not alter either model’s parameters, and it does not convert the amateur model into a verifier. The decoder combines two token distributions and then selects from the resulting scores.

Artificial Intelligence 13 Sep 2026 7 min read

Contrast Expert and Amateur Models During Decoding

A language model can assign high probability to a token for two different reasons: the token may fit the prompt particularly well, or it may simply be common under many contexts. Contrastive decoding tries to separate those effects by comparing the next-token distributions of two models. A stronger expert supplies the main distribution, while a weaker amateur supplies a signal for patterns that do not require the expert’s extra capability.

Artificial Intelligence 13 Sep 2026 8 min read

Balance Sparse Mixture-of-Experts Routing Under Capacity Limits

A sparse mixture-of-experts layer does not send every token through every parameter block. A router scores the available experts, selects a small subset for each token, and dispatches token representations only to those selected experts. That conditional computation is the main attraction of sparse MoE designs, but it also creates a resource-allocation problem inside the model. The router can prefer the same experts for many tokens. Hardware, meanwhile, has finite buffers and communication capacity. A routing policy that looks reasonable from token scores alone can therefore create overloaded experts, idle experts, uneven communication, or discarded assignments.