Skip to content

Archive

Inference

87 articles
Artificial Intelligence 06 Sep 2026 10 min read

Control LLM Repetition with Token Penalties

Language models sometimes repeat a phrase, return to the same point, or fall into a short loop even when the prompt asks for a concise answer. A common response is to increase randomness, but temperature changes the whole next-token distribution. That can reduce repetition while also making unrelated choices less predictable. Token penalties provide a more targeted control. They adjust the scores of tokens that have already appeared, making some repeated tokens less likely before the decoder chooses the next token. This can be useful for open-ended generation, but it is not a general quality switch: repeated tokens are often exactly what correct text requires.

Artificial Intelligence 06 Sep 2026 8 min read

Constrain LLM Output with Grammar-Guided Decoding

Asking a language model to return JSON, SQL, or another structured format creates a failure mode that ordinary prompting cannot remove: the model can understand the requested format and still generate a token that makes the output syntactically invalid. For applications that immediately parse model output, one missing quote or delimiter can turn an otherwise useful answer into an error. Retrying helps, but it spends more inference time without guaranteeing that the next attempt will parse.

Artificial Intelligence 06 Sep 2026 10 min read

Accelerate LLM Generation with Speculative Decoding

Autoregressive language models generate text one token at a time. Even when an accelerator has substantial parallel compute available, the model normally cannot determine token 12 until token 11 is known. That dependency makes generation latency difficult to reduce simply by adding more parallel hardware. Speculative decoding attacks this bottleneck by doing cheap work ahead of the expensive model. A faster draft model proposes several future tokens. The full target model then evaluates those proposals together and accepts the portion that is consistent with its own distribution. With the appropriate acceptance-and-correction algorithm, this changes how generation is computed without changing the distribution that the target model defines.

Artificial Intelligence 05 Sep 2026 12 min read

Use Early Exits for Adaptive Neural Network Inference

A conventional neural network uses the same depth for every input. An obvious example and a difficult edge case both pass through all layers before the model returns a prediction. That fixed computation is simple to operate, but it can waste work when intermediate representations are already sufficient for some inputs. Early-exit inference makes computation adaptive. The model attaches prediction heads to intermediate layers. At each head, an exit policy decides whether the current prediction is reliable enough to return or whether the input should continue through deeper layers. Easy inputs can therefore use less computation while difficult inputs retain access to the full network.

Artificial Intelligence 05 Sep 2026 9 min read

Reduce LLM Decoding Latency with Speculative Decoding

Large language models generate text autoregressively: each new token depends on the tokens that came before it. That dependency makes ordinary decoding sequential. Even when a GPU has enough compute to process many token positions in parallel, the model normally discovers only one new token per decoding step. Speculative decoding tries to turn some of that sequential work into parallel verification. A faster draft model proposes several future tokens. The larger target model then scores those proposed positions together and decides which proposals can be accepted. When the draft predicts well, one expensive target-model pass can advance generation by multiple tokens.

Artificial Intelligence 05 Sep 2026 12 min read

Contrastive Decoding with Expert and Amateur Models

A language model can assign high probability to text that is fluent but unhelpfully generic, repetitive, or too close to an easy pattern. Changing temperature or top-p changes how tokens are sampled from one model’s distribution, but it does not ask a different question: which candidate tokens are especially characteristic of a stronger model rather than a weaker one? Contrastive decoding asks exactly that. It uses two language models at inference time: a stronger expert and a weaker amateur. A candidate is favored when the expert scores it well relative to the amateur, while a plausibility constraint prevents the decoder from choosing bizarre tokens merely because the amateur dislikes them even more.

Artificial Intelligence 05 Sep 2026 9 min read

Classifier-Free Guidance in Diffusion Models

A conditional diffusion model may understand a prompt and still produce samples that only weakly reflect it. During generation, developers therefore often want a way to push the denoising trajectory toward the condition without training a separate classifier for every prompt or label. Classifier-free guidance (CFG) is a widely used way to do that. At each denoising step, the model is evaluated with the condition and without it. The difference between those predictions gives a direction associated with the condition, and a guidance scale controls how strongly sampling moves along that direction.

Artificial Intelligence 05 Sep 2026 10 min read

Attention Sinks for Stable Streaming LLM Inference

Autoregressive language models normally reuse the keys and values of earlier tokens while generating the next token. This KV cache avoids recomputing the entire prefix at every decoding step, but its memory use grows with the cached sequence. A long-running chat, agent, or stream can therefore accumulate more cached state than a serving system wants to keep. A tempting fix is a sliding window: retain only the most recent tokens and evict everything older. For models trained with ordinary dense attention, however, abruptly dropping all early tokens can damage generation quality even when those old tokens do not appear semantically important.

Artificial Intelligence 03 Sep 2026 10 min read

Tokenization in Language Models

Language models do not read text as a sequence of words. Before a model can process a prompt, a tokenizer converts the text into a sequence of token IDs from a fixed vocabulary. The model operates on those IDs, and generated IDs are later converted back into text. This extra layer is easy to ignore because most model APIs accept ordinary strings. But tokenization affects how much text fits in a context window, how usage-based costs are calculated, how text is truncated or split, and why seemingly small formatting changes can alter model behavior.

Artificial Intelligence 03 Sep 2026 10 min read

Speculative Decoding for Faster LLM Inference

Autoregressive language models generate text sequentially. After processing the prompt, the model predicts a next token, appends that token to the sequence, and repeats the process. That dependency makes generation difficult to parallelize across time: token 101 cannot normally be generated until token 100 is known. Speculative decoding changes the amount of useful work performed during each expensive target-model step. A cheaper draft process proposes several future tokens, then the target model verifies those proposals together. When enough proposals are accepted, the application can advance by multiple tokens while invoking the large model fewer times.

Artificial Intelligence 03 Sep 2026 7 min read

LLM Quantization for Efficient Inference

Large language models can require substantial memory bandwidth and compute during inference. Quantization reduces those requirements by representing some model values with fewer bits than the floating-point formats commonly used during training. The idea sounds simple: store numbers with lower precision. In practice, quantization trades memory, latency, hardware support, implementation complexity, and model quality. Choose the configuration from measurements rather than assuming that fewer bits are always better. What quantization changes A neural network contains many numerical values, especially weights. A model stored with 16-bit weights needs roughly two bytes per weight before accounting for runtime buffers and other overhead. If those weights can instead be represented with 8 or 4 bits, their raw storage requirement falls substantially.

Artificial Intelligence 03 Sep 2026 7 min read

KV Caching in LLM Inference

Large language models generate text one token at a time. Without an optimization, every new token would force the model to repeat attention calculations for tokens it has already processed. KV caching avoids much of that repeated work. During inference, the model stores the key and value representations produced by attention layers for previous tokens. When generating the next token, it can reuse those stored representations instead of recomputing them from the beginning.

Artificial Intelligence 03 Sep 2026 9 min read

Beam Search for Sequence Generation

A model that generates text or another sequence makes a series of local decisions. At each step, it assigns scores or probabilities to possible next tokens. The simplest decoder chooses the most likely token, appends it, and repeats. That strategy is called greedy decoding. It is cheap and easy to understand, but an early choice that looks best by itself can lead to a worse complete sequence. Once greedy decoding commits to that choice, it cannot reconsider it.

Artificial Intelligence 03 Sep 2026 9 min read

Batch LLM Inference for Better Throughput

An LLM server can receive many requests at the same time, yet processing every request independently is often an inefficient way to use an accelerator. GPUs and similar hardware are designed to perform large amounts of parallel numerical work. A single small request may leave part of that capacity unused. Batching combines work from multiple requests so the model can process more of it together. This can improve total throughput, but it introduces an important trade-off: waiting to form a batch can delay individual requests, and requests with different sequence lengths do not all consume the same amount of work.

Artificial Intelligence 02 Sep 2026 5 min read

Control LLM Randomness with Temperature and Top-p

Large language models usually generate text one token at a time. At each step, the model assigns scores to possible next tokens, those scores become probabilities, and a decoding strategy chooses what comes next. Two common controls in that process are temperature and top-p. They are often described as creativity settings, but that description is incomplete. They change how the model samples from its probability distribution, which affects repeatability, diversity, and the chance of selecting lower-probability tokens.