Skip to content

Archive

LLM

76 articles
Artificial Intelligence 06 Sep 2026 10 min read

Rotary Position Embeddings in Transformers

A Transformer attention layer needs to know more than which tokens are present. Order matters: dog bites man and man bites dog contain the same words but express different relationships. Yet the dot products used by self-attention do not inherently know whether two token representations came from adjacent positions or opposite ends of a sequence. Rotary position embedding, usually shortened to RoPE, adds position information by rotating parts of the query and key vectors before their attention scores are computed. The useful consequence is subtle: each token receives a transformation based on its absolute position, while the dot product between two transformed vectors depends on their relative position.

Artificial Intelligence 06 Sep 2026 10 min read

Reduce KV Cache Size with Grouped-Query Attention

Autoregressive language models generate one token at a time. To avoid recomputing attention keys and values for every previous token at every step, inference systems usually keep those tensors in a key-value cache, or KV cache. This saves computation, but the cache grows with sequence length and can become a major memory cost when serving long contexts or many requests at once. One architectural choice has a direct effect on that cost: how many separate key and value heads the attention layer stores. Standard multi-head attention gives every query head its own key and value head. Grouped-query attention (GQA) keeps multiple query heads but lets groups of them share key and value heads.

Artificial Intelligence 06 Sep 2026 10 min read

Control LLM Repetition with Token Penalties

Language models sometimes repeat a phrase, return to the same point, or fall into a short loop even when the prompt asks for a concise answer. A common response is to increase randomness, but temperature changes the whole next-token distribution. That can reduce repetition while also making unrelated choices less predictable. Token penalties provide a more targeted control. They adjust the scores of tokens that have already appeared, making some repeated tokens less likely before the decoder chooses the next token. This can be useful for open-ended generation, but it is not a general quality switch: repeated tokens are often exactly what correct text requires.

Artificial Intelligence 06 Sep 2026 8 min read

Constrain LLM Output with Grammar-Guided Decoding

Asking a language model to return JSON, SQL, or another structured format creates a failure mode that ordinary prompting cannot remove: the model can understand the requested format and still generate a token that makes the output syntactically invalid. For applications that immediately parse model output, one missing quote or delimiter can turn an otherwise useful answer into an error. Retrying helps, but it spends more inference time without guaranteeing that the next attempt will parse.

Artificial Intelligence 06 Sep 2026 10 min read

Accelerate LLM Generation with Speculative Decoding

Autoregressive language models generate text one token at a time. Even when an accelerator has substantial parallel compute available, the model normally cannot determine token 12 until token 11 is known. That dependency makes generation latency difficult to reduce simply by adding more parallel hardware. Speculative decoding attacks this bottleneck by doing cheap work ahead of the expensive model. A faster draft model proposes several future tokens. The full target model then evaluates those proposals together and accepts the portion that is consistent with its own distribution. With the appropriate acceptance-and-correction algorithm, this changes how generation is computed without changing the distribution that the target model defines.

Artificial Intelligence 05 Sep 2026 9 min read

Reduce LLM Decoding Latency with Speculative Decoding

Large language models generate text autoregressively: each new token depends on the tokens that came before it. That dependency makes ordinary decoding sequential. Even when a GPU has enough compute to process many token positions in parallel, the model normally discovers only one new token per decoding step. Speculative decoding tries to turn some of that sequential work into parallel verification. A faster draft model proposes several future tokens. The larger target model then scores those proposed positions together and decides which proposals can be accepted. When the draft predicts well, one expensive target-model pass can advance generation by multiple tokens.

Artificial Intelligence 05 Sep 2026 10 min read

Pack Training Sequences to Reduce Padding Waste

Language-model training often processes sequences in fixed-size tensors. When examples have very different lengths, padding makes those tensors easy to batch but can leave many token positions doing little useful work. A batch that physically contains 8,000 positions may contain far fewer than 8,000 real training tokens. Sequence packing reduces this waste by placing multiple shorter examples into the same fixed-length training sequence. The idea is simple; the semantics are not. If packing accidentally lets one example attend to another, predicts across boundaries that should be independent, or assigns incorrect position IDs, the training objective changes rather than merely becoming more efficient.

Artificial Intelligence 05 Sep 2026 9 min read

Diversify RAG Context with Maximum Marginal Relevance

A retrieval-augmented generation (RAG) system can retrieve highly relevant passages and still build poor context. The problem appears when several top results say almost the same thing. Sending all of them to the language model consumes context without adding much evidence, while a slightly lower-ranked passage containing a different useful fact may be excluded. Maximum marginal relevance (MMR) is a selection strategy for this situation. Instead of choosing passages only by their relevance to the query, MMR repeatedly chooses a passage that is both relevant and sufficiently different from passages already selected.

Artificial Intelligence 05 Sep 2026 12 min read

Contrastive Decoding with Expert and Amateur Models

A language model can assign high probability to text that is fluent but unhelpfully generic, repetitive, or too close to an easy pattern. Changing temperature or top-p changes how tokens are sampled from one model’s distribution, but it does not ask a different question: which candidate tokens are especially characteristic of a stronger model rather than a weaker one? Contrastive decoding asks exactly that. It uses two language models at inference time: a stronger expert and a weaker amateur. A candidate is favored when the expert scores it well relative to the amateur, while a plausibility constraint prevents the decoder from choosing bizarre tokens merely because the amateur dislikes them even more.

Artificial Intelligence 04 Sep 2026 9 min read

Perplexity for Language Model Evaluation

A language model can assign high probability to likely text and low probability to unlikely text, but developers still need a compact way to summarize that behavior across many tokens. Perplexity is one common metric for this job. Perplexity is useful when comparing probabilistic language models on the same evaluation data under compatible tokenization and scoring rules. It is much less useful as a general score for whether generated answers are correct, helpful, safe, or well written.

Artificial Intelligence 04 Sep 2026 9 min read

Interpret Language Model Perplexity Correctly

A language model can improve on its training objective while still leaving an important question unanswered: how well does it predict text it did not train on? Perplexity is a compact way to measure that predictive fit for autoregressive language models, but the number is easy to misuse. A lower perplexity can mean that a model assigns higher probability to held-out text. It does not automatically mean that the model follows instructions better, reasons more reliably, hallucinates less, or produces more useful answers. Comparisons can also become misleading when tokenization, evaluation data, or context handling differs.

Artificial Intelligence 03 Sep 2026 6 min read

Use Few-Shot Prompting with Effective Examples

A prompt can explain a task with instructions, but sometimes examples communicate the desired behavior more precisely. Few-shot prompting places a small number of input-output demonstrations in the model’s context before the real input. This technique is useful when a task has a specific output format, subtle classification boundary, naming convention, or transformation rule that is difficult to describe completely in prose. The model is not retrained by these examples. Instead, it uses the demonstrations as part of the current context when generating the next response.

Artificial Intelligence 03 Sep 2026 7 min read

Tokenization in Large Language Models

Large language models do not read text as words or characters in the way people do. Before text reaches the model, a tokenizer converts it into a sequence of discrete units called tokens and maps those tokens to numerical identifiers. Tokenization is easy to overlook because most model APIs perform it automatically. Yet token boundaries affect context-window usage, inference cost, truncation, multilingual behavior, and even whether two visually similar strings are represented in similar ways. These details explain many LLM behaviors that otherwise look inconsistent.

Artificial Intelligence 03 Sep 2026 10 min read

Speculative Decoding for Faster LLM Inference

Autoregressive language models generate text sequentially. After processing the prompt, the model predicts a next token, appends that token to the sequence, and repeats the process. That dependency makes generation difficult to parallelize across time: token 101 cannot normally be generated until token 100 is known. Speculative decoding changes the amount of useful work performed during each expensive target-model step. A cheaper draft process proposes several future tokens, then the target model verifies those proposals together. When enough proposals are accepted, the application can advance by multiple tokens while invoking the large model fewer times.

Artificial Intelligence 03 Sep 2026 7 min read

Reduce LLM Hallucinations with Grounding and Verification

Large language models can produce fluent answers that contain incorrect facts, invented details, or unsupported claims. This behavior is commonly called hallucination. Hallucinations are not simply random mistakes. A language model generates tokens that are plausible given its input and learned parameters. Plausible text is not necessarily true text, especially when the model lacks reliable evidence for the question being asked. For developers, the practical goal is therefore not to find a single setting that eliminates hallucinations. It is to design the application so that factual claims are grounded in appropriate evidence, uncertainty is handled explicitly, and important outputs are verified.

Artificial Intelligence 03 Sep 2026 9 min read

Mixture-of-Experts Models

A neural network does not have to use every parameter for every input. A mixture-of-experts (MoE) layer takes advantage of this idea by keeping several expert networks and using a router to select only a small subset for each token. This separates total parameter count from the number of parameters active for one token. That can increase model capacity without making the arithmetic performed for every token grow in direct proportion to the total number of expert parameters.

Artificial Intelligence 03 Sep 2026 10 min read

LoRA for Parameter-Efficient Fine-Tuning

Fine-tuning a large model does not always require updating every model parameter. Low-Rank Adaptation (LoRA) takes advantage of this idea by keeping the original model weights frozen and learning much smaller matrices that modify selected layers. For developers, LoRA changes what must be trained, stored, and moved between experiments—not just the size of the fine-tuning job. That distinction determines when it is useful, what it does not save, and how adapter choices affect model behavior.

Artificial Intelligence 03 Sep 2026 7 min read

LLM Quantization for Efficient Inference

Large language models can require substantial memory bandwidth and compute during inference. Quantization reduces those requirements by representing some model values with fewer bits than the floating-point formats commonly used during training. The idea sounds simple: store numbers with lower precision. In practice, quantization trades memory, latency, hardware support, implementation complexity, and model quality. Choose the configuration from measurements rather than assuming that fewer bits are always better. What quantization changes A neural network contains many numerical values, especially weights. A model stored with 16-bit weights needs roughly two bytes per weight before accounting for runtime buffers and other overhead. If those weights can instead be represented with 8 or 4 bits, their raw storage requirement falls substantially.

Artificial Intelligence 03 Sep 2026 7 min read

KV Caching in LLM Inference

Large language models generate text one token at a time. Without an optimization, every new token would force the model to repeat attention calculations for tokens it has already processed. KV caching avoids much of that repeated work. During inference, the model stores the key and value representations produced by attention layers for previous tokens. When generating the next token, it can reuse those stored representations instead of recomputing them from the beginning.

Artificial Intelligence 03 Sep 2026 7 min read

Improve RAG Retrieval with Reranking

Retrieval-augmented generation (RAG) depends on finding useful evidence before asking a language model to answer. A vector search can retrieve candidates quickly, but the nearest vectors are not always the passages that best answer the user’s question. Reranking adds a second relevance step. The system first retrieves a reasonably broad candidate set with a fast method, then applies a more precise model to reorder those candidates before selecting context for the LLM.

Artificial Intelligence 03 Sep 2026 10 min read

Design Reliable LLM Tool Calls with Structured Outputs

Giving a language model access to tools changes what an AI application can do. Instead of only producing text, the model can request a database lookup, search a document index, calculate a value, or trigger an application operation. The difficult part is not exposing a function. It is deciding which responsibilities belong to the model and which must remain under application control. A useful mental model is: model proposes an action application validates the proposal application decides whether to execute it tool returns data model explains or uses the result This separation makes tool-calling systems easier to reason about. The model handles interpretation and selection. Deterministic application code handles authorization, validation, execution, and state changes.

Artificial Intelligence 03 Sep 2026 6 min read

Choose Between Prompting, RAG, and Fine-Tuning

When an AI application produces weak results, teams often jump directly to fine-tuning. That can be the right choice, but many problems are cheaper and easier to solve with better prompting or retrieval-augmented generation (RAG). The three approaches change different parts of the system. Prompting changes the instructions and context given at inference time. RAG supplies relevant external information at inference time. Fine-tuning changes the model’s learned parameters through additional training.

Artificial Intelligence 03 Sep 2026 9 min read

Batch LLM Inference for Better Throughput

An LLM server can receive many requests at the same time, yet processing every request independently is often an inefficient way to use an accelerator. GPUs and similar hardware are designed to perform large amounts of parallel numerical work. A single small request may leave part of that capacity unused. Batching combines work from multiple requests so the model can process more of it together. This can improve total throughput, but it introduces an important trade-off: waiting to form a batch can delay individual requests, and requests with different sequence lengths do not all consume the same amount of work.

Artificial Intelligence 02 Sep 2026 5 min read

Semantic Caching for LLM Applications Without Serving Stale Answers

Large language model requests are expensive compared with ordinary cache lookups. When users repeatedly ask questions with slightly different wording, an exact string cache misses even though the intended answer may be identical. A semantic cache uses vector similarity to decide whether a new request is close enough to a previous request that its answer can be reused. The idea is attractive, but the difficult part is not storing embeddings. It is deciding when reuse is actually safe.