Skip to content

Topic archive

Artificial Intelligence

Artificial intelligence includes machine learning, deep learning, and systems that apply models to practical tasks.

396 articles
Artificial Intelligence 23 Sep 2026 5 min read

Attention Entropy Measures Concentration, Not Causal Importance

An attention head can place most of its probability mass on one position and still provide little evidence that this position controls the final model output. The entropy of its attention weights captures concentration, not causal influence. That distinction matters when attention maps are inspected as diagnostics. Entropy can reveal whether a head spreads mass broadly or focuses it narrowly for a given query. It cannot, by itself, establish that the highest-weight token carries the feature responsible for a downstream prediction.

Artificial Intelligence 23 Sep 2026 6 min read

Additive Logit Bias Changes Token Odds Before Sampling

A decoder can favor or suppress a token without changing model weights. Add a constant to that token’s logit before softmax, and its probability changes relative to the rest of the vocabulary. The operation is simple, but its effect depends on where the bias enters the decoding pipeline and on every transformation that follows it. This makes additive logit bias useful as an inference control, but not as a general semantic constraint. It changes a score used by the decoder. It does not rewrite the model’s internal representation or guarantee that a concept disappears from generated text.

Artificial Intelligence 23 Sep 2026 5 min read

Activation Outliers Can Set the Scale for an Entire Quantization Group

A quantizer with a fixed integer width has only a finite set of representable codes. When many activations share one scale, a single value with much larger magnitude can force that scale to cover a wider real-valued range. The remaining values then occupy fewer useful code intervals around the region where they are concentrated. This behavior is not a generic statement that quantization fails in the presence of large numbers. It follows from a specific coupling: values inside the same quantization group share parameters that map real numbers to integer codes.

Artificial Intelligence 23 Sep 2026 5 min read

Activation Outliers Can Dominate Per-Tensor Quantization Scale

A per-tensor quantizer maps every value in an activation tensor through one shared scale. That coupling matters when most activations occupy a narrow interval but a few values have much larger magnitude. The large values can determine the scale, while the dense central region is represented with coarser spacing than its own range would require. This is not a statement that every large activation is erroneous or removable. An outlier may carry useful model state. The issue is numerical: one scale has to cover values with very different magnitudes.

Artificial Intelligence 22 Sep 2026 5 min read

Top-p Sampling Rebuilds Its Candidate Set at Every Token

Top-p sampling does not keep a fixed shortlist of tokens throughout generation. At each decoding step, the model produces a new logit vector, that vector becomes a probability distribution, and the sampler forms a new candidate set whose cumulative probability mass reaches the configured threshold. The consequence is easy to miss in serving code: the same top_p value can admit two tokens at one step and dozens at another. The parameter controls probability mass, not candidate count.

Artificial Intelligence 22 Sep 2026 6 min read

Speculative Decoding Couples Draft Speed with Acceptance Rate

Autoregressive generation normally advances one accepted token at a time. Each new token extends the prefix, so the next target-model evaluation depends on the token selected at the preceding position. Speculative decoding changes the execution schedule: a cheaper draft process proposes several future tokens, then the target model evaluates those positions together and decides how much of the proposal can be retained. That rearrangement can reduce the number of serial target-model calls per emitted token. It does not make verification free, and a longer draft block is not automatically better. The useful operating point depends on how quickly proposals are produced, how often they survive target verification, and what the serving stack spends on rejected work.

Artificial Intelligence 22 Sep 2026 7 min read

Prefix KV Cache Reuse Depends on Exact Token History

A KV cache entry is not a reusable representation of arbitrary text that happens to look similar. For an autoregressive transformer, cached keys and values are intermediate states produced for a specific token prefix under a specific execution context. Reusing them is valid only when the new request reaches the same state boundary. That boundary is stricter than matching visible characters. Tokenization, token order, position handling, model identity, adapter state, and other inputs that affect hidden states can all determine whether a cached prefix still represents the computation required by the new request.

Artificial Intelligence 22 Sep 2026 6 min read

Padding Masks Do Not Remove Padding Compute in Dense Attention

A batch can contain two prompts with very different token counts yet represent both with the same rectangular tensor. The shorter prompt is extended with padding so its tensor shape matches the longest sequence in the batch. An attention mask can stop those padded positions from contributing to attention probabilities, but that semantic exclusion does not imply that dense kernels skip every operation associated with the padded rows and columns.

Artificial Intelligence 22 Sep 2026 6 min read

Cosine Similarity Discards Embedding Magnitude

Two embedding vectors can point in the same direction while having very different norms. Cosine similarity gives those vectors the same directional score. A raw dot product does not. That distinction becomes an implementation boundary when a retrieval system changes index metrics, normalizes vectors at ingestion, or mixes embeddings produced by different pipelines. The issue is not that one metric is universally preferable. The relevant question is whether vector magnitude carries information that the scoring contract intends to preserve. Once vectors are normalized to unit length, that information is removed from the similarity calculation.

Artificial Intelligence 19 Sep 2026 8 min read

Where PEGASUS-XSum Inference Time Goes

google/pegasus-xsum is a summarization checkpoint, not a compact text utility. Its latency follows directly from the work performed by a large Transformer encoder-decoder: first encode the source document, then run the decoder repeatedly until the summary is complete. On a CPU, the second phase is usually the part that makes a short output feel disproportionately expensive. The original PEGASUS work describes a Transformer encoder-decoder pretrained with Gap Sentences Generation and reports a 568M-parameter best model. The XSum checkpoint is fine-tuned for highly abstractive single-document summarization. That combination is useful when summary quality matters, but it also means inference has substantially more machinery than extracting a few source sentences or running a small classifier.

Artificial Intelligence 19 Sep 2026 6 min read

Prefix Caching Reuses KV State Only Across Identical Prompt Prefixes

Autoregressive transformer serving often repeats the same prompt prefix across requests: a system message, a long document header, or a fixed tool schema may precede user-specific text. Prefix caching stores the key-value state produced by that shared prefix so a later request can resume computation from the cached boundary instead of recomputing the entire prefix. The useful boundary is narrower than “similar prompts.” Reuse depends on the exact token sequence and on model state that affects the cached activations. A one-character text edit may preserve most tokens, shift tokenization near the edit, or change every token after a formatting boundary. The cache can only reuse the portion whose effective input is still identical.

Artificial Intelligence 19 Sep 2026 6 min read

Ollama Gemma 3 270M Memory Is More Than the Model File

Ollama lists gemma3:270m at about 292 MB. That number is useful for storage planning, but it is not a RAM requirement. It describes the packaged model data for the default Ollama variant, which uses Q8_0 quantization. Once inference starts, the runtime also needs memory for model metadata, execution buffers, token state, and the key-value cache used by attention. That distinction matters on small machines. A device with 512 MB of RAM may appear large enough when compared only with a 292 MB model file, yet the remaining memory must also accommodate Ollama and the operating system. The context configuration can move the total substantially.

Artificial Intelligence 19 Sep 2026 7 min read

Monocular Depth Estimation with MiDaS and DPT

A single RGB camera records image coordinates and color, not the physical distance from the lens to every visible surface. Monocular depth models infer the missing depth structure from visual evidence learned during training. That distinction matters when using MiDaS or DPT: a convincing depth map does not automatically mean that pixel values are distances in meters. MiDaS is an open-source project for robust monocular relative depth estimation. DPT, or Dense Prediction Transformer, is an architecture for dense prediction that has also been used as a backbone in MiDaS models. Both make single-camera depth estimation practical, but neither changes the geometric ambiguity inherent in one unconstrained RGB image.

Artificial Intelligence 19 Sep 2026 8 min read

Chunk Long Prefills to Limit Decode Stalls in LLM Serving

A long prompt can occupy an accelerator for a much larger scheduling interval than a single decode iteration. When a serving engine mixes new prefills with requests that are already generating tokens, that difference can show up as irregular time between output tokens. The model has not changed; the interference comes from how two distinct inference phases share execution time. Prefill processes a prompt and builds the key-value state required by later causal attention. Decode then extends the sequence autoregressively, usually one new token per active request per iteration. Those phases place different pressure on hardware, so treating them as interchangeable scheduling units can produce avoidable stalls.

Artificial Intelligence 18 Sep 2026 5 min read

Treat the Logit Lens as a Readout, Not a Causal Trace

A transformer can expose an intermediate residual state that strongly favors a token under the model’s final vocabulary projection, then produce a different token after later blocks run. The logit lens makes that intermediate preference visible. It does not establish that the preference caused the final output. That boundary matters when developers use layer-by-layer token rankings to inspect model behavior. The logit lens is a readout: it asks what the model’s output head would report if applied to an intermediate representation. The actual forward pass asks a different question because every remaining block can transform that representation before the final readout.

Artificial Intelligence 18 Sep 2026 5 min read

Recompute BatchNorm Statistics After Weight Averaging

Averaging two neural-network checkpoints can produce a useful parameter vector, yet leave BatchNorm running statistics tied to a different network. The weights define one set of activations; the stored running means and variances may describe activations produced by earlier weights. Inference then combines state from two different points in parameter space. This mismatch is easy to miss because BatchNorm running statistics are buffers rather than trainable parameters in common implementations. A parameter-averaging routine can handle every weight correctly and still produce an internally inconsistent inference state.

Artificial Intelligence 17 Sep 2026 6 min read

Use Selective Classification to Trade Coverage for Error Rate

A classifier usually returns a label for every input, even when its score distribution is nearly tied or the input sits far from familiar data. Selective classification changes that interface: the system may return a prediction or abstain. The acceptance rule then determines both how many inputs receive predictions and how often those accepted predictions are wrong. This is not the same as making the classifier intrinsically more accurate. Abstention moves some cases out of the automatic-decision set. Its value depends on whether the selection score ranks difficult cases well enough for rejected inputs to contain a disproportionate share of errors.

Artificial Intelligence 17 Sep 2026 6 min read

Teacher Forcing Creates a Prefix Distribution Gap

Autoregressive models predict the next token from a prefix. During teacher-forced training, that prefix usually comes from the reference sequence. During generation, it contains tokens emitted by the model itself. A prediction error can therefore change the context used for every later prediction. This difference is often called exposure bias. The useful engineering detail is more specific: training and generation can present different prefix distributions to the same conditional predictor. Token-level validation on clean reference prefixes does not fully characterize behavior after the model enters a prefix that its training data rarely presented.

Artificial Intelligence 17 Sep 2026 6 min read

Retain Attention Sinks in Bounded KV Caches

A bounded KV cache seems to invite a simple eviction rule: once the cache is full, discard the oldest key-value pair and keep the most recent tokens. For some decoder-only transformers, that rule can degrade generation sharply after the sequence moves beyond the retained window. The failure is not explained only by missing old semantic content. Early tokens can receive substantial attention even when their text carries little useful information for the current prediction.

Artificial Intelligence 17 Sep 2026 6 min read

Limit Parameter Drift with Elastic Weight Consolidation

Fine-tuning a neural network on a new task can move parameters away from values that supported an earlier task. The new objective has no inherent reason to preserve those earlier behaviors when the earlier data is absent from the update. Elastic weight consolidation, commonly abbreviated EWC, adds a parameter-space constraint intended to reduce that drift. The constraint is selective rather than uniform. Parameters estimated to be more consequential for the earlier task receive a larger penalty for moving, while parameters assigned lower importance can change more freely. That distinction is the central mechanism; EWC is not simply weight decay around zero.

Artificial Intelligence 17 Sep 2026 5 min read

Evaluate Probabilistic Classifiers with the Brier Score

Two classifiers can produce the same predicted labels and the same accuracy while assigning very different probabilities to those labels. A system that emits 0.51 for every correct binary decision is not making the same probabilistic claim as one that emits 0.99, even though thresholded accuracy may treat them identically. The Brier score keeps that distinction visible. It measures squared error between predicted probabilities and observed outcomes, so both the selected class and the probability assigned to each outcome affect the result. This makes it useful when downstream code consumes probabilities for ranking, thresholds, abstention, or expected-cost decisions.

Artificial Intelligence 17 Sep 2026 6 min read

Conformal Prediction Sets Need Exchangeable Calibration Data

A classifier that emits 0.93 for one class does not automatically provide a statistical statement that the class is correct with probability 0.93. Conformal prediction takes a different route: it uses held-out labeled examples to construct a set of candidate labels with a target marginal coverage level. For split conformal classification, the base model can remain fixed. The guarantee comes from ranking a test example’s conformity or nonconformity score against scores computed on an exchangeable calibration sample. That assumption is the part that gives the coverage statement its scope.

Artificial Intelligence 17 Sep 2026 6 min read

Bound Gradient Updates with Global Norm Clipping

A single optimization step can contain gradients whose combined magnitude is far larger than the surrounding steps. If those gradients are passed directly to an optimizer, the resulting parameter update can move the model into a very different region of parameter space. Global norm clipping places a bound on the gradient magnitude before the optimizer consumes it. The mechanism is simple, but its behavior is easy to misread. It does not cap every gradient element independently, and it does not guarantee a fixed parameter-update norm for adaptive optimizers. It rescales the collected gradient vector when a chosen norm crosses a threshold.

Artificial Intelligence 17 Sep 2026 7 min read

Balance Token Routing in Sparse Mixture-of-Experts Models

A sparse mixture-of-experts layer can contain many expert networks while evaluating only a small subset for each token. The router makes that sparsity possible: it assigns scores to experts, selects a limited set, and sends each token through the selected computation paths. That selection is not only an optimization detail. If many tokens concentrate on a few experts, some devices can receive much more work than others, capacity limits can discard or redirect assignments, and experts that receive little traffic get fewer task gradients. Router balance therefore affects both computation and the function represented by the model.