Skip to content

Archive

Artificial Intelligence

391 articles
Artificial Intelligence 08 Sep 2026 9 min read

Use Test-Time Augmentation for More Stable Predictions

A classifier can change its prediction because an object moved a few pixels, an image was cropped differently, or another harmless transformation changed the input representation. If those transformations should not change the correct answer, that sensitivity is undesirable. Test-time augmentation (TTA) addresses this problem by running the same trained model on several meaning-preserving versions of an input and combining their predictions. Instead of asking the model for one view of the evidence, TTA asks it to evaluate several valid views.

Artificial Intelligence 08 Sep 2026 10 min read

Use Self-Conditioning in Diffusion Models

A diffusion model repeatedly turns a noisy state into a cleaner one. Each denoising call normally receives the current noisy sample, a noise level or timestep, and any external condition such as a text embedding. Yet the previous call has already produced useful information about what the clean sample may look like. Throwing that estimate away means the next call must reconstruct similar information again from the new noisy state.

Artificial Intelligence 08 Sep 2026 10 min read

Transformer Attention Weights: What They Show and What They Do Not

Transformer attention maps are visually compelling. A token appears to assign most of its attention to another token, so it is tempting to conclude that the second token caused the model’s prediction. That conclusion is stronger than the data supports. An attention weight has a precise local meaning: inside one attention operation, it controls how strongly a query mixes information from available value vectors. A complete transformer prediction, however, also depends on value vectors, residual connections, feed-forward layers, normalization, later layers, and often many attention heads. A large weight is therefore evidence about one routing operation, not a complete causal explanation.

Artificial Intelligence 08 Sep 2026 10 min read

Train Neural Networks with Curriculum Learning

Most training pipelines treat the dataset as a fixed pool and repeatedly shuffle it. That is a strong default: it is simple, exposes the model to the full data distribution, and avoids assumptions about which examples should come first. But some learning problems have a useful notion of progression. A model may learn basic patterns more reliably before it is asked to handle noisy, ambiguous, or structurally difficult examples. Curriculum learning makes that progression explicit. Instead of changing the model architecture or the loss, it changes which training examples are emphasized at different stages of training. A common curriculum begins with easier examples and gradually introduces harder ones.

Artificial Intelligence 08 Sep 2026 11 min read

Trade Compute for Memory with Activation Checkpointing

A neural network can fit comfortably in accelerator memory for inference and still run out of memory during training. The reason is that training needs more than model weights. Backpropagation also needs intermediate values from the forward pass, and those activations can consume a large share of memory in deep models or with long sequences and large batches. Activation checkpointing reduces that memory pressure by deliberately not keeping every intermediate activation. Instead, training saves selected checkpoints and recomputes missing forward-pass values when the backward pass needs them. The trade is straightforward: keep fewer activations in memory, but perform extra computation.

Artificial Intelligence 08 Sep 2026 9 min read

Trace Token Influence with Attention Rollout

Looking at one Transformer attention matrix can answer a local question: which positions a token attends to in that layer. It does not directly tell you how much an input token can influence a representation several layers later. The reason is mixing. After one layer, a token representation already contains information gathered from other positions. The next layer attends to those mixed representations, not to untouched input tokens. Residual connections add another path that carries each representation forward. Reading only the final layer therefore skips the paths through earlier layers.

Artificial Intelligence 08 Sep 2026 9 min read

The Softmax Bottleneck in Language Models

A language model can have a powerful network behind it and still be constrained by the layer that turns its hidden state into next-token probabilities. In the common linear-softmax output layer, that constraint has a precise form: across many contexts, the model can represent only a limited family of log-probability patterns. This limitation is known as the softmax bottleneck. It is not a claim that softmax itself is defective, nor does it mean every modern language model is visibly harmed by it. It is a structural result about a particular output parameterization.

Artificial Intelligence 08 Sep 2026 9 min read

Stabilize Neural Network Weights with Exponential Moving Averages

A neural network’s parameters rarely move smoothly toward their final values. Mini-batch training produces noisy updates: one batch may push a weight in one direction, while the next pushes it partly back. The final checkpoint therefore represents one point on a noisy training path, not necessarily the most useful point near the end of that path. An exponential moving average (EMA) of model weights keeps a second set of parameters that changes more slowly than the actively trained model. Recent training states contribute more than old ones, but no single update immediately replaces the averaged weights.

Artificial Intelligence 08 Sep 2026 13 min read

Rewrite RAG Queries Without Losing User Intent

A retrieval-augmented generation system often searches with the user’s latest message. That works for self-contained questions, but conversational questions frequently depend on earlier turns. Consider a support assistant. The user first asks about a failed database migration, discusses PostgreSQL for several turns, and then asks: Does the rollback command work on version 16 too? Searching that sentence literally may retrieve pages about unrelated rollback commands because the query does not say what is being rolled back. A query rewriter can turn the conversational message into a self-contained retrieval query such as:

Artificial Intelligence 08 Sep 2026 11 min read

Retrieve Multi-Hop Evidence with Graph RAG

Retrieval-augmented generation (RAG) usually starts with a simple idea: find text chunks similar to a question, place the most relevant chunks in the model’s context, and ask the model to answer from that evidence. This works well when the answer is stated in one passage or in several passages that are independently easy to retrieve. Some questions are harder because the useful evidence is connected by relationships, not just by similar wording. A developer may need to answer, “Which service depends on the library maintained by the team that owns the payment API?” No single chunk has to contain all of those words. The answer may require following several links across services, libraries, teams, and APIs.

Artificial Intelligence 08 Sep 2026 8 min read

Residual Connections in Deep Neural Networks

Making a neural network deeper gives it more transformations to work with, but depth alone does not make optimization easy. A stack of layers must learn useful transformations while gradients travel backward through every stage. As the stack grows, that optimization path can become difficult even when the deeper model has enough capacity to represent a good solution. Residual connections change what a block is asked to learn. Instead of making the block produce an entirely new representation, they let it learn a change to the representation it already received. The original input travels along a shortcut and is added back to the learned branch.

Artificial Intelligence 08 Sep 2026 9 min read

Regularize Residual Networks with Stochastic Depth

Deep residual networks can overfit even when their skip connections make optimization manageable. Standard dropout can regularize individual activations, but residual architectures offer another useful unit to randomize: the entire residual branch. Stochastic depth randomly removes selected residual branches during training while keeping the skip path intact. A training example may therefore pass through a slightly shallower effective network on one step and the full set of blocks on another. At inference time, every residual branch is normally active.

Artificial Intelligence 08 Sep 2026 10 min read

Model Soups for Combining Fine-Tuned Models

A hyperparameter sweep often leaves you with several fine-tuned models that are individually useful. The usual workflow keeps the checkpoint with the best validation score and discards the rest. An ensemble can use several checkpoints, but then every request may require multiple model evaluations, increasing inference cost and operational complexity. A model soup offers a third option: average the parameters of compatible fine-tuned models and deploy the resulting parameter set as one model. The technique is simple, but its simplicity can be misleading. Parameter averaging is meaningful only when the checkpoints are sufficiently compatible, and the averaged model still needs independent evaluation.

Artificial Intelligence 08 Sep 2026 12 min read

Migrate Embedding Models Without Breaking Retrieval

Changing an embedding model can look like a routine dependency upgrade. Replace the model identifier, deploy the service, and continue querying the existing vector index. That approach can silently damage retrieval. An embedding is meaningful relative to the representation space produced by its model. If stored document vectors came from one model while new query vectors come from another, their coordinates generally do not have a shared meaning. Matching dimensions are not enough to make the vectors compatible.

Artificial Intelligence 08 Sep 2026 9 min read

Masked Autoencoders for Visual Representation Learning

Labeled image datasets are expensive to build, but unlabeled images are often plentiful. A useful pretraining strategy is therefore to create a learning signal from each image itself instead of asking a human to annotate it. A masked autoencoder (MAE) does this by hiding part of an image and training a model to reconstruct the missing content. The reconstruction task is not usually the final product. Its purpose is to make the encoder learn visual representations that can later support tasks such as classification or detection.

Artificial Intelligence 08 Sep 2026 9 min read

Load Balancing in Mixture-of-Experts Models

A mixture-of-experts model can contain many expert networks while activating only a small subset for each token. That sparse computation is attractive because the model can have more parameters without evaluating every parameter for every token. But sparsity creates a new problem: the router can send too many tokens to the same experts. If one expert receives most of a batch while others sit nearly idle, the model does not get the practical benefit that its expert count suggests. In systems with fixed expert capacity, overloaded experts can also overflow, so some token-to-expert assignments cannot be processed as intended.

Artificial Intelligence 08 Sep 2026 11 min read

Filter Synthetic Training Data with Rejection Sampling

Generating synthetic examples is easy; generating synthetic examples that are worth training on is harder. A language model can produce thousands of candidate answers, but blindly adding them to a training set can reinforce factual errors, weak reasoning, unwanted style, or artifacts of the generator itself. Rejection sampling provides a simple mental model for controlling that pipeline: generate one or more candidates, evaluate each candidate with an acceptance rule, and keep only candidates that pass. The acceptance rule might use deterministic checks, a learned reward model, another language model, human review, or a combination of signals.

Artificial Intelligence 08 Sep 2026 8 min read

Estimate LLM Uncertainty with Semantic Entropy

A language model can produce a fluent answer even when it is uncertain. Token probabilities help describe uncertainty during generation, but they can be misleading at the answer level because many different strings can express the same meaning. Consider a question whose correct answer is Paris. A model might generate Paris, The answer is Paris, and France's capital is Paris. These strings differ, yet they represent the same answer. Treating them as three unrelated outcomes exaggerates the apparent uncertainty.

Artificial Intelligence 08 Sep 2026 10 min read

Dynamical Isometry in Neural Networks

A deep neural network can look well scaled one layer at a time and still be difficult to optimize. Signals pass through many transformations, and small expansions or contractions can multiply with depth. By the time a gradient travels through the whole network, some directions may have nearly disappeared while others have been amplified dramatically. Dynamical isometry gives a precise way to reason about this problem. Instead of asking only whether the average gradient magnitude is reasonable, it asks how the network transforms different directions in its input space. The relevant object is the input-output Jacobian, and its singular values provide the main measurements.

Artificial Intelligence 08 Sep 2026 9 min read

Detect Out-of-Distribution Inputs with Energy Scores

A classifier can be highly accurate on its test set and still behave confidently on inputs that are unlike anything it was trained to recognize. A product classifier trained on shoes, bags, and watches may receive a photo of a bicycle and still be forced to choose one of its known classes. That creates a deployment problem: ordinary classification answers which known class looks most likely, but many systems also need to ask whether this input resembles the data on which the classifier was validated.

Artificial Intelligence 08 Sep 2026 8 min read

Deep Ensembles for Model Uncertainty

A neural network can return a confident prediction even when an input is unfamiliar or ambiguous. Looking only at one model’s largest probability can therefore hide an important question: would another plausible model, trained on the same task, make the same decision? A deep ensemble helps answer that question by training several neural networks independently and combining their predictions. The combined prediction can improve robustness in some settings, while disagreement among members provides a practical uncertainty signal. It is not a guarantee that the prediction is correct, and it does not detect every kind of uncertainty.

Artificial Intelligence 08 Sep 2026 10 min read

Avoid Tokenization Boundary Failures in LLM Generation

A language model application usually treats a prompt as text: provide a prefix, then ask the model to continue it. The model sees something more specific. Its tokenizer first converts that text into tokens, and the end of the prompt forces the last token to end at exactly that position. That detail can matter when the prompt ends at a character position that would normally fall inside a larger token if the prompt and its continuation were tokenized together. The resulting tokenization boundary problem, also called the partial token problem, can make an otherwise natural continuation unexpectedly unlikely.

Artificial Intelligence 07 Sep 2026 12 min read

Xavier and He Initialization for Neural Networks

A deep neural network can fail before learning has had a fair chance. If its initial weights make activations or gradients shrink layer after layer, useful signals can become tiny. If those quantities grow too much, training can become unstable. The optimizer may receive a problem that is unnecessarily difficult even though the architecture and data are otherwise reasonable. Weight initialization tries to start the network in a numerically useful regime. Two common schemes are Xavier initialization, also called Glorot initialization, and He initialization, also called Kaiming initialization. Both choose the scale of random weights from the size of a layer, but they make different assumptions about how signals pass through the activation function.

Artificial Intelligence 07 Sep 2026 8 min read

Weight Tying in Language Models

A language model needs to solve two related problems with vocabulary-sized parameters. At the input, it must turn each token ID into a vector. At the output, it must turn a hidden state into one score for every possible next token. A straightforward architecture gives these two operations separate parameter matrices. That works, but it can be expensive when the vocabulary and hidden dimension are large. Weight tying is a simple architectural idea: use the same learned matrix for the input token embeddings and the output token projection when their shapes and semantics are compatible.