Skip to content

Topic archive

Artificial Intelligence

Artificial intelligence includes machine learning, deep learning, and systems that apply models to practical tasks.

396 articles
Artificial Intelligence 06 Sep 2026 10 min read

Steer Language Models by Editing Hidden Activations

Prompting changes what a language model reads. Fine-tuning changes its parameters. There is another, more experimental way to influence generation: change the model’s internal activations while it runs. This technique is commonly called activation steering or activation engineering. A simple version measures how hidden representations differ between examples that express opposite properties, turns that difference into a steering vector, and adds a scaled version of the vector during inference. The model weights stay unchanged.

Artificial Intelligence 06 Sep 2026 10 min read

Stabilize Neural Network Evaluation with Exponential Moving Average Weights

A neural network’s final training step is not necessarily its most useful checkpoint. Stochastic optimization keeps moving the parameters as it follows noisy mini-batch gradients, so two nearby checkpoints can behave slightly differently even when training is otherwise healthy. An exponential moving average (EMA) of model weights gives you a second set of parameters that changes more smoothly. Instead of evaluating only the latest training weights, you maintain a weighted history in which recent weights matter most and older weights gradually fade away.

Artificial Intelligence 06 Sep 2026 11 min read

Sequence Parallelism for Lower Transformer Activation Memory

Large transformer training can run out of accelerator memory even after the model’s weights are split across several devices. The reason is easy to miss: tensor parallelism can shard expensive matrix multiplications while some intermediate activations remain replicated on every worker in the tensor-parallel group. Sequence parallelism removes part of that replication. For operations that work independently on each token, it partitions activations along the sequence dimension so each tensor-parallel worker keeps only a slice of the tokens. The workers temporarily reconstruct or reduce data where the tensor-parallel computation requires communication, then return to sequence-sharded activations.

Artificial Intelligence 06 Sep 2026 10 min read

Rotary Position Embeddings in Transformers

A Transformer attention layer needs to know more than which tokens are present. Order matters: dog bites man and man bites dog contain the same words but express different relationships. Yet the dot products used by self-attention do not inherently know whether two token representations came from adjacent positions or opposite ends of a sequence. Rotary position embedding, usually shortened to RoPE, adds position information by rotating parts of the query and key vectors before their attention scores are computed. The useful consequence is subtle: each token receives a transformation based on its absolute position, while the dot product between two transformed vectors depends on their relative position.

Artificial Intelligence 06 Sep 2026 10 min read

Reduce Transformer Inference with Early Exits

A transformer classifier normally spends the same number of layers on every input. A straightforward support ticket and an ambiguous one both travel through the entire network, even when an intermediate representation may already contain enough information for the easy case. Early exiting changes that fixed-compute rule. It adds prediction points inside the model and lets sufficiently confident inputs stop before the final layer. Harder inputs continue through more layers. The result is input-dependent computation: the model can reduce average work without forcing every request to use a smaller network.

Artificial Intelligence 06 Sep 2026 11 min read

Reduce Transformer Inference Cost with Early Exiting

A transformer classifier normally spends the same number of layers on every input. A clear support request and an ambiguous one both pass through the entire network, even when an intermediate representation already contains enough information to classify the easy case correctly. Early exiting changes that fixed-compute rule. It attaches prediction heads to intermediate layers and lets an input stop once a chosen exit rule considers the prediction sufficiently reliable. Easy inputs can use less computation, while harder inputs continue through deeper layers.

Artificial Intelligence 06 Sep 2026 8 min read

Reduce Language Model Parameters with Weight Tying

Language models need to turn token IDs into vectors before processing them and turn hidden vectors back into vocabulary scores before predicting the next token. A straightforward design gives those two operations separate parameter matrices. When the vocabulary and hidden dimension are large, each matrix can contain many parameters. Weight tying removes that duplication by reusing one parameter matrix for both roles. The input side reads rows from the matrix as token embeddings; the output side uses the same learned vectors to score candidate tokens, usually through the matrix transpose.

Artificial Intelligence 06 Sep 2026 10 min read

Reduce KV Cache Size with Grouped-Query Attention

Autoregressive language models generate one token at a time. To avoid recomputing attention keys and values for every previous token at every step, inference systems usually keep those tensors in a key-value cache, or KV cache. This saves computation, but the cache grows with sequence length and can become a major memory cost when serving long contexts or many requests at once. One architectural choice has a direct effect on that cost: how many separate key and value heads the attention layer stores. Standard multi-head attention gives every query head its own key and value head. Grouped-query attention (GQA) keeps multiple query heads but lets groups of them share key and value heads.

Artificial Intelligence 06 Sep 2026 11 min read

Length-Normalized Log Probabilities for Comparing Generated Sequences

A language model assigns a probability to each next token, but applications often need to compare complete candidate sequences. A reranker may choose among generated answers. A decoder may keep several partial hypotheses. An evaluator may compare alternative completions under the same prompt. The obvious approach is to multiply each candidate’s token probabilities, or equivalently add their log probabilities. That gives the probability the model assigns to the whole continuation. It also creates an important bias: every additional token contributes a probability no greater than 1, so longer sequences usually accumulate lower raw scores even when their individual tokens are highly plausible.

Artificial Intelligence 06 Sep 2026 9 min read

Inspect Transformer Predictions with the Logit Lens

A transformer language model produces its next-token prediction only after many layers of computation. When that prediction is wrong or surprising, developers often want a more specific question answered: how did the model’s candidate tokens change as the input moved through the network? The logit lens is a simple interpretability technique for exploring that question. Instead of waiting for the final layer, it takes an intermediate representation and passes it through the model’s final decoding machinery to obtain vocabulary logits. Repeating this across layers gives a rough view of how token predictions evolve with depth.

Artificial Intelligence 06 Sep 2026 9 min read

Inspect Language Model Uncertainty with Token Entropy

A language model can produce fluent text even when several continuations look similarly plausible to the model. Looking only at the selected token hides that ambiguity: a token chosen with probability 0.90 and one chosen from a nearly even 0.51 versus 0.49 split both appear as a single output token. Token entropy summarizes how spread out the model’s next-token probability distribution is. It can help developers inspect uncertain generation steps, compare decoding behavior under controlled conditions, and build diagnostic signals for evaluation. But entropy is not a probability that the model is correct, and using it as one leads to unreliable decisions.

Artificial Intelligence 06 Sep 2026 10 min read

Direct Preference Optimization for LLM Alignment

Supervised fine-tuning can teach a language model to imitate good answers, but many alignment problems are easier to express as comparisons: given two responses to the same prompt, which one is better? A preference dataset captures that signal as triples containing a prompt, a preferred response, and a rejected response. The challenge is turning those comparisons into model updates without treating a subjective preference as an ordinary next-token target. Direct Preference Optimization (DPO) provides one practical answer. It trains a policy model to increase its relative preference for chosen responses over rejected responses while measuring that change against a fixed reference model. Unlike a common reinforcement-learning-from-human-feedback pipeline, standard DPO does not require training a separate reward model and then running a reinforcement-learning optimizer.

Artificial Intelligence 06 Sep 2026 12 min read

Cross-Entropy Loss for Classification

A classifier needs more than a way to count correct answers. During training, it needs a signal that says not only whether a prediction was wrong, but also how the model’s scores should change. Suppose the correct class is cat. A model that assigns cat probability 0.49 and another class 0.51 is wrong, but it is close to the decision boundary. A model that assigns cat probability 0.001 is also wrong, and much more confident in that mistake. Treating those predictions as equally bad throws away useful information.

Artificial Intelligence 06 Sep 2026 10 min read

Control LLM Repetition with Token Penalties

Language models sometimes repeat a phrase, return to the same point, or fall into a short loop even when the prompt asks for a concise answer. A common response is to increase randomness, but temperature changes the whole next-token distribution. That can reduce repetition while also making unrelated choices less predictable. Token penalties provide a more targeted control. They adjust the scores of tokens that have already appeared, making some repeated tokens less likely before the decoder chooses the next token. This can be useful for open-ended generation, but it is not a general quality switch: repeated tokens are often exactly what correct text requires.

Artificial Intelligence 06 Sep 2026 11 min read

Contrastive Learning for Text Embeddings

A text embedding model turns text into a vector so that software can compare meaning numerically. The difficult part is not producing vectors. A neural network can produce vectors for almost any input. The difficult part is teaching the geometry of those vectors so that distances correspond to the relationships your application cares about. Contrastive learning provides a practical way to do that. Instead of asking a model to predict a class label, you show it examples that should be close together and examples that should be farther apart. Training adjusts the encoder so that those relationships become easier to recover from the resulting vectors.

Artificial Intelligence 06 Sep 2026 8 min read

Constrain LLM Output with Grammar-Guided Decoding

Asking a language model to return JSON, SQL, or another structured format creates a failure mode that ordinary prompting cannot remove: the model can understand the requested format and still generate a token that makes the output syntactically invalid. For applications that immediately parse model output, one missing quote or delimiter can turn an otherwise useful answer into an error. Retrying helps, but it spends more inference time without guaranteeing that the next attempt will parse.

Artificial Intelligence 06 Sep 2026 11 min read

Compress Embeddings with Scalar Quantization

Embedding systems can become expensive for a reason that has little to do with the embedding model itself: storing and scanning the vectors. A collection of millions of dense vectors can consume gigabytes even before an index adds its own data structures. Moving those vectors through memory can also become part of query latency. Scalar quantization reduces that cost by representing each embedding coordinate with fewer bits. Instead of storing every coordinate as a 32-bit floating-point value, a system might map it to an 8-bit integer and keep enough information to approximately reconstruct or compare the original value.

Artificial Intelligence 06 Sep 2026 10 min read

Compare Model Distributions with KL Divergence

AI systems often produce probability distributions rather than single answers. A classifier assigns probabilities to classes, a language model assigns probabilities to possible next tokens, and a teacher model can provide a soft target distribution for a smaller student. In all of these cases, developers need a way to ask: how different is one probability distribution from another? Kullback-Leibler divergence, usually shortened to KL divergence, is one answer. It measures how much a comparison distribution Q differs from a reference distribution P, with the differences weighted by what P considers important.

Artificial Intelligence 06 Sep 2026 9 min read

Combine Fine-Tuned Models with Weight Averaging

Fine-tuning the same model for different datasets or objectives can leave a team with several useful checkpoints. Serving all of them as an ensemble may improve robustness, but it also multiplies inference work. Choosing only one checkpoint avoids that cost but discards what the others learned. Weight averaging offers a third option: combine compatible checkpoints by averaging their parameters, then serve the result as one model. The arithmetic is simple. The important question is whether the checkpoints occupy a compatible region of parameter space so that interpolation preserves useful behavior rather than destroying it.

Artificial Intelligence 06 Sep 2026 9 min read

Choose Pooling Strategies for Text Embeddings

A transformer usually produces one contextual representation for every input token. Many applications, however, need one vector for an entire sentence, query, or document. Semantic search, clustering, and similarity systems commonly compare these fixed-size vectors rather than every token representation separately. The operation that turns a variable number of token vectors into one vector is called pooling. It can look like a minor implementation detail, but changing it changes the representation being compared. Averaging every meaningful token, selecting a designated token, or emphasizing particular positions encodes different assumptions about where useful information lives.

Artificial Intelligence 06 Sep 2026 10 min read

Accelerate LLM Generation with Speculative Decoding

Autoregressive language models generate text one token at a time. Even when an accelerator has substantial parallel compute available, the model normally cannot determine token 12 until token 11 is known. That dependency makes generation latency difficult to reduce simply by adding more parallel hardware. Speculative decoding attacks this bottleneck by doing cheap work ahead of the expensive model. A faster draft model proposes several future tokens. The full target model then evaluates those proposals together and accepts the portion that is consistent with its own distribution. With the appropriate acceptance-and-correction algorithm, this changes how generation is computed without changing the distribution that the target model defines.

Artificial Intelligence 05 Sep 2026 10 min read

Use Mixup to Regularize Neural Network Classifiers

A neural network can fit its training examples very well while learning decision boundaries that behave poorly between them. Ordinary augmentation helps by creating plausible variations of individual examples, but there is another useful idea: train the model on points that lie between pairs of examples. Mixup does this by interpolating both the inputs and their targets. If one image is labeled cat and another is labeled dog, mixup can create a synthetic input that is partly each image and a target that is partly each class. The model is then trained to produce a correspondingly mixed prediction.

Artificial Intelligence 05 Sep 2026 12 min read

Use Early Exits for Adaptive Neural Network Inference

A conventional neural network uses the same depth for every input. An obvious example and a difficult edge case both pass through all layers before the model returns a prediction. That fixed computation is simple to operate, but it can waste work when intermediate representations are already sufficient for some inputs. Early-exit inference makes computation adaptive. The model attaches prediction heads to intermediate layers. At each head, an exit policy decides whether the current prediction is reliable enough to return or whether the input should continue through deeper layers. Easy inputs can therefore use less computation while difficult inputs retain access to the full network.

Artificial Intelligence 05 Sep 2026 9 min read

Use Class-Weighted Loss for Imbalanced Classification

A classifier trained on imbalanced data can achieve a low average loss while learning the minority class poorly. If 99% of training examples belong to one class, errors on the remaining 1% contribute relatively little to an unweighted objective simply because they occur less often. Class-weighted loss changes that training signal. Instead of treating every example’s loss equally, it gives examples from selected classes more influence on parameter updates. This is useful when class frequency and the importance of learning each class are badly misaligned.