Skip to content

Archive

Artificial Intelligence

391 articles
Artificial Intelligence 05 Sep 2026 9 min read

Tie Input and Output Embeddings in Language Models

A language model needs token representations in two places. At the input, it converts token IDs into vectors. At the output, it converts a hidden vector into one score for every token in the vocabulary. A straightforward design gives these two operations separate parameter matrices, even though both matrices associate vocabulary items with vectors. Weight tying removes that duplication by using the same matrix for both roles. The model looks up input embeddings from the matrix and later uses its transpose to produce output logits. This can remove a large block of parameters, but it also couples two parts of the model that would otherwise learn independently.

Artificial Intelligence 05 Sep 2026 10 min read

Stop Neural Network Training with Early Stopping

Neural network training does not automatically become more useful because it runs for more epochs. Training loss can continue falling while performance on unseen data stops improving or begins to deteriorate. That creates a practical question: when should training stop? A fixed epoch count is easy to configure, but it cannot know whether a particular run converged early or still needs more optimization. Early stopping answers this by monitoring performance on validation data during training. When the chosen validation metric stops improving for long enough, training ends. Used carefully, it can reduce wasted computation and limit unnecessary overfitting while preserving the checkpoint that performed best on validation data.

Artificial Intelligence 05 Sep 2026 12 min read

Split Conformal Classification for Prediction Sets

A classifier usually returns one label or a vector of scores. That is convenient when the application must choose one answer, but it hides an important distinction: some inputs strongly support one class, while others leave several classes plausible. Conformal prediction provides a way to expose that ambiguity. For classification, it can return a prediction set containing one or more labels instead of forcing every input into a single choice. With an appropriate calibration procedure and statistical assumptions, the method can target a long-run coverage level such as 90%: roughly speaking, the true label should appear in the prediction set for at least that proportion of future examples.

Artificial Intelligence 05 Sep 2026 10 min read

Reduce Transformer Padding with Length Bucketing

Transformer training often starts with a simple batching rule: shuffle the examples, take the next B sequences, and pad every sequence in the batch to the length of the longest one. The rule is correct, but it can waste substantial computation when sequence lengths vary widely. A batch containing a 900-token document and several 100-token documents must usually represent every sequence with 900 token positions. Attention masks prevent padding from acting like real input, but they do not necessarily make the padded positions free to process.

Artificial Intelligence 05 Sep 2026 9 min read

Reduce LLM Decoding Latency with Speculative Decoding

Large language models generate text autoregressively: each new token depends on the tokens that came before it. That dependency makes ordinary decoding sequential. Even when a GPU has enough compute to process many token positions in parallel, the model normally discovers only one new token per decoding step. Speculative decoding tries to turn some of that sequential work into parallel verification. A faster draft model proposes several future tokens. The larger target model then scores those proposed positions together and decides which proposals can be accepted. When the draft predicts well, one expensive target-model pass can advance generation by multiple tokens.

Artificial Intelligence 05 Sep 2026 9 min read

Prevent Target Leakage in Machine Learning Evaluation

A machine learning model can score extremely well in offline evaluation and fail as soon as it reaches production. Sometimes the model is not the main problem. The evaluation accidentally gave it information that would not exist when a real prediction is made. This failure is called target leakage: information related to the outcome enters the model inputs in a way that makes the target easier to predict than it will be at inference time. Leakage can produce impressive metrics because the model is solving an easier, unrealistic problem.

Artificial Intelligence 05 Sep 2026 8 min read

Pooling Token Embeddings into Sequence Representations

Transformer encoders usually produce one vector for every input token. Many downstream tasks, however, need one vector for the whole input: a classifier may need a single representation of a support ticket, and a retrieval system may need one vector for an entire passage. The step that converts a variable number of token vectors into one fixed-size vector is pooling. It looks simple, but the choice of pooling rule changes what information survives, how padding must be handled, and whether the resulting vector matches the way a model was trained.

Artificial Intelligence 05 Sep 2026 8 min read

Pool Token Embeddings into Text Representations

A Transformer usually produces one contextual vector for every input token. Many downstream tasks, however, need one vector for the whole text. Semantic search may need one vector per document, clustering needs one vector per item, and similarity scoring often expects two fixed-size vectors to compare. Pooling is the step that turns a variable number of token vectors into one fixed-size representation. The operation looks simple, but small implementation choices can change the resulting geometry. Averaging padding tokens, assuming the first token is meaningful for every model, or changing pooling at deployment time can make an otherwise correct embedding pipeline behave poorly.

Artificial Intelligence 05 Sep 2026 10 min read

Pack Training Sequences to Reduce Padding Waste

Language-model training often processes sequences in fixed-size tensors. When examples have very different lengths, padding makes those tensors easy to batch but can leave many token positions doing little useful work. A batch that physically contains 8,000 positions may contain far fewer than 8,000 real training tokens. Sequence packing reduces this waste by placing multiple shorter examples into the same fixed-length training sequence. The idea is simple; the semantics are not. If packing accidentally lets one example attend to another, predicts across boundaries that should be independent, or assigns incorrect position IDs, the training objective changes rather than merely becoming more efficient.

Artificial Intelligence 05 Sep 2026 9 min read

Normalize Embeddings Before Dot-Product Similarity

Embedding systems often compare vectors with cosine similarity or a dot product. The formulas look similar enough that it is easy to treat the two metrics as interchangeable. They are not interchangeable for arbitrary vectors. A dot product depends on both the angle between two vectors and their magnitudes. Cosine similarity removes magnitude and compares direction only. That difference can change nearest-neighbor rankings, retrieval results, and similarity thresholds. This article builds a practical mental model for deciding whether to normalize embeddings. You will see why L2 normalization makes dot product equivalent to cosine similarity, how inconsistent normalization breaks comparisons, and when preserving vector magnitude may be intentional.

Artificial Intelligence 05 Sep 2026 10 min read

Model Ensembling for Combining Predictions

A machine learning model can fail because of patterns specific to its training run: its initialization, sampled batches, training data, architecture, or hyperparameters. Training another model may produce different mistakes. Ensembling uses that disagreement by combining predictions from multiple models instead of trusting one model alone. The idea is simple, but useful ensembles require more than averaging everything available. Models that make nearly identical errors provide little complementary information, while diverse models can improve predictions at the cost of additional training, memory, and inference work.

Artificial Intelligence 05 Sep 2026 10 min read

Measure Attention Concentration with Entropy

Transformer attention is often inspected as a matrix of weights. That works for a few examples, but it becomes difficult when you need to compare many heads, layers, tokens, or model runs. A useful summary is attention entropy: a number that describes how concentrated or spread out one attention distribution is. Entropy can answer a narrow but practical question: does this query place most of its attention mass on a few available positions, or distribute that mass broadly? It does not tell you whether the model is correct, whether a token caused the prediction, or whether a head is important. Used with those limits in mind, it is a compact diagnostic for attention behavior.

Artificial Intelligence 05 Sep 2026 9 min read

Mask Padding Tokens in Transformer Attention

Transformer batches often contain sequences with different lengths. To store them in one rectangular tensor, shorter sequences are usually extended with padding tokens. Padding solves a shape problem, but it creates a modeling problem: those extra positions are not part of the original input. If attention treats padding like ordinary content, real tokens can assign probability to positions that carry no useful information. The result may be wasted attention, representations that depend on how much padding was added, and training behavior that differs unnecessarily across batches.

Artificial Intelligence 05 Sep 2026 10 min read

Handle Label Noise in Supervised Learning

Supervised learning assumes that training examples come with useful target labels. Real datasets rarely satisfy that assumption perfectly. A support ticket may be assigned to the wrong queue, an image may receive the wrong class, or two annotators may interpret an ambiguous policy differently. These errors create label noise: the recorded target does not reliably represent the target the model is supposed to learn. Enough noise can teach a model contradictory patterns, distort evaluation, and make apparently difficult modeling problems into data-quality problems.

Artificial Intelligence 05 Sep 2026 11 min read

Evaluate Neural Networks with Exponential Moving Average Weights

Neural network parameters do not move smoothly toward a final solution. Stochastic optimization updates them using noisy minibatch gradients, so the weights used after one training step can differ slightly from those used after the next. Saving only the final step therefore makes one particular point on that training path responsible for evaluation and deployment. An exponential moving average, or EMA, keeps a second copy of the parameters that changes more gradually. Instead of replacing this copy with every new set of training weights, each update blends the previous average with the current parameters.

Artificial Intelligence 05 Sep 2026 12 min read

Estimate Neural Network Uncertainty with Monte Carlo Dropout

A neural network can produce a confident-looking prediction even when the input is unlike the data it learned from. A single output such as 0.93 tells you what one forward pass predicts; by itself, it does not tell you how sensitive that prediction is to uncertainty in the learned model. Monte Carlo dropout is a practical way to obtain an additional uncertainty signal from some neural networks that were trained with dropout. Instead of disabling dropout at inference time, you keep it active, run the same input through the network multiple times, and inspect how much the predictions vary.

Artificial Intelligence 05 Sep 2026 9 min read

Diversify RAG Context with Maximum Marginal Relevance

A retrieval-augmented generation (RAG) system can retrieve highly relevant passages and still build poor context. The problem appears when several top results say almost the same thing. Sending all of them to the language model consumes context without adding much evidence, while a slightly lower-ranked passage containing a different useful fact may be excluded. Maximum marginal relevance (MMR) is a selection strategy for this situation. Instead of choosing passages only by their relevance to the query, MMR repeatedly chooses a passage that is both relevant and sufficiently different from passages already selected.

Artificial Intelligence 05 Sep 2026 9 min read

Direct Preference Optimization for Language Models

A language model can learn to imitate examples with supervised fine-tuning, but imitation alone does not directly express a common requirement: for the same prompt, one acceptable response may be preferable to another. Preference data represents that requirement as comparisons. A training record contains a prompt, a chosen response, and a rejected response. Direct preference optimization (DPO) uses those pairs to adjust a language model so that the chosen response becomes more favored relative to the rejected one, while comparing the update with a fixed reference model.

Artificial Intelligence 05 Sep 2026 12 min read

Contrastive Decoding with Expert and Amateur Models

A language model can assign high probability to text that is fluent but unhelpfully generic, repetitive, or too close to an easy pattern. Changing temperature or top-p changes how tokens are sampled from one model’s distribution, but it does not ask a different question: which candidate tokens are especially characteristic of a stronger model rather than a weaker one? Contrastive decoding asks exactly that. It uses two language models at inference time: a stronger expert and a weaker amateur. A candidate is favored when the expert scores it well relative to the amateur, while a plausibility constraint prevents the decoder from choosing bizarre tokens merely because the amateur dislikes them even more.

Artificial Intelligence 05 Sep 2026 10 min read

Combine Retrieval Rankings with Reciprocal Rank Fusion

A retrieval-augmented generation (RAG) system often has more than one useful way to find evidence. Keyword retrieval is good at exact names, identifiers, and rare terms. Embedding retrieval can find passages that express the same idea with different wording. Using both can improve candidate coverage, but it creates a practical problem: their scores usually do not mean the same thing. A keyword score of 12.4 and a cosine similarity of 0.81 cannot be safely averaged just because both are numbers. Their scales, distributions, and even direction conventions depend on the retrieval methods and implementations.

Artificial Intelligence 05 Sep 2026 9 min read

Classifier-Free Guidance in Diffusion Models

A conditional diffusion model may understand a prompt and still produce samples that only weakly reflect it. During generation, developers therefore often want a way to push the denoising trajectory toward the condition without training a separate classifier for every prompt or label. Classifier-free guidance (CFG) is a widely used way to do that. At each denoising step, the model is evaluated with the condition and without it. The difference between those predictions gives a direction associated with the condition, and a guidance scale controls how strongly sampling moves along that direction.

Artificial Intelligence 05 Sep 2026 10 min read

Attention Sinks for Stable Streaming LLM Inference

Autoregressive language models normally reuse the keys and values of earlier tokens while generating the next token. This KV cache avoids recomputing the entire prefix at every decoding step, but its memory use grows with the cached sequence. A long-running chat, agent, or stream can therefore accumulate more cached state than a serving system wants to keep. A tempting fix is a sliding window: retain only the most recent tokens and evict everything older. For models trained with ordinary dense attention, however, abruptly dropping all early tokens can damage generation quality even when those old tokens do not appear semantically important.

Artificial Intelligence 04 Sep 2026 8 min read

Weight Decay in Neural Network Training

A neural network can keep reducing its training loss while learning parameter values that generalize poorly. Weight decay is one way to regularize training: it applies a small pressure that shrinks selected parameters as optimization proceeds. The idea sounds similar to adding an L2 penalty to the loss, and for plain stochastic gradient descent the two can be made equivalent by matching their scaling. With adaptive optimizers such as Adam, however, adding an L2 penalty to the gradient and directly decaying the weights are not generally the same operation. That distinction is why optimizers such as AdamW use decoupled weight decay.

Artificial Intelligence 04 Sep 2026 9 min read

Use Padding Masks for Variable-Length Transformer Batches

Transformer inputs rarely have identical lengths. One sentence may contain 8 tokens while another contains 30, yet efficient training and inference usually process multiple sequences in rectangular tensors. The usual solution is to add padding tokens to shorter sequences until their shapes match. Padding solves the shape problem but creates a semantic one: the added positions are not real input. If the model treats them like ordinary tokens, they can influence attention, pooling, and training loss. A padding mask tells the computation which positions are valid and which exist only to make the batch rectangular.