Skip to content

Archive

Language Models

40 articles
Artificial Intelligence 13 Sep 2026 7 min read

Contrast Expert and Amateur Models During Decoding

A language model can assign high probability to a token for two different reasons: the token may fit the prompt particularly well, or it may simply be common under many contexts. Contrastive decoding tries to separate those effects by comparing the next-token distributions of two models. A stronger expert supplies the main distribution, while a weaker amateur supplies a signal for patterns that do not require the expert’s extra capability.

Artificial Intelligence 10 Sep 2026 10 min read

Distill Sequence Models with Teacher-Generated Outputs

Distill Sequence Models with Teacher-Generated Outputs A large text generator may produce useful outputs but still be too expensive for the latency, memory, or throughput budget of a deployment. Training a smaller model on the original dataset is the obvious baseline, but it throws away information encoded in the larger model’s behavior. Sequence-level knowledge distillation offers another option: let a capable teacher generate target sequences, then train a smaller student to reproduce those sequences. The student learns from concrete examples of what the teacher tends to produce rather than only from the original human targets or from the teacher’s next-token probabilities.

Artificial Intelligence 08 Sep 2026 9 min read

The Softmax Bottleneck in Language Models

A language model can have a powerful network behind it and still be constrained by the layer that turns its hidden state into next-token probabilities. In the common linear-softmax output layer, that constraint has a precise form: across many contexts, the model can represent only a limited family of log-probability patterns. This limitation is known as the softmax bottleneck. It is not a claim that softmax itself is defective, nor does it mean every modern language model is visibly harmed by it. It is a structural result about a particular output parameterization.

Artificial Intelligence 08 Sep 2026 13 min read

Rewrite RAG Queries Without Losing User Intent

A retrieval-augmented generation system often searches with the user’s latest message. That works for self-contained questions, but conversational questions frequently depend on earlier turns. Consider a support assistant. The user first asks about a failed database migration, discusses PostgreSQL for several turns, and then asks: Does the rollback command work on version 16 too? Searching that sentence literally may retrieve pages about unrelated rollback commands because the query does not say what is being rolled back. A query rewriter can turn the conversational message into a self-contained retrieval query such as:

Artificial Intelligence 08 Sep 2026 11 min read

Filter Synthetic Training Data with Rejection Sampling

Generating synthetic examples is easy; generating synthetic examples that are worth training on is harder. A language model can produce thousands of candidate answers, but blindly adding them to a training set can reinforce factual errors, weak reasoning, unwanted style, or artifacts of the generator itself. Rejection sampling provides a simple mental model for controlling that pipeline: generate one or more candidates, evaluate each candidate with an acceptance rule, and keep only candidates that pass. The acceptance rule might use deterministic checks, a learned reward model, another language model, human review, or a combination of signals.

Artificial Intelligence 07 Sep 2026 8 min read

Weight Tying in Language Models

A language model needs to solve two related problems with vocabulary-sized parameters. At the input, it must turn each token ID into a vector. At the output, it must turn a hidden state into one score for every possible next token. A straightforward architecture gives these two operations separate parameter matrices. That works, but it can be expensive when the vocabulary and hidden dimension are large. Weight tying is a simple architectural idea: use the same learned matrix for the input token embeddings and the output token projection when their shapes and semantics are compatible.

Artificial Intelligence 06 Sep 2026 8 min read

Reduce Language Model Parameters with Weight Tying

Language models need to turn token IDs into vectors before processing them and turn hidden vectors back into vocabulary scores before predicting the next token. A straightforward design gives those two operations separate parameter matrices. When the vocabulary and hidden dimension are large, each matrix can contain many parameters. Weight tying removes that duplication by reusing one parameter matrix for both roles. The input side reads rows from the matrix as token embeddings; the output side uses the same learned vectors to score candidate tokens, usually through the matrix transpose.

Artificial Intelligence 06 Sep 2026 11 min read

Length-Normalized Log Probabilities for Comparing Generated Sequences

A language model assigns a probability to each next token, but applications often need to compare complete candidate sequences. A reranker may choose among generated answers. A decoder may keep several partial hypotheses. An evaluator may compare alternative completions under the same prompt. The obvious approach is to multiply each candidate’s token probabilities, or equivalently add their log probabilities. That gives the probability the model assigns to the whole continuation. It also creates an important bias: every additional token contributes a probability no greater than 1, so longer sequences usually accumulate lower raw scores even when their individual tokens are highly plausible.

Artificial Intelligence 06 Sep 2026 9 min read

Inspect Transformer Predictions with the Logit Lens

A transformer language model produces its next-token prediction only after many layers of computation. When that prediction is wrong or surprising, developers often want a more specific question answered: how did the model’s candidate tokens change as the input moved through the network? The logit lens is a simple interpretability technique for exploring that question. Instead of waiting for the final layer, it takes an intermediate representation and passes it through the model’s final decoding machinery to obtain vocabulary logits. Repeating this across layers gives a rough view of how token predictions evolve with depth.

Artificial Intelligence 06 Sep 2026 10 min read

Direct Preference Optimization for LLM Alignment

Supervised fine-tuning can teach a language model to imitate good answers, but many alignment problems are easier to express as comparisons: given two responses to the same prompt, which one is better? A preference dataset captures that signal as triples containing a prompt, a preferred response, and a rejected response. The challenge is turning those comparisons into model updates without treating a subjective preference as an ordinary next-token target. Direct Preference Optimization (DPO) provides one practical answer. It trains a policy model to increase its relative preference for chosen responses over rejected responses while measuring that change against a fixed reference model. Unlike a common reinforcement-learning-from-human-feedback pipeline, standard DPO does not require training a separate reward model and then running a reinforcement-learning optimizer.

Artificial Intelligence 05 Sep 2026 9 min read

Tie Input and Output Embeddings in Language Models

A language model needs token representations in two places. At the input, it converts token IDs into vectors. At the output, it converts a hidden vector into one score for every token in the vocabulary. A straightforward design gives these two operations separate parameter matrices, even though both matrices associate vocabulary items with vectors. Weight tying removes that duplication by using the same matrix for both roles. The model looks up input embeddings from the matrix and later uses its transpose to produce output logits. This can remove a large block of parameters, but it also couples two parts of the model that would otherwise learn independently.

Artificial Intelligence 05 Sep 2026 9 min read

Direct Preference Optimization for Language Models

A language model can learn to imitate examples with supervised fine-tuning, but imitation alone does not directly express a common requirement: for the same prompt, one acceptable response may be preferable to another. Preference data represents that requirement as comparisons. A training record contains a prompt, a chosen response, and a rejected response. Direct preference optimization (DPO) uses those pairs to adjust a language model so that the chosen response becomes more favored relative to the rejected one, while comparing the update with a fixed reference model.

Artificial Intelligence 04 Sep 2026 9 min read

Perplexity for Language Model Evaluation

A language model can assign high probability to likely text and low probability to unlikely text, but developers still need a compact way to summarize that behavior across many tokens. Perplexity is one common metric for this job. Perplexity is useful when comparing probabilistic language models on the same evaluation data under compatible tokenization and scoring rules. It is much less useful as a general score for whether generated answers are correct, helpful, safe, or well written.

Artificial Intelligence 04 Sep 2026 9 min read

Interpret Language Model Perplexity Correctly

A language model can improve on its training objective while still leaving an important question unanswered: how well does it predict text it did not train on? Perplexity is a compact way to measure that predictive fit for autoregressive language models, but the number is easy to misuse. A lower perplexity can mean that a model assigns higher probability to held-out text. It does not automatically mean that the model follows instructions better, reasons more reliably, hallucinates less, or produces more useful answers. Comparisons can also become misleading when tokenization, evaluation data, or context handling differs.

Artificial Intelligence 04 Sep 2026 8 min read

Choose Between Causal and Masked Language Modeling

Two language models can use similar transformer components yet learn from text in very different ways. One may predict the next token from everything to its left. Another may hide selected tokens and reconstruct them from surrounding text. That training choice changes what information is available during learning and strongly influences which tasks the resulting model naturally supports. These objectives are called causal language modeling and masked language modeling. Understanding the distinction helps when choosing a pretrained model, interpreting its outputs, or designing a training objective for a language task.

Artificial Intelligence 03 Sep 2026 10 min read

Tokenization in Language Models

Language models do not read text as a sequence of words. Before a model can process a prompt, a tokenizer converts the text into a sequence of token IDs from a fixed vocabulary. The model operates on those IDs, and generated IDs are later converted back into text. This extra layer is easy to ignore because most model APIs accept ordinary strings. But tokenization affects how much text fits in a context window, how usage-based costs are calculated, how text is truncated or split, and why seemingly small formatting changes can alter model behavior.