Skip to content

Archive

LLM

76 articles
Artificial Intelligence 02 Sep 2026 5 min read

Control LLM Randomness with Temperature and Top-p

Large language models usually generate text one token at a time. At each step, the model assigns scores to possible next tokens, those scores become probabilities, and a decoding strategy chooses what comes next. Two common controls in that process are temperature and top-p. They are often described as creativity settings, but that description is incomplete. They change how the model samples from its probability distribution, which affects repeatability, diversity, and the chance of selecting lower-probability tokens.

Artificial Intelligence 02 Sep 2026 5 min read

Budget LLM Context Windows Without Losing Critical Instructions

Large-language-model applications rarely fail because a prompt is one token too long. They fail because context growth is handled without priorities. Chat history expands, retrieval returns more passages, tool results become verbose, and eventually the application truncates whichever text happens to be easiest to cut. A safer design treats the context window as a budget with explicit allocations. The goal is not to fill every available token. The goal is to preserve the information that controls behavior while leaving enough room for a complete answer.

Artificial Intelligence 01 Sep 2026 5 min read

Validate LLM Output with Structured Contracts

Large language models are useful when software needs to turn ambiguous text into a structured decision, extraction, or plan. The dangerous shortcut is to treat a model response as if it were already trusted application data. Even when a provider can constrain output to JSON or a schema, the result can still be semantically wrong: a date can be impossible, an identifier can refer to a nonexistent record, or a supposedly positive amount can be negative. Reliable integrations therefore need a contract boundary between model output and the rest of the system.

Artificial Intelligence 01 Sep 2026 5 min read

Evaluating RAG Systems with a Small Golden Dataset

Retrieval-augmented generation (RAG) is easy to demo and surprisingly hard to evaluate. A fluent answer can hide weak retrieval, while a good retriever can be blamed for an answer model that ignores its evidence. A useful evaluation process separates those failure modes. You do not need thousands of examples to begin. A carefully maintained golden dataset of 30 to 100 representative questions can catch many regressions before users do. Define what the system is supposed to do Start with the product contract rather than a model metric. For a documentation assistant, useful requirements might be: