Diversify RAG Retrieval with Maximal Marginal Relevance
A retrieval-augmented generation (RAG) system can retrieve highly relevant passages and still build a poor context. The problem appears when the top results repeat the same fact in slightly different wording. Suppose five retrieved chunks all explain how to reset an API token, while a lower-ranked chunk explains the permission change that must happen afterward. Filling the context window with the five near-duplicates gives the language model less useful evidence than selecting a smaller set that covers both parts of the task.