Skip to content

Where Should a Document Live: Context, Representations, or Parameters?

Sep 2026 · 0 citations · 43 references
Computer Science

TL;DR

A controlled comparison of representation-based (KV-cache based) and parametric (fine-tuning-based) adaptation methods on five knowledge-intensive benchmarks shows that in the oracle setting, Cartridges are the most accurate injection method at nearly every storage budget, outperforming parametric methods by 10 points.

Abstract

To answer questions outside of their pre-training data, large language models (LLMs) need access to new information, which can be presented in the context window as documents, encoded into the model's parameters, or injected as latent representations. However, each of these methods comes with different efficiency, cost, and performance trade-offs, with no single winner. We present a controlled comparison of representation-based (KV-cache based) and parametric (fine-tuning-based) adaptation methods on five knowledge-intensive benchmarks. We show that in the oracle setting, Cartridges (KV) are the most accurate injection method at nearly every storage budget, outperforming parametric methods by 10 points. Compaction (KV) matches Cartridges only at low compression rates, lagging behind the parametric methods by 10 points at rates higher than $50\times$. In the more realistic multi-document retrieval scenario, Cartridges are the only method that matches in-context learning (ICL), leading the parametric methods by 29 points and Compaction by 15 points. Nonetheless, Cartridges are also the only method, besides full fine-tuning and large MLP adapters, that suffers from catastrophic forgetting, i.e., a 6% performance degradation on control benchmarks, with 13% in coding.

View source

Similar papers

#machine learning Preprint Sep 2026

Cartridges++: KV Cache Compression without Off-Context Derailment

Serving long documents to a Large Language Model (LLM) repeatedly is expensive: computations grow with context length, and the memory footprint of the key-value (KV) cache balloons. Compressed KV (CKV) representations aim to mimic the cache of a document and are typically computed once and for all, ahead of inference t...

Sonia Laguna, João Monteiro, Marco Cuturi et al. · 0 citations
#artificial intelligence Preprint Sep 2026

An Empirical Study of VLM Pipelines for Long-Document QA

Vision-Language Models (VLMs) are increasingly used for long-document processing, where the inputs combine text with charts, tables, figures, and complex layouts. Deploying them means choosing how to feed the document to the model, which retriever to use when only a subset of pages is sent, and whether to run the model...

K. E. Ak, Jay Mohta, Gwang Lee et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Infinite-Parameter LLMs: Generating and Adapting Weights from Live Data

Scaling laws hold that language models grow more capable with more parameters and more training data. Mixture-of-Experts (MoE) architectures are a remarkable demonstration of these laws, activating only a fraction of an enormous parameter bank for each token. But this success is built on static pretraining data --- the...

Jin-Lin Hu, Ross M. Clarke, Yi-Chuan Zhang et al. · 0 citations
Open access Sep 2026

Large language models as digital libraries: a multi-benchmark and multi-model study

Querying LLMs as digital libraries is feasible, but its effectiveness depends on model strength, deployment conditions, dataset structure, and execution strategy, and Galois remains valuable when relational discipline and controlled query execution are required.

Mirco Cazzaro, G. Silvello · 0 citations
#natural language process... Preprint Sep 2026

Mixture-of-Experts Language Models Can Be Strong and Efficient Retrievers

Recent work has shown that fine-tuning decoder-only large language models (LLMs) for retrieval yields strong first-stage retrievers, with effectiveness improving as backbones grow in size. However, every query and document must pass through the full model, so encoding cost increases with model size. Mixture-of-Experts...

Anubhav Shrestha, Safal Shrestha, Minwu Kim et al. · 0 citations
Book Open access Aug 2026

NumCache: KV Cache Compression and Retrieval for Financial Document QA

The proposed NumCache, which compresses SEC filings into KV caches initialized from numerically dense regions and trained directly on financial QAs, is evaluated, which highlights cache-based retrieval with number-preserving representations as an effective approach for long-context financial QA.

Eftychia Makri, Peiwen Li, Yi-Dong Jiang et al. · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.