ReRead: Retrieval-Guided Efficient Reading for Long-Context Question Answering
Abstract
Long-context question-answering tasks require models to locate a small amount of scattered evidence within lengthy, sparse, and noisy inputs, and to integrate information across multiple fragments during reading. Existing recurrent reading and memory-updating methods provide an effective paradigm for handling ultra-long contexts by processing the context chunk by chunk while continuously maintaining an intermediate memory. However, these methods typically adopt a complete sequential reading strategy, in which all chunks are treated equally and fed into the memory-updating process. As the context length increases, this strategy leads to rapidly growing large language model (LLM) invocations, token consumption, and inference time, while also increasing the risk that irrelevant information contaminates the memory state. To better balance answer quality and inference cost, we propose ReRead, a retrieval-guided selective reading method for long-context question-answering. ReRead first divides the long context into ordered reading units and performs question-aware hybrid retrieval using BM25 and a dense retriever, then merges and deduplicates the retrieved candidates, and reorganizes them into a selective reading sequence according to their original order in the context. Finally, the model progressively updates its natural-language memory along this sequence and generates the final answer based on the resulting memory state. Experiments on RULER-HQA and various long-context generalization tasks demonstrate that ReRead maintains answer quality comparable to complete sequential recurrent reading, while effectively reducing LLM invocations, token consumption, and end-to-end inference time, achieving a superior efficiency-performance trade-off.