MoNe: Modular Neural Memory for Efficient Long Context Inference
MoNe is a lightweight modular neural memory that attaches to any frozen pretrained Transformer to enable long-context inference without retraining, achieving strong performance on needle-in-a-haystack and word extraction benchmarks from RULER, where ICL degrades sharply.
Won-Yong Cho, Kyubyung Chae, Tribhuvanesh Orekondy et al.
· 0 citations