Applying the DFA-WK method to evaluate long-context retention in large language models
The growth of context windows in Large Language Models (LLMs), now exceeding millions of tokens, has made evaluating long-range dependency retention a central challenge. Traditional evaluation approaches, such as pointwise retrieval tests (needle-in-a-haystack) and average perplexity, capture only part of this dynamic....