MUDDLE: Measuring Understanding of Documents under Distractor and Length Effects
In the complete markdown sweep, hard negatives lower accuracy more than length-matched random documents at both context sizes for gpt-5-mini, while random documents stay near the no-distractor baseline, and for gpt-5-mini hard negatives significantly underperform length-matched random distractors when pooled across context sizes.