How Robust Are Automated Fact-Checking Systems? A Cross-Benchmark Evaluation
On ClimateCheck claim-only and fine-tuned models outperform zero-shot LLM and top-performing AVeriTeC 2025 systems, highlighting that noisy evidence can degrade veracity prediction and replacing retrieved evidence with gold annotations improves veracity accuracy, confirming retrieval remains primary bottleneck.