Book
Open access
Jun 2025
TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models
This paper proposes a taxonomy to systematically measure both shallow table understanding abilities and deep table reasoning abilities, and designs an evaluation framework to robustly measure table reasoning capabilities with three distinct inference modes.
Ce Li, Xiaofan Liu, Zhiyan Song et al.
· Annual International ACM SIG... · 3 citations