SQLStructEval is introduced, a framework that analyzes this behavior through canonical abstract syntax tree representations and adopts a pipeline that first generates structured intermediate representations and then deterministically compiles them into SQL, improving execution accuracy and structural agreement among co...
Yi-Xi Zhou, Fan Zhang, Zhiyu Guo et al.· arXiv.org· 1 citation
Jev has the lowest cost and median response time among the evaluated configurations, while hosted language models achieve higher baseline accuracy and Rankings by baseline accuracy differ from rankings by correctness across every condition and repeat, although small differences in the latter do not establish a general...
Fan Zhang, Yan-Kai Chen, Zhuo-Han Xie et al.· 6 citations· ⚡1
FinMMEval 2026 Task 1 evaluates multilingual financial multiple-choice question answering in English, Chinese, Arabic, and Hindi. The task tests whether systems can select the correct answer to finance questions involving domain terminology, numerical interpretation, and conceptual financial reasoning across languages...
Zhuohan Xie, Yu-Yang Dai, R. Elbadry et al.· arXiv.org· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.