SQLStructEval is introduced, a framework that analyzes this behavior through canonical abstract syntax tree representations and adopts a pipeline that first generates structured intermediate representations and then deterministically compiles them into SQL, improving execution accuracy and structural agreement among co...
Yi-Xi Zhou, Fan Zhang, Zhiyu Guo et al.· arXiv.org· 1 citation
Claims paired with AI-generated images are a rapidly growing form of misinformation. Existing automated fact-checking (AFC) methods mainly treat this as a provenance problem, detecting low-level synthesis artifacts to decide whether an image is AI-generated. However, such methods do not verify what human fact-checkers...
Rui-Hong Zeng, Jonathan Tonglet, Preslav Nakov et al.· 0 citations
Jev has the lowest cost and median response time among the evaluated configurations, while hosted language models achieve higher baseline accuracy and Rankings by baseline accuracy differ from rankings by correctness across every condition and repeat, although small differences in the latter do not establish a general...
Fan Zhang, Yan-Kai Chen, Zhuo-Han Xie et al.· 6 citations· ⚡1
Co-FactChecker is proposed, a framework for human-AI collaborative claim verification that translates expert feedback into trace-edits that introduce targeted modifications to the trace, sidestepping the shortcomings of dialogue-based interaction.
It is suggested that robust meme understanding requires multimodal training, varies substantially across semantic categories, and remains sensitive to over-specialization when models are fine-tuned on individual datasets rather than trained in a unified setting.
Ali Ezzat Shahroor, Mohamed Bayan Kmainasi, A. Hasnat et al.· arXiv.org· 5 citations
This paper introduces a novel two-step multi-class classification system to identify varying degrees of machine involvement in Bulgarian text. As Large Language Models (LLMs) proliferate, distinguishing original human writing from machine-assisted or machine-generated content is crucial to prevent misinformation and pr...
Boyan Bogdanov, D. Georgiev, D. Dimitrov et al.· Computational Linguistics in...· 0 citations
This work introduces FinCUABuildBench, a benchmark for evaluating financial CUA task construction, and introduces FinCUABuildAgent, a multi-agent system for automatically constructing dynamic financial CUA evaluation tasks.
Jing-Pu Yang, Feng-Xian Ji, Jinri Guo et al.· 0 citations
Neuron steering reinforces dialect in some varieties when the prompt is already dialectal but cannot induce it from MSA prompts, whereas vector steering succeeds in both settings, and Arabic dialects are therefore steerable mainly through distributed rather than localized representations.
K. Elozeiri, Mervat T. Abassy, Omar Kallas et al.· arXiv.org· 0 citations
Jais 2 is a family of Arabic-centric large language models developed jointly by MBZUAI, Cerebras, and Inception, designed to advance Arabic-centric language modeling, with strong performance across the Arabic and culturally grounded benchmarks evaluated in this report. The family includes, to our knowledge, the largest...
Mohamed Anwar, A. Freihat, George Ibrahim et al.· 3 citations
A semantic correctness taxonomy is introduced that assigns open-ended answers to eight ordered classes, separating verbose-but-correct answers from those contaminated by hallucinated content and CAP (Context-Aware Precision), a reference-based metric that scores question-conditioned statements using bidirectional NLI.
Elitsa Yotkova, Violeta Kastreva, Petar Velkov et al.· 0 citations
This work introduces EDRAC, the first large-scale benchmark for dialectal Arabic machine reading comprehension (MRC) and generative QA, covering five major dialects: Egyptian, Moroccan, Emirati, Syrian, and Saudi Arabic, and benchmarks Arabic-centric and multilingual LLMs on EDRAC using lexical and semantic metrics.
Noor Abo Mokh, K. Chirkunov, Teresa Lynn et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.