When a question has valid answers under different normative frameworks, a language model must decide which framework to use and whether it can answer correctly within it. We call this setting normative pluralism and study it in Islamic finance using a four-choice taxonomy that separates framework selection from within-...
R. Elbadry, Ahmed Heakl, Saeed Almheiri et al.· 0 citations
MIRAGE, a training-free, model-agnostic defense for long-form RAG, is introduced, a training-free, model-agnostic defense for long-form RAG that consistently restores factuality under mixed and fully polluted evidence and outperforms prior robust-RAG methods.
Saadeldine Eletter, Ruihong Zeng, Yuxia Wang et al.· arXiv.org· 0 citations
TradeLens is introduced, a trace-grounded diagnostic toolkit for evaluating agentic trading systems from their trading records, runtime traces, and deployment configurations, which reframe the evaluation of LLM-based trading agents from capability-centric performance ranking to trace-grounded diagnosis of intelligence-...
Qiqi Duan, Changlun Li, Chen Wang et al.· arXiv.org· 0 citations
Think with Structured Grounding (TwSG), a novel fine-grained image perception framework designed to internalize complex images's tool-use capabilities within the model, is proposed, endowing models with native fine-grained region description and flexible reasoning capabilities.
Chang-Jiang Jiang, Qian-Nian Zhao, Lei Xin et al.· 2 citations
SurakshaEval is introduced, a novel safety benchmark composed of human-written prompts spanning real-world scenarios, explicitly designed for ten major Indian languages - Assamese, Bengali, Gujarati, Hindi, Kannada, Malayalam, Marathi, Punjabi, Tamil, and Telugu - along with English.
Debopriyo Banerjee, K. R. Kavitha, Angana Borah et al.· 0 citations
FinMMEval 2026 Task 1 evaluates multilingual financial multiple-choice question answering in English, Chinese, Arabic, and Hindi. The task tests whether systems can select the correct answer to finance questions involving domain terminology, numerical interpretation, and conceptual financial reasoning across languages...
Zhuohan Xie, Yu-Yang Dai, R. Elbadry et al.· arXiv.org· 1 citation
Large Language Models (LLMs) achieve strong results on many medical benchmarks, but their clinical reasoning remains difficult to evaluate reliably. A central risk is an evaluation illusion: fluent and well-structured explanations can appear clinically convincing even when the final diagnosis is incorrect. We introduce...
Abin Roy, Afthab Salam Kanniyan, Jawadh Abdul Kabeer et al.· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.