Large language models (LLMs) have been gradually applied to tasks such as specialized question answering, educational assistance, clinical decision support, and knowledge integration in stomatology. However, general-purpose LLMs still suffer from critical limitations, including outdated knowledge, hallucinatory generat...
This work introduces FORTE, a training-free framework that addresses the challenge of adaptive relevance scoring and global keyframe optimization through two stages: adaptive relevance scoring and global keyframe optimization.
Haifeng Huang, Bi-Yin Xu, Chun-Sheng Xin et al.· 0 citations
SSP first abstracts the underlying question structure into a skeleton, retrieves a small set of same-skeleton questions from a lightweight question pool, then prompts the VLM to analyze their shared negation pattern and synthesize a single-sentence answering strategy.
Yu-Liang Cai, Mohammad Rostami, Jesse Thomason· 0 citations
Keyword-matching benchmarks can credit small models for tool use they never perform. We document such a false positive in a matched-architecture pair of Spanish security language models and propose a ladder of strict, cheap diagnostics. A 661.6M parameter model (approx. 65% code/technical text; no dedicated SFT) and a...
Automated ideation systems are often evaluated on the novelty of the ideas they produce, and that judgment is increasingly delegated to large language models. Such judges are typically built ad hoc and validated, if at all, on human-authored papers rather than on the generated ideas they are meant to score. So, how do...
Noy Sternlicht, Simra Shahid, P. Jansen et al.· 0 citations
Yor\`ub\'a is a widely spoken tonal language that depends on diacritics to avoid lexical ambiguity. However, it is often written without these diacritics, thereby hindering downstream Natural Language Processing (NLP) tasks. In this paper, we introduce Yo-ByT5, a byte-level Automatic Diacritic Restoration (ADR) model f...
Ahmad Samuel Gali, Shamsuddeen Hassan Muhammad University of Lagos, Bayero Unuversity Kano et al.· 0 citations
Large language models are often asked which input factors influenced their outputs. For structured inputs, such reports can be checked by counterfactual perturbation, but each factor must be queried multiple times to estimate its effect, so verification is usually budget-limited. We study how this limited-budget settin...
Tao-Lin Zhang, Han-Yu Wang, Jiu-Heng Wan et al.· 0 citations
LLM arenas turn pairwise human preferences into model rankings. Those preferences may reflect how an answer is presented as well as what it says. We take a stylometric approach to 137,293 decisive French-language votes from the July 2026 Compar:IA release; the primary formatting analysis includes 137,113 battles across...
Simonas Zilinskas, Maayeesha Farzana, Christophe Benavent· 0 citations
Large language models (LLMs) are widely used for claim verification, yet remain brittle for numerical reasoning: even small changes in value can sharply degrade accuracy. We show that this brittleness persists in frontier LLMs, but can be mitigated through adversarial fine-tuning on numerically perturbed examples. Usin...
Bengali secondary education lacks curriculum-grounded benchmarks for STEM question-solving, and general-purpose language models struggle with the precise terminology, unit conventions, and derivations that physics problems demand. We introduce PhysicsMate, a benchmark of 1834 question-answer pairs built from the Nation...
Language models learn content and style jointly, making stylistic variation in their outputs difficult to identify and control. We study whether recurring styles in model responses can be discovered without supervision and explicitly controlled. We design an algorithm that learns to separate representations of content...
Ioana Marinescu, E. Oermann, Kyunghyun Cho· 0 citations