A model-based evaluation framework that combines multidimensional item response theory model with question contexts to predict LLM performance on unseen questions is studied, suggesting that context-aware psychometric modeling is a promising direction for efficient and interpretable LLM evaluation and highlighting cross-scenario generalization as a central open challenge.
Abstract
Evaluation of large language models (LLMs) increasingly requires predicting how a model will perform on new questions or tasks before collecting large amounts of new annotations. This problem is challenging because question difficulty, scenario, and underlying capability demands can vary substantially. Simple retrospective averages may confound model ability with item characteristics. In this paper, we study a model-based evaluation framework that combines multidimensional item response theory model with question contexts to predict LLM performance on unseen questions. The framework represents LLMs through latent capability profiles while using question content to inform item characteristics, allowing information to transfer beyond previously observed items. Empirically, we find that for within-scenario evaluation, incorporating question embeddings improves prediction relative to model-free baselines, and that multidimensional latent structure provides a richer description of capability variation than unidimensional alternatives. At the same time, our results reveal an important limitation that the generalizability does not necessarily translate into reliable prediction under cross-scenario shift. These findings suggest that context-aware psychometric modeling is a promising direction for efficient and interpretable LLM evaluation, while also highlighting cross-scenario generalization as a central open challenge.
While large language models (LLMs) are having a transformative impact on human society, evaluating them remains challenging. Standard benchmarks usually rely on single point estimates that obscure response stochasticity and variability in question difficulty. Here, we introduce a hierarchical Bayesian Beta-Binomial framework for uncertainty-aware evaluation of LLMs on multiple-choice datasets. Our approach moves beyond single accuracy metrics by modeling correct responses binomially and decomposing performance variation into intra-question stochasticity (response variability for a given question) and inter-question heterogeneity (variation in difficulty across questions) using separate priors. The framework yields a probabilistic assessment, providing full posterior distributions and credible intervals for mean accuracy, inter-question heterogeneity, and mean intra-question response variability, enabling rigorous uncertainty quantification. We demonstrate its utility by evaluating multiple LLMs across diverse benchmarks, including under semantic perturbations like question rephrasing. This analysis reveals nuanced model robustness insights and uncovers distinct behaviors across model classes (e.g., reasoning vs. non-reasoning) often missed by traditional descriptive statistics. This methodology offers a statistically grounded and powerful Bayesian lens for analyzing LLM performance, providing deeper insights into accuracy, consistency, and heterogeneity essential for reliable model benchmarking.
Giannis Manousaridis, John G. Samuelsson, B. Emir et al.· Frontiers in Applied Mathema...· 0 citations
The results demonstrate that demographic prompting is not a monolithic intervention: its utility is highly context-dependent, shaped by attribute signal quality, task characteristics, and model architecture.
M. Kamruzzaman, Shrabony Das, Gene Louis Kim· arXiv.org· 0 citations
Multimodal Large Language Models (MLLMs) have achieved strong performance on structured visual understanding tasks such as chart and document question answering. However, existing benchmarks typically evaluate these domains in isolation, leaving underexplored a key capability: whether models can use textual context to determine how chart evidence should be selected, interpreted, and aggregated. We introduce DocHop, a benchmark for integrated chart--context reasoning in document-style images. In DocHop, the document narrative specifies multi-step compositional constraints, while charts provide the corresponding data values. Questions are grounded on a semantic reference label defined in the narrative, requiring models to resolve target entities from context before aggregating evidence across multiple charts. To enable systematic evaluation, we construct DocHop via a stochastic logic-first generation pipeline with controllable reasoning depth and visual density, covering 2,074 examples across six task categories. Experiments on a wide range of proprietary and open-source MLLMs show a substantial gap to human performance: annotators achieve over 90% accuracy, while the best model reaches only 62.83%. Reasoning-enhanced models consistently show improved results, but performance degrades as reasoning complexity increases. Overall, DocHop provides a controlled testbed for challenging multi-hop document reasoning.
Zhuoran Yu, Le Thien Phuc Nguyen, Jaden Park et al.· 0 citations
Large Language Models (LLMs) have demonstrated widespread utility in various tasks, including reasoning over social questions. While LLMs are primarily designed for open-domain applications, recent efforts have adapted them for closed-domain tasks in fields like finance and healthcare to enhance relevance and precision. However, closed-domain LLMs face challenges, including data scarcity, high costs, and limited capacity in addressing intersectional domains that require knowledge across multiple areas. To address these limitations, this study introduces TopicTune, a topic-based prompt-tuning framework that improves the reasoning abilities of open-domain LLMs on social questions by aligning their responses with the specific topics of each query. TopicTune employs a reinforcement learning (RL) model to refine the selection of the top k relevant topics for each query, ensuring the responses align with the query's context. The framework then tailors prompts based on topic-specific nuances, enabling LLMs to generate well-reasoned and contextually relevant responses. We evaluate TopicTune using real-world, open-ended questions from Reddit and Lemmy, demonstrating an improvement in reasoning over social questions across three LLMs. The motivation of this study is to enhance LLM reasoning performance by enabling AI to generate more contextually relevant responses that better meet human needs.
Maryam Amirizaniani, Baktash Ansari, Chirag Shah et al.· International Conference on...· 0 citations
Multidimensional item response theory relies on calibrated item parameters, such as discrimination and category threshold values, which are usually estimated from large samples of human test responses. This study investigates whether the directional loadings of these parameters can be recovered directly from item text using pre-trained sentence embeddings, avoiding the need for initial item calibration. Using the open-source IPIP Big-Five dataset ($n=19{,}719$; 50 items), we built a multidimensional computerized adaptive testing (CAT) simulation using D-optimal item selection. We compared three item loading sources: fitted graded response model parameters, semantic text embeddings, and a lexical baseline. In simulation, semantic embeddings recovered latent trait profiles almost as accurately as fitted parameters (correlation $0.825$ vs. $0.857$), performing noticeably better than simple word overlap ($0.752$). However, the embedding-based model produced inflated posterior variance, showing nearly four times higher measurement uncertainty despite accurate point estimates. We attribute this to collinearity across dimensions, as embedding-derived loadings pointed in similar directions across traits (condition number $137$ vs. $1.0$; mean trait cosine $0.90$). This outcome reflects the shared vocabulary common in personality items. We propose a simple diagnostic metric based on the loading matrix condition number to evaluate whether an item bank is suitable for text-derived loadings prior to testing.
Cross-scale heterogeneous MLLM fusion is recast as selective language-side reasoning transfer within a narrow, low-interference regime, rather than broad capability inheritance, to recast cross-scale capability transfer within a narrow, low-interference regime.
Yinghao Hou, Jiahe Fan, Yuanhao Pu et al.· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.