It is established that LLM outputs should be treated as draws from a distribution rather than as fixed measurements, as well as guidance for data editors and authors.
Abstract
Large language models (LLMs) are increasingly used to generate data for research. Typical use cases are classifications, annotations, information extraction, and generation of numerical scores. Unlike conventional measurements, LLM outputs can vary across repeated requests even when the prompt and apparent model settings remain unchanged. This variation arises from deliberate sampling, silent model updates, numerical rounding, or expert routing. Setting a dedicated temperature parameter to zero removes deliberate sampling when that option is available, but it does not eliminate the other sources of randomness. Exact reproduction is therefore generally not possible when using proprietary application programming interfaces. Local execution of open-weight models offers greater control, but reproducibility still depends on the complete hardware and software stack. We illustrate these issues through sentiment classifications of corporate filings and examine their consequences for downstream regression results. We then propose a reporting standard for articles and replication packages, as well as guidance for data editors and authors. Together, these findings and recommendations establish that LLM outputs should be treated as draws from a distribution rather than as fixed measurements.
Large language models (LLMs) are increasingly used in settings where textual summaries, numerical observations, and external tool outputs may provide conflicting evidence. We study how LLMs arbitrate between such sources when they support opposing decisions. To do so, we introduce a controlled synthetic benchmark in which latent risk trajectories generate both numerical time series and natural language summaries, allowing us to construct conflicts where exactly one evidence source is aligned with the ground-truth label. This design lets us independently manipulate modality, temporal recency, source reliability, and evidence provenance. Across open-weight instruction-tuned models, we find that arbitration behaviour is systematic rather than random: models exhibit distinct text-versus-number preferences, follow temporal recency more consistently than explicit reliability cues, and can over-rely on external forecasts even when they conflict with direct contextual evidence. These results suggest that current LLMs often rely on heuristic arbitration strategies when integrating heterogeneous evidence, highlighting a failure mode for tool-augmented decision systems.
Mattia Carletti, Edward Phillips, Fredrik K. Gustafsson et al.· 0 citations
A crossed random-effects (generalizability-theory) decomposition is specified that partitions the total variance of a response-level brand outcome into these four sources, and embeds the components in a decision-study allocation that returns how many repeats, paraphrases, models, and languages to buy for a target reliability.
The results show that the LLM's capability dissolves with dimension in a way no noise-corrupted classical learner mimics - which explains why LLMs, so capable elsewhere, keep losing to fifty-year-old baselines on tables, while leaving the mechanism of the prediction as an open question.
As large language models (LLMs) become increasingly integrated into analytical workflows, an urgent question arises: Can these models replace the trained statistician? This paper presents a controlled experiment to directly test the statistical reasoning ability of LLMs. We employ Monte Carlo simulations to generate datasets with known ground‐truth parameters and pose four canonical statistical questions to five commercially prominent models across three linguistically distinct prompt formulations and five sampling temperature settings, yielding 3000 observations in a full‐factorial design. In order to find an answer to our question, we design prompts that reflect how decision‐makers with varying degrees of statistical knowledge would query AI in the absence of a trained statistician. We find that prompts that portray higher statistical competency can result in higher accuracy for some (but not all) LLMs; we also find that LLMs can fail catastrophically on tasks requiring quantitative precision. We connect these findings to architectural differences among models and to recent literature on epistemic mirroring in LLMs and argue that the observed patterns reveal models are performing linguistic pattern matching on statistically flavoured text rather than genuine statistical reasoning. We conclude that current LLMs cannot replace the statistician, though certain architectures approach useful performance on pattern‐recognition subtasks.
Wolfgang Jank, Bernhard Klingenberg, Sonal Prabhune et al.· International Statistical Re...· 1 citation
Paired equivalence testing at a declared margin is supply: paired equivalence testing at a declared margin, with certification tables giving the items an evaluation needs, computed from disagreement observed under compression, not from independent-binomial variance.
For applications that require per-persona outputs, the same model that cannot sample from a distribution can describe it accurately in a single call, and is proposed Prompt-Perturbed Argyle (PPA), which reduces the same error by 21% at no added cost.
Chaemi Jang, Dongman Lee, Jihee Kim· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.