Skip to content
Conference Open access

What Tokens Truly Matter? The Logit Conflation Problem in LLM Sampling

2026 · Annual Meeting of the Association for Computational Linguistics · pp. 36943-36961 · 0 citations · 32 references
Computer Science

TL;DR

This work identifies the Logit Conflation Problem, where a to-ken’s logit aggregates prompt-independent factors, including linguistic fluency and parametric associations, with prompt-relevance, and proposes SEAL-Sampling to isolate this component through attention-weighted attribution.

Abstract

Sampling methods for large language models select candidate tokens based on logit statistics, implicitly assuming that high log-its indicate desirable outputs. We identify the Logit Conflation Problem , where a to-ken’s logit aggregates prompt-independent factors, including linguistic fluency and parametric associations, with prompt-relevance. However, only prompt-relevance determines instruction-following quality. We propose SEAL-Sampling ( S ignal E xtraction for A ctive Re L evance) to isolate this component through attention-weighted attribution. Our framework defines prompt-relevance as the causal effect of prompt content on token logits and establishes attention patterns as an efficient proxy. Experiments on LLaMA-3 demonstrate significant improvements over top-nσ , with gains of 1.8% on AlpacaEval 2.0 and 2.2% on IFEval. Furthermore, attribution scores correlate weakly with raw logits, confirming the extraction of an orthogonal signal. The method is training-free and introduces minimal latency, adding less than 12ms overhead per token.

Read PDF

Similar papers

Review Jul 2026

ConfidenceBench: Evaluating Confidence Calibration in Large Language Models

ConfidenceBench, a calibration benchmark that evaluates verbalized confidence estimates in 15 frontier LLMs using the Brier score, a proper scoring rule that incentivises truthful probability reporting shows that verbalized confidence calibration is a distinct and practically important axis of LLM reliability, complementary to standard accuracy-based evaluation.

M. ffrench-Constant, Daniel Yang, Xinmeng Huang et al. · 1 citation · ⚡1
Jul 2026

Where Does the Noise Come From? A Variance-Components Decomposition of Non-Determinism in LLM Brand Answers

A crossed random-effects (generalizability-theory) decomposition is specified that partitions the total variance of a response-level brand outcome into these four sources, and embeds the components in a decision-study allocation that returns how many repeats, paraphrases, models, and languages to buy for a target reliability.

D. Żatuchin · 3 citations · ⚡1
Jul 2026

Would You Walk to the Car Wash? Revealing the Salience Bias of Large Language Models in Commonsense Reasoning

The findings relocate the bottleneck of commonsense reasoning failures from model competence to elicitation, and release SaliTrap as a testbed for this blind spot, to show that lightweight, inference-time prompting alone substantially closes the gap without any retraining.

Zheng Wu, Chenhao Xue, Shijie Zheng et al. · 0 citations
Open access Jul 2026

How Does Prompt Anchoring Affect Large Language Model Outputs?

The study identifies prompt anchoring as a source of methodological variation in LLM-assisted content analysis, indicating that anchoring strategies should be explicitly specified, justified, and reported as part of the study methodology.

Eungi Kim · 0 citations
#artificial intelligence Preprint Aug 2026

When Linguistic and Internal Confidence Diverge in Large Language Models

Regression analyses show that distributional properties of confidence scores explain much of the observed alignment pattern, with model metadata playing a smaller role after controls, and support a lossy-channel view of linguistic confidence.

Hefan Zhang, Bing-Quan Zhang, Ming Cheng et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.