Skip to content

Calibrated e-CUSUM Decoding for Quantized Reasoning Models: Why Token Log-Probability Is the Wrong Observable for Decoding Monitors

Jul 2026 · arXiv.org · Vol abs/2607.11317 · 0 citations · 27 references
Computer Science Mathematics

TL;DR

A training-free decoding controller that combines a degeneration-aware alarm score fusing token uncertainty with explicit verbatim repetition and a calibrated e-process-inspired sequential detector, and finds that non-termination, rather than looping, is the dominant failure mode on GSM8K.

Abstract

Low-bit quantization makes small reasoning models inexpensive to deploy but can degrade their chains of thought. This motivates decoder-side monitors that intervene when generation becomes unreliable. We show that a natural candidate, the centered token log-probability increment $\log p(w_t)+H_t$, is the wrong observable for this purpose. Under the model's own sampling law it is a mean-zero martingale by construction, so it measures sampling self-consistency rather than trajectory health and is nearly silent during confident repetition, where both $\log p(w_t)$ and entropy are close to zero. We introduce a training-free decoding controller that combines (i) a degeneration-aware alarm score fusing token uncertainty with explicit verbatim repetition and (ii) a calibrated e-process-inspired sequential detector. The raw product process is Ville-valid under a conditional-mean null, while the deployed CUSUM-floored statistic is treated as an empirical change detector because the score is history-dependent and autocorrelated. On GSM8K with DeepSeek-R1-Distill-Qwen-1.5B in FP16 and INT4, calibration turns a monitor that fires on 93--95% of generations into a selective detector of failing traces ($\phi \approx 0.3$, precision $\approx 0.6$ against a 0.38 base rate). In this pilot, the controller reduces measured verbatim-degeneration signals and yields a positive but statistically inconclusive INT4 accuracy change from 63% to 69% (paired McNemar $p=0.18$, $n=100$), at a 28% token-budget cost. We also find that non-termination, rather than looping, is the dominant failure mode on GSM8K. The main contribution is methodological: an explanation of why centered token log-probability is inadequate for decoder monitoring and a calibrated, cautiously evaluated replacement.

View source

Similar papers

#artificial intelligence Preprint Aug 2026

HSRM: Hidden-State Reward Models for Test-Time Verification

HSRM is introduced, a lightweight hidden-state reward model that verifies candidate solutions by directly reading the generator's internal representations rather than re-processing its text, providing an efficient alternative to text-only verification by reusing representations already computed during generation.

Xianzhi Li, Xiao-Dan Zhu · 0 citations
Preprint Aug 2026

Attend to Your Own Thoughts: Breaking the Barrier for Post-Training Quantization of Reasoning LLMs through the Lens of 1.58-Bit Quantization

We propose ScaleQ-1.58, a scalable ternary post-training quantization (PTQ) framework for reasoning LLMs. Its core insight stems from an empirical finding: although modern LLMs are typically trained to exhibit chain-of-thought reasoning capabilities, in the PTQ regime, even the latest CAT-Q method based on learning-bas...

Shigeng Wang, Chao Li, Yangyuxuan Kang et al. · 3 citations
Preprint Aug 2026

Which Decisions Low-Bit Quantization Breaks, and How to Predict Them

This work tracks quantization across 16 models from 8 families under round-to-nearest, seven under AWQ, two under GPTQ and one under GGUF, at 8 down to 2 bits, and measures the margin, the picked option's score minus its best alternative's, which removes the protection a large margin affords.

Zekun Wu, Swati Dhiman, A. Koshiyama · 1 citation
Jul 2026

Flat Score, Amplified Failures: How the Error Budget Masks Damage in Quantized LLM Agents

This claim for multi-turn, tool-calling agents, where it now matters most, is tested for post-training quantization to 4-bit weights and diagnostics, the per-channel error rate and success under a shrinking budget come from logs benchmarks already collect.

Jiwon Jang, Kisu Yang, Heuiseok Lim et al. · 1 citation
Preprint Aug 2026

Attention-Path Fragility as an Uncertainty Signal in Large Language Models

It is proposed that a model's uncertainty about a token is reflected not only in the breadth of its output distribution but also in whether a confident prediction is \emph{fragile} under perturbation of its attention pathways, a training-free estimator that masks attention heads and measures the BALD mutual information...

Minsoo Kim, Sungyoung Ji, Kisung Moon et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.