Skip to content

Category

natural language processing

6,429 papers

#machine learning Preprint Open access Oct 2026

Continuous Semantic Caching for Low-Cost LLM Serving

As Large Language Models (LLMs) become increasingly popular, caching responses so that they can be reused by users with semantically similar queries has become a vital strategy for reducing inference costs and latency. Existing caching frameworks have proposed to decide which query responses to cache by assuming a fini...

Baran Atalar, Xutong Liu, Jinhang Zuo et al. · 0 citations
#machine learning Preprint Open access Oct 2026

Just on Time: Token-Level Early Stopping for Diffusion Language Models

Diffusion language models generate text through iterative refinement, a process that is often computationally inefficient because many tokens reach stability long before the final denoising step. We introduce a training-free, token-level early stopping approach that identifies convergence independently at each position...

Zakhar Kohut, Severyn Shykula, Mykola Vysotskyi et al. · 0 citations
#machine learning Preprint Open access Oct 2026

Rephrase Before You Act: Characterizing and Mitigating Language Sensitivity in Vision-Language-Action Models

Vision-language-action models (VLAs) are strikingly sensitive to instruction phrasing and do not inherit the language robustness of the vision-language models they are built on. A one-word edit can move success by tens of points: $\pi_{0.5}$ turns on a LIBERO stove 100% of the time for "switch on the stove" and 2% for...

Mikey Watts (Independent Researcher), Yuchen Cui (University of California, Los Angeles) · 0 citations
#machine learning Preprint Open access Oct 2026

Constitution-Guided Watermarking

Watermarking enables language model providers to identify text generated by their models. However, its desired properties can conflict (\ie~stronger watermark signals can degrade text quality), while designs that resist editing may also facilitate forgery. Providers address these trade-offs by choosing configurations t...

Toluwani Aremu, Samuele Poppi, Nils Lukas · 0 citations
#machine learning Preprint Open access Oct 2026

Certified by Abstention: Distribution-Free Guarantees for Chain-of-Thought Verifiers at Small Calibration Budgets

Signals that predict whether a chain-of-thought (CoT) trace is correct are compared by AUC, but deploying one requires a threshold with a guarantee. We ask what distribution-free selective guarantees deliver for CoT verifiers at realistic calibration budgets of tens to a few hundred labelled problems, using seven open...

Arjun Balaji · 0 citations
#machine learning Preprint Open access Oct 2026

The Persona Hierarchy Model: Understanding Contextual Generalization in Fine-Tuning LLMs

Language models are routinely fine-tuned under a fixed context, such as a generic system prompt, persona or domain-specific instruction, yet the learned behavior sometimes stays confined to that context and sometimes broadly generalizes to unseen contexts. We propose the Persona Hierarchy Model to explain this: a share...

Jiachen Zhao, Zhengxuan Wu, David Bau et al. · 0 citations
#machine learning Preprint Open access Oct 2026

Quantize by Drift: Label-Free Mixed-Precision Post-Training Quantization for Text Embedders

Mixed-precision post-training quantization needs a per-module sensitivity signal; for a text embedder the obvious one -- the retrieval quality a module costs when quantized -- needs relevance labels that deployments rarely have. We measure a label-free substitute: quantization-induced representation drift, obtained by...

Hyojung Han, Jongmin Kim, Seung-Hun Jeon · 0 citations
#machine learning Preprint Open access Oct 2026

Same Text, Different Prediction: Serving-Context Nondeterminism in Text Classifiers

Deterministic inference is essential for reliable and trustworthy machine learning. Prior studies of text generation have shown that changing factors such as batch size, batch composition, hardware, or inference engine can alter the generated text, even when the prompt, model parameters, and sampling randomness are fix...

Santhosh Kumar Kasa, Siva Rajesh Kasa, Sumit Negi · 0 citations
#machine learning Preprint Open access Oct 2026

CARE: Certifying Acceleration for Vision-Language-Action Inference

While vision-language-action (VLA) models have advanced rapidly, running them at every control step remains expensive. Prior work accelerates VLA inference using techniques like action chunking and visual-token pruning, typically evaluating based on latency and average task success. However, acceleration may discard in...

Rui Liu, Tong Zheng, Jindong Gu et al. · 0 citations
#machine learning Preprint Open access Oct 2026

CoDR: Training-Free Confidence-Drift Remasking for Diffusion Language Models

Masked diffusion language models (MDLMs) decode by repeatedly committing tokens to masked positions, but these commitments are usually irreversible. A token chosen under sparse, partial context is kept fixed, even when later context no longer supports it. Existing samplers mainly decide when to commit a token, but rare...

Yue Wu, Qinghe Zhang, Yu Zhang et al. · 0 citations
#machine learning Preprint Open access Oct 2026

Which Rollout Taught It That? BehaviorTrace and the Limits of Training-Data Attribution in Online RL

When reinforcement learning teaches a language model a new behavior, can we find the training rollouts that taught it? And when an attribution method says it can, how do we know the answer is real? We study both questions on online RL fine-tuning with GRPO, using a planted behavior with a known cause. We release Behavi...

Amit Nautiyal · 0 citations
#machine learning Preprint Oct 2026

Input-Blind Controls Produce Substantial Oracle Headroom for Layer Programs in Multiple-Choice Evaluation

Adaptive computation aims to improve language-model inference by tailoring execution to each input. For layer programs, oracle evaluations use known answers to estimate the potential gain from this flexibility, before a practical selector is available. However, a gain from selection does not by itself explain why the c...

Yi-Bei Guo, Rui Liu · 0 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.