Skip to content

Category

natural language processing

6,613 papers

#natural language process... Preprint Open access Oct 2026

Spend Bytes on Breadth: Precision-Count Trade-offs for Decode-Time KV Compression in Long Chain-of-Thought Reasoning

Reasoning models write most of their KV cache while decoding long chains of thought (CoT), so the cache has to be compressed online under a fixed memory budget. Decode-time methods mostly decide which tokens to evict. We ask how a fixed byte budget should be split between the number of cached tokens and their precision...

Runguo Li · 0 citations

Automatic Speech Recognition for Low-Resource Sinhala: A Critical Review of Methods, Challenges, and Future Directions

Automatic speech recognition (ASR) for low-resource languages remains a major challenge. Sinhala, the primary language of Sri Lanka with about 16 million speakers, illustrates the difficulty: agglutinative morphology, a 54-phoneme inventory, subject-object-verb (SOV) syntax and scarce annotated speech data limit both c...

Chanuka Dinuwan, Sanath Jayasena, Buddhika Karunarathne · 0 citations
#computer vision Preprint Open access Oct 2026

Atomic Visual Entailment: Enhancing Zero-Shot Vision-Language Reasoning through Atomic Fact Decomposition and Learned Selection

Visual entailment (VE) asks whether an image supports, contradicts, or leaves undecided a textual hypothesis. Strong results come from fine-tuning large vision-language models on labelled data, while zero-shot and hybrid approaches remain far behind. A VE hypothesis often bundles several visual claims, yet existing zer...

Nallathambi Vethiappan, Derya Soydaner, Gijs Wijnholds · 0 citations
#natural language process... Preprint Oct 2026

More Than Words: Compositional Tokenization for Efficient Language Models

Language models process and generate text sequentially in token units, and the tokenizer determines how much text each inference step covers. Under standard tokenization, a short English phrase such as"On the table."is usually produced as four separate predictions for the preposition (On), article (the), noun (table),...

Yuval Reif, Guy Kaplan, Roy Schwartz · 0 citations
#natural language process... Preprint Open access Oct 2026

What Is a Repeated Token Worth? The Scaling Geometry of Multi-Epoch Pretraining

As pretraining increasingly repeats data, every run faces three questions: how many epochs to take, how that number should change with model size, and whether anything besides the epoch count matters. We answer them by pricing a repeated token against two references: one epoch on the same data, which gives its value, a...

Yekun Chai, Haoyi Xiong · 0 citations
#natural language process... Preprint Open access Oct 2026

Dataset Signatures in Human-LLM Interactions and User Modeling

Human--LLM interaction datasets shape our understanding of AI use and provide a foundation for downstream research, including training and evaluation of user models. In recent years, a growing number of datasets have sought to capture a representative picture of human--LLM interactions. But how different are the pictur...

Joseph Suh, Serina Chang · 0 citations
#natural language process... Preprint Open access Oct 2026

Writing as a Self-Organized Critical Process

We explain autocorrelation decay power laws omnipresent in texts by self-organized criticality. Specifically, we analyze the recently released KLiCKe keystroke dataset and show that not only the final texts' autocorrelations form a manifold that adheres to a power law with a finite-size scaling, but also the text revis...

Nikolay Mikhaylovskiy · 0 citations
#natural language process... Preprint Oct 2026

The Hidden States Cookbook: A Large-Scale Ablation Study for Noise-Robust Conversational Intent Classification in Industry

Conversational database interfaces face a critical challenge: users naturally embed queries in conversational noise (greetings, politeness, off-topic remarks), which degrades intent classification accuracy and wastes computational resources. Despite advances in orchestration and retrieval strategies, a fundamental ques...

Bogdan Bogachov, Nikita Letov, Yaoyao Fiona Zhao · 0 citations
#natural language process... Preprint Oct 2026

Harness-Search: Guiding Long-Horizon Search through Multi-Agent Coordination

Long-horizon search requires agents to gather evidence across multiple steps and synthesize it into well-supported answers. The recent agent harnesses provide a natural and promising framework to support such long-running search processes. As interaction histories grow, one single agent in harnesses might get stuck and...

Shan-Yong Wang, Zhen-Wen Ji, Lei Jin et al. · 0 citations
#natural language process... Preprint Oct 2026

Towards Unbiased On-Policy Distillation for Block Diffusion Language Models

On-policy distillation (OPD) has emerged as an effective post-training paradigm for language models, with recent efforts extending it to block diffusion language models (BDLMs). However, existing studies focus almost exclusively on small block sizes, leaving distillation into student models with larger blocks underexpl...

Zai-Quan Yang, Fei Wei, Yong Wang et al. · 0 citations
#natural language process... Preprint Open access Oct 2026

The \`{I}r\`{o}y\`{i}nSpeech Text Corpus: 24,905 Curated Yor\`ub\'a Sentences for Speech and Language Technology

\`{I}r\`{o}y\`{i}nSpeech is a 42-hour, 80-speaker Yor\`ub\'a read-speech corpus whose audio has been distributed by ELRA since 2024. This paper describes the release of its text component: 24,905 unique, hand-verified, tone-marked Yor\`ub\'a sentences (275,897 tokens; 15,687 types), curated in 2022 as recording prompts...

Kola Tubosun, Aanuoluwapo Aremu, Tolulope Ogunremi et al. · 0 citations
#natural language process... Preprint Open access Oct 2026

MemStrata: 95% and 90.91% Source-Aware Accuracy on LongMemEval-500 and LoCoMo-1540 with a Local Qwen 3.8 27B Q4_K_M Reader

An adequate conversational answer may differ from a short or incomplete benchmark reference. To measure adequacy against the recorded history we prefer source-aware grading, in which the judge checks the reference against the full source before assessing system-blinded answers; original reference-only grading is report...

Neeraj Yadav (Called It Inc.) · 0 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.