Skip to content

Category

natural language processing

6,613 papers

#natural language process... Preprint Open access Oct 2026

Hint-Guided Diversified Policy Optimization for LLM Reasoning

Recent developments in Large Language Models (LLMs) have showcased impressive reasoning capabilities, with Reinforcement Learning with Verifiable Rewards (RLVR) being a promising enhancement strategy. However, existing reward mechanisms are constrained to the outcome-level correctness and lack explicit signals to guide...

Zhiyu Cao, Kaixin Wu, Mingjie Zhong et al. · 0 citations
#natural language process... Preprint Open access Oct 2026

Encoded but Not Routed: Explaining the Table-Chart Gap in Scientific Claim Verification

Multimodal LLMs are increasingly used to assist scientific peer review, where a core requirement is verifying whether claims in a paper are supported by its evidence. Prior work has shown that models perform substantially better at this task when the evidence is a table than when it is a chart of the same underlying da...

Sunisth Kumar, Xanh Ho, Tim Schopf et al. · 0 citations
#natural language process... Preprint Open access Oct 2026

Specializing Without Forgetting: Analyzing Knowledge Preservation in Multilingual Model Adaptation

While continual pretraining (CPT) is a practical way to extend large language models to new languages, na\"ive finetuning often erodes existing capabilities through catastrophic forgetting. We investigate which model layers drive this trade-off, and whether interventions at these layers can guide knowledge preservation...

Sanchit Ahuja, Terra Blevins · 0 citations
#natural language process... Preprint Open access Oct 2026

HyperLogic: A Hard, Forward-Authored Chinese Logical Reasoning Benchmark with Execution-Derived Answers

Existing logic benchmarks primarily measure models' ability to answer reasoning questions directly. Scalable benchmarks often generate text from formal structures, which makes answers easy to compute but fixes the formalization before the problem is written. Forward construction preserves the challenge of finding a fai...

Ming Zhang, Qiyuan Peng, Yinxi Wei et al. · 0 citations
#natural language process... Preprint Open access Oct 2026

Rank-Turbulence Delta and Interpretable Approaches to Stylometric Delta Metrics

This article introduces two new measures for authorship attribution - Rank-Turbulence Delta and Jensen-Shannon Delta - which generalise Burrows's classical Delta by applying distance functions designed for probabilistic distributions. We first set out the theoretical basis of the measures, contrasting centred and uncen...

Dmitry Pronin, Evgeny Kazartsev · 0 citations
#natural language process... Preprint Open access Oct 2026

Can We Trust LLMs on Memristors? Diving into Reasoning Ability under Non-Ideality

Memristor-based analog compute-in-memory (CIM) architectures provide a promising substrate for the efficient deployment of Large Language Models (LLMs), owing to superior energy efficiency and computational density. However, these architectures suffer from precision issues caused by intrinsic non-idealities of memristo...

Taiqiang Wu, Yuxin Cheng, Chenchen Ding et al. · 0 citations
#natural language process... Preprint Open access Oct 2026

RT-SFT: Text Style Transfer from Non-Parallel Corpora by Roundtrip Translation

Text style transfer (TST) is naturally a supervised task - rewrite a sentence in a target style while preserving its meaning - yet the parallel corpora that supervision requires exist for only a handful of style domains. A common workaround is to *normalize* an input into a style-agnostic intermediate and then *stylize...

Ruoxi Liu, Philipp Koehn · 0 citations
#natural language process... Preprint Open access Oct 2026

Last Layer Logits to Logic: Empowering LLMs with Logic-Consistent Structured Knowledge Reasoning

Large Language Models (LLMs) achieve excellent performance in natural language reasoning tasks through pre-training on vast unstructured text, enabling them to understand the logic in natural language and generate logic-consistent responses. However, the representational differences between unstructured and structured...

Songze Li, Zhiqiang Liu, Zhaoyan Gong et al. · 0 citations
#natural language process... Preprint Open access Oct 2026

Enrich-on-Graph: Query-Graph Alignment for Complex Reasoning with LLM Enriching

Large Language Models (LLMs) exhibit strong reasoning capabilities in complex tasks. However, they still struggle with hallucinations and factual errors in knowledge-intensive scenarios like knowledge graph question answering (KGQA). We attribute this to the semantic gap between structured knowledge graphs (KGs) and un...

Songze Li, Zhiqiang Liu, Zhengke Gui et al. · 0 citations
#computer vision Preprint Open access Oct 2026

The Percept-V Challenge: Can Multimodal LLMs Crack Simple Perception Problems?

Cognitive science research treats visual perception, the ability to understand and make sense of a visual input, as one of the early developmental signs of intelligence. Its TVPS-4 framework categorizes and tests human perception into seven skills such as visual discrimination, and form constancy. Do Multimodal Large L...

Samrajnee Ghosh, Ashish Goswami, Naman Agarwal et al. · 0 citations
#natural language process... Preprint Open access Oct 2026

BioMol-MQA: A Multi-Modal Question Answering Dataset For LLM Reasoning Over Bio-Molecular Interactions

Retrieval augmented generation (RAG) has shown great power in improving Large Language Models (LLMs). However, most existing RAG-based LLMs are dedicated to retrieving single modality information, mainly text; while for many real-world problems, such as healthcare, information relevant to queries can manifest in variou...

Saptarshi Sengupta, Shuhua Yang, Paul Kwong Yu et al. · 0 citations
#natural language process... Preprint Open access Oct 2026

Automatic register identification for the open web using multilingual deep learning

This article presents multilingual deep learning models for identifying web registers -- text varieties such as news reports and discussion forums -- across 16 languages. We introduce the Multilingual CORE corpora, which contain over 72,000 documents annotated with a hierarchical taxonomy of 25 registers designed to co...

Erik Henriksson, Amanda Myntti, Saara Hellstr\"om et al. · 0 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.