Skip to content

Category

natural language processing

6,613 papers

#artificial intelligence Preprint Open access Oct 2026

Thought-Like-Pro: Enhancing Reasoning of Large Language Models through Self-Bootstrapped Prolog-based Chain-of-Thought

Large language models have demonstrated remarkable capabilities as general-purpose assistants, excelling in a wide range of reasoning tasks and supporting various aspects of daily web usage. This achievement represents a significant step toward achieving artificial general intelligence. Despite these advancements, the...

Xihe Qiu (Shanghai University of Engineering Science, National University of Singapore), Yongxin Deng (Shanghai University of Engineering Science et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Predicting Alignment Generalization with Value Representations

LLM developers post-train their models to exhibit prosocial values and behavioral traits, which are enumerated in an alignment target. However, while recent post-training developments have yielded models that score highly on alignment evaluations, training models on sets of narrow behaviors still influences their behav...

Andy Liu, Mehar Bhatia, Karolina Stanczak et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

All Verdicts are Not Equal: Rethinking LLM Judge Reliability

LLM-as-a-Judge is the standard paradigm for NLP evaluation, yet its systemic reliability remains poorly understood despite being widely treated as a deterministic ground truth. We present a comprehensive reliability audit, stresstesting six frontier models across four benchmarks, five prompt formats, two presentation o...

Vineet Kumar, Darshita Rathore, Anindya Moitra · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Examining Social Attribution in LLM Reasoning: A Theory-Guided Probing Methodology

Large language models (LLMs) are increasingly deployed in sociotechnical systems where social attribution, the reasoning process attributing external events to the causes and reasons of agents' social behaviors, plays a critical role. These processes involve judgments of social cause, responsibility, and blame/credit t...

Zhaoxin Yu, Qingchao Kong, Dajun Zeng et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Agentic-TTT: Training test-time policy for test-time training

Test-time training (TTT) adapts an LLM's parameters using signals derived from test inputs, and can make striking improvements in pre-specified settings such as IMO competitions or designated open problems. By turning deployment experience into parameter updates, TTT provides a direct mechanism for model-level self-imp...

Jiahao Lu, Mohan Kankanhalli · 0 citations
#artificial intelligence Preprint Open access Oct 2026

DataVista: Diagnosing Multimodal LLMs on Data Video Understanding

Data video is a media form that integrates data visualization with video narrative, widely adopted in news reporting and business analysis. Compared with general video understanding, data video understanding places greater emphasis on accurately reading data from animated charts, integrating evidence across charts and...

Yupeng Xie, Zhenyang Wang, Jiayi Zhu et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Forms of LLM-Integrated Applications from LLM-Chats to Autonomous AI Agent System

Large language models (LLMs) are increasingly embedded as components in software systems, marketed under labels such as chatbot, copilot, retrieval-augmented generation, workflow, coding agent and AI agent. Whether these labels denote genuine architectural forms or serve as branding has not been assessed systematically...

Irene Weber (University of Applied Sciences Kempten, Germany) · 0 citations
#artificial intelligence Preprint Open access Oct 2026

GRPODropout: Less is More for Online Reinforcement Learning Rollouts

Reinforcement learning (RL) methods such as GRPO substantially improve large language model reasoning but often suffer from policy entropy collapse: the loss of sampling diversity weakens exploration and limits further improvement. Existing methods address this issue either through algorithm-level interventions, such a...

Hexuan Deng, Zihao Yan, Xuebo Liu et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Same Outcome, Different Evidence: Intent Recovery in LLM Safety Evaluation

Safety evaluations of large language models commonly summarize harmful-output behavior with attack success rate (ASR). Yet the same non-harmful outcome can arise for very different reasons. A model may recover a harmful task and refuse it, fail to recover the task, or respond to something else entirely. Distinguishing...

Haitong Jiang, Chunlin Liu, Sihan Tang et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

DIAL-OPD: Learning More from Fewer Tokens in On-Policy Distillation

On-policy distillation (OPD) supervises student-generated trajectories with token-level teacher signals. Its sampled-token variant avoids the cost of full-vocabulary probabilities. Yet we find that training on fewer tokens can outperform full-token OPD, challenging the intuition that more supervision improves learning....

Anhao Zhao, Haoran Xin, Junlong Tong et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Large Language Model Turnover Undermines Screening for Artificial Intelligence-Assisted Scientific Writing

Journals and conferences have begun to screen submitted manuscripts for text written using large language models (LLMs). The reliability of this screening rests on benchmark evaluations against a fixed set of LLM versions, while the versions in actual use keep changing. Here we quantify how this LLM turnover affects th...

Kazuki Nakajima, Takayuki Mizuno · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Adapting English Quality Classifiers for Multilingual LLM Pretraining Data Selection

Recent advances in large language model (LLM) pretraining highlight the role of high-quality training data in improving performance. While model-based filtering has proven effective in selecting high-quality subsets from web-scale corpora, especially for high-resource languages, low-resource languages face challenges d...

Vinko Sabol\v{c}ec, Bettina Messmer, Yassine Turki et al. · 0 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.