Skip to content

Category

natural language processing

6,613 papers

#artificial intelligence Preprint Oct 2026

FSPO: Policy-Consistent Risk and Pareto-Feasible Control for Budgeted LLM RL Post-Training

Adaptive LLM reinforcement-learning post-training changes multiple training actuators online, including rollout temperature, group size, clipping, KL regularization, verifier allocation, and update budget. Three coupled issues remain unresolved. A future-risk model trained from behavior trajectories need not estimate t...

Miao-Bo Hu, Shu-Hao Hu, Xiao-Bo Guo et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

ROUTEAUDIT: Interaction-Aware Identification for Budgeted Multi-Verifier Routing

Adaptive multi-verifier systems are commonly compared through endpoint quality-cost gaps, even when the verifier catalog, availability, accounting, information filtration, or scorer changes with the policy. We formulate verifier routing as a contract-conditioned identification problem. The contract records request supp...

Miaobo Hu, Shuhao Hu, Xiaobo Guo et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Large language models exhibit unreliable updating of clinical judgment as patient evidence evolves

Large language models (LLMs) are increasingly explored for clinical reasoning, but whether they appropriately revise judgments as patient evidence evolves remains unclear. We evaluated longitudinal belief updating using matched intensive-care trajectories from electronic health records. Across diverse LLMs, conditionin...

Min Zeng, Rui Zhang · 0 citations
#artificial intelligence Preprint Oct 2026

VERSE: Verified Self-Evolving Optimizer for Agent Harnesses

Harness evolution improves an LLM agent's prompts, tools, and workflow, while the optimizer's own tools and procedures often remain fixed. We study whether an optimizer can improve another agent more effectively by also improving how it diagnoses failures, develops edits, and tests their effects. Two observations guide...

Ze-Kai Wang, Ying-Qiang Ge, Ze-Kun Wang et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Right Order, Wrong Scale: Auditing LLM Judges for Occupational AI Measurement

LLM judges are increasingly used to assess whether AI outputs meet workplace requirements, but agreement on response rankings does not establish agreement on acceptance rates or occupational aggregates. We introduce O*NET-BENCH, an audit suite derived from an existing survey of 45,796 worker ratings, and evaluate 33 pr...

Harry Lyu, Neil Thompson · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Fast Models, Slow Evidence: A Paired and Self-Audited Evaluation of System-1 Decision Models for LLM Agent Harnesses

Agent harnesses make many small, typed decisions per task: which model to call, which tool to use, whether retrieved text is relevant, whether an input carries an injection. System-1 decision models answer such questions in a single forward pass with class probabilities, promising large cost and latency savings over LL...

Jiawei Li · 0 citations
#natural language process... Open access Oct 2026

Knowledge boundary probing and demand-guided intervention for LLM-based power system code generation

Large language models (LLMs) can turn grid-analysis requests into executable programs for power-system simulation, but utilities and research laboratories often require on-premise deployment. In this setting, first-pass failures frequently arise at an API-knowledge boundary , through hallucinated functions, misused par...

hui wu, Xiaoyang Wang, Zhong Fan · 0 citations
#natural language process... Preprint Open access Oct 2026

A Situational Speech Synthesizer for Yoruba: System Design, Phonological Rule Architecture, and Orthographic Extensions for Contour

We present TTSYoruba, a rule-based concatenative diphone speech synthesizer for Yoruba, deployed at online as part of the YorubaName.com open dictionary of Yoruba personal names. The system takes tone-marked Yoruba text as input and produces audio output by applying a hand-crafted phonological rule system to a recorded...

Kola Tubosun, Adedayo Oluokun, Hafiz Adewuyi et al. · 0 citations
#natural language process... Preprint Aug 2026

Tangut Word Segmentation under Extreme Resource Scarcity: Integrating Traditional Lexicons and Unlabeled Text

Tangut is an extinct language whose script does not explicitly mark word boundaries. We present the first systematic study of Tangut word segmentation using 2,750 expert-annotated segments (31,893 tokens), traditional lexicons, and unlabeled text. Our framework combines a reliability-calibrated lexicon-lattice represen...

Lifan Deng, Yongwei Zhang, Sen Sun et al. · 0 citations
#natural language process... Preprint Open access Oct 2026

ReNikud: Audio-Supervised Hebrew Grapheme-to-Phoneme Conversion

Grapheme-to-phoneme (G2P) conversion for Modern Hebrew is needed for applications like text-to-speech (TTS), but is challenging due to the language's abjad writing system, which leaves vowels largely unwritten, creating substantial ambiguity. Standard approaches first predict vowel diacritics (nikud) to produce Interna...

Maxim Melichov, Yakov Kolani, Morris Alper · 0 citations
#natural language process... Preprint Open access Oct 2026

The Piggyback Hypothesis of Generalization: Explaining and Mitigating Emergent Misalignment

The mechanisms behind LLMs' broad over-generalization beyond training examples remain unclear. Emergent misalignment (EM) offers a striking case study: finetuning on narrow tasks induces broad misalignment to semantically-unrelated test domains. In this work, we propose the Piggyback Hypothesis: the chat-template token...

Jiachen Zhao, Zhengxuan Wu, Aryaman Arora et al. · 0 citations
#natural language process... Preprint Open access Oct 2026

DECK: A Consistency x Confidence Taxonomy of LLM Hallucinations

Existing hallucination taxonomies classify LLM errors by what is wrong with the output -- memorised misconceptions, reasoning failures, fluent fabrications -- but cannot answer a different question: which uncertainty scorer would have caught this error? We propose a complementary taxonomy that classifies errors by thei...

Mohit Singh Chauhan · 0 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.