Skip to content

Category

natural language processing

6,613 papers

#artificial intelligence Preprint Open access Oct 2026

TextReg: Mitigating Prompt Distributional Overfitting via Regularized Text-Space Optimization

Large language models (LLMs) are highly sensitive to the prompts used to specify task objectives and behavioral constraints. Many recent prompt optimization methods iteratively rewrite prompts using LLM-generated feedback, but the resulting prompts often become longer, accumulate narrow sample-specific rules, and gener...

Lucheng Fu, Ye Yu, Yiyang Wang et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Less Back-and-Forth: A Comparative Study of Structured Prompting

Large language models (LLMs) are widely used for open-ended tasks, but underspecified prompts can lead to low-quality answers and additional interaction. This paper studies whether structured prompt design improves response quality while reducing user effort. We compare three prompt conditions: a raw prompt, a checklis...

Saurav Ghosh · 0 citations
#artificial intelligence Preprint Open access Oct 2026

An Assessment of Human vs. Model Uncertainty in Soft-Label Learning and Calibration

Central to human-aligned AI is understanding the benefits of human-elicited labels over synthetic alternatives. While human soft-labels improve calibration by capturing uncertainty, prior studies conflate these benefits with the implicit correction of mislabeled data (mode shifts), obscuring true effects of soft-labels...

Maja Pavlovic, Silviu Paun, Massimo Poesio · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Text Knows What, Tables Know When: Clinical Timeline Reconstruction via Retrieval-Augmented Multimodal Alignment

Clinical language models increasingly operate over electronic health records (EHRs), yet patient records are not stored as temporally grounded trajectories. Clinical notes describe symptoms, assessments, and disease progression, but often compress or narratively reorder events. Structured EHR rows provide timestamps fo...

Sayantan Kumar, Shahriar Noroozizadeh, Juyong Kim et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

When Can Digital Personas Reliably Approximate Human Survey Findings?

Digital personas powered by Large Language Models (LLMs) are increasingly proposed as substitutes for human survey respondents, yet it remains unclear when they can reliably approximate human survey findings. We answer this question using the LISS panel, constructing personas from respondents' background variables and...

Mumin Jia, Yilin Chen, Divya Sharma et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

NCO: A Versatile Plug-in for Handling Negative Constraints in Decoding

Controlling Large Language Models (LLMs) to prevent the generation of undesirable content, such as profanity and personally identifiable information (PII), has become increasingly critical. While earlier approaches relied on post-processing or resampling, recent research has shifted towards constrained decoding methods...

Hyundong Jin, Yo-Sub Han · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Is Escalation Worth It? On the Depth of LLM Cascades

LLM cascades, in which a cheap model defers to an expensive one on low-confidence queries, are widely used to reduce inference cost. Given a pool of models, a practitioner must decide how many models to include and where to set each deferral threshold. We derive first-order optimality conditions showing that, at an opt...

Dylan Bouchard · 0 citations
#artificial intelligence Preprint Open access Oct 2026

PSI-Bench: Interpretable and Clinically Meaningful Evaluation of Depression Patient Simulators

Patient simulators are gaining traction in mental health training by providing scalable exposure to complex and sensitive patient interactions. Simulating depressed patients is challenging, as safety constraints and high patient variability complicate simulations and underscore the need for simulators that capture dive...

Nguyen Khoi Hoang, Shuhaib Mehri, Tse-An Hsu et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

$\pi^2$: Structure-Originated Reasoning Data Improves Long-Context Reasoning Ability of Large Language Models

We study a QA curation pipeline for improving long-context complex reasoning in large language models (LLMs). Our approach, $\pi^2$, constructs high-quality reasoning data through rigorous QA curation: 1) extracting and expanding tables from Wikipedia, 2) from the collected tables together with relevant metadata, gener...

Quyet V. Do, Thinh Pham, Nguyen Nguyen et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

MDKeyChunker: What Does One LLM Call per Chunk Buy for Markdown Retrieval?

Markdown carries structure a parser reads for free: headers, section paths, and block boundaries. Many RAG pipelines also spend LLM calls per chunk on generated metadata. We ask what one LLM call per chunk buys over that free structure. MDKeyChunker splits Markdown into header-led chunks without splitting any block; ma...

Bhavik Mangla · 0 citations
#artificial intelligence Preprint Open access Oct 2026

COMIC: Agentic Sketch Comedy Generation

We propose a fully automated AI system that produces short comedic videos similar to sketch shows such as Saturday Night Live. Starting from character references, the system employs a population of agents loosely modeled on roles in real production studios and structured to optimize the quality and diversity of ideas a...

Susung Hong, Brian Curless, Ira Kemelmacher-Shlizerman et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Vulnerability of LLMs' Stated Beliefs? LLMs Belief Resistance Check Through Strategic Persuasive Conversation Interventions

Large Language Models (LLMs) are increasingly employed in various question-answering tasks. However, recent studies showcase that LLMs are susceptible to persuasion and could adopt counterfactual beliefs. We present a systematic evaluation of LLM susceptibility to persuasion under the \emph{Source--Message--Channel--Re...

Fan Huang, Haewoon Kwak, Jisun An · 0 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.