Skip to content

Category

natural language processing

6,613 papers

#artificial intelligence Preprint Open access Oct 2026

Grammar Concept Annotation at Scale: Deployed Fine-Tuned Small Language Models Outperform Prompted Frontier Models

Corrective feedback is among the best-evidenced drivers of second-language acquisition, yet corrections delivered during lessons rarely accumulate into an actionable view of grammar mastery. Prompted frontier models can provide such a view from learner--tutor lesson transcripts, but they are costly at scale. We close t...

Marjan Celikik, Ana Peleteiro Ramallo, Javier Morales · 0 citations
#artificial intelligence Preprint Open access Oct 2026

NavGPT-3: Harnessing Context in a Hierarchical Navigation Runtime

Language models trained with long-horizon agentic reinforcement learning can generalize knowledge through reasoning, express precise actions, and pursue goals over many steps, raising the ceiling on what an embodied agent can understand and decide. Physical interaction, however, remains the domain of action policies, w...

Gengze Zhou, Yicong Hong, Jiazhao Zhang et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Clarify, Then Focus: Statement Normalization for Conversation Analytics at Scale

Enterprise conversation analytics asks many questions of millions of interactions. Each question can require reconstructing what people mean and identifying which information matters, repeating costly interpretive work across the same transcripts. We propose a simple principle: clarify the text, then focus the reader....

Mikhail L. Arbuzov (Independent researcher), Karan Dave (Independent researcher), Evgeniya Dontsova (Independent researcher) et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Conversational Task Disambiguation over Tabular Data: Leakage-Aware Formulation, Benchmark Suite, and Training

Conversational task disambiguation over tabular data uses dialogue to resolve missing information about a user's intended task before producing a solution over tables or databases. Existing evaluation and training lack a leakage-aware foundation. Task success mixes the agent's disambiguation and solution-generation cap...

Nafiseh Ghoroghchian, Luis Scoccola, Tina Sedaghat et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

WorldBench: Evaluating LLMs on Three.js Voxel World Generation

Large language models can now write complete, interactive 3D worlds as code, but grading those worlds automatically is unreliable. Existing judges take one view of the output: a vision-language model scores a few rendered snapshots, or a language model reads the source. On worlds written by five frontier models we find...

Krish Bakshi · 0 citations
#artificial intelligence Preprint Open access Oct 2026

OnTrack: Real-Time Monitoring and Intervention in LLM Agent Trajectories via Streaming Structure-Aware Optimal Transport

Agents are deployed in applications from trip planners and stock trading to IT incident triage. In most cases, LLM agents work autonomously with minimal rule-based safeguarding, leading to cost and safety issues from irreversible actions. Recent works resolve this either by using a safeguard agent to monitor behavior o...

Babak Barazandeh, Connor Swanson, Chinmay Kulkarni et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Cited but Not Consulted: A Counterfactual Audit of Legal Chain-of-Thought Faithfulness

Large language models increasingly justify legal decisions by naming the statute or precedent behind a verdict, treated as evidence that the decision follows from it. We test this directly: holding case facts fixed, we substitute the named legal authority for an unrelated one and decode a model's evolving verdict from...

Saisab Sadhu, Shreeyans Arora, Pratinav Seth · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Accurate but Not Humble: Evaluating Epistemic Humility in LLM Agents under Knowledge Conflict

When retrieved evidence contradicts an agent's prior beliefs, does it revise its answer, acknowledge uncertainty, or persist with an incorrect conclusion? Existing evaluations of agentic systems focus primarily on task success, offering limited insight into how agents handle such conflicts. We propose to evaluate agent...

Kaiser Sun, Bernal Jimenez Gutierrez, Hongjun Liu et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Overcoming Prior Barriers: Supervised Fine-Tuning under Long-Tail Distribution

Supervised fine-tuning (SFT) adapts pretrained large language models (LLMs) to downstream tasks, but the required concepts can receive substantially different levels of pretrained support. Frequent concepts are more likely to be well learned, whereas rare concepts may remain weakly represented. We introduce a novel not...

Haohui Wang, Jiahao Xu, Wangzhi Zhan et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Can AI Agents Learn Their Way to the Top? Evaluating Heuristic Learning in a Long-Running Game Agent Competition

Adversarial games have driven advances from heuristic search to reinforcement learning, yet learning and adapting strategies from limited samples remain challenging. AI agents offer an alternative by turning game experience into revisions of executable policies. Building on heuristic learning (HL), we formalize Adversa...

Kaisen Yang, Qingle Liu, Kejin Wang et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Verdict Without the Rule: Diagnosing and Auditing Regulatory Rule Sensitivity in LLM Compliance Systems

Large language model compliance systems are deployed on the assumption that a verdict depends on the regulatory rule it is given. We test this directly across five models and 20 regulatory and platform-policy domains: delete, swap, or negate the governing rule while holding the case fixed, and check whether the verdict...

Saisab Sadhu, Aadit Sengupta, Vinay kumar Sankarapu et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

When Should Agents Think? Adaptive Reasoning via Cross-Turn Estimation

Large language model (LLM)-based agents have demonstrated strong capabilities on complex tasks. They typically perform reasoning before each action throughout an interaction trajectory. However, reasoning may not be necessary at every turn, as reasoning produced earlier can continue to support subsequent actions. A key...

Yiruo Cheng, Shen Huang, Xiaoshuai Song et al. · 0 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.