Skip to content

Category

natural language processing

6,613 papers

#artificial intelligence Preprint Open access Oct 2026

Just for FUNS: LLM-Guided Spatio-Temporal Graph Node Generation for Forecasting Unobserved Node States

Spatio-temporal forecasting is a cornerstone of logistics, urban planning, and intelligent transportation systems. However, constrained by deployment costs and maintenance resources, sensor networks often lack comprehensive spatial coverage, rendering Forecast Unobserved Node States (FUNS) a critical yet formidable cha...

Shuhao Li, Weidong Yang, Changan Liu et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Tokka-Bench: Evaluating Tokenizers Across 100 Natural and 20 Programming Languages

Large language models rely on subword tokenizers whose quality varies across languages, yet no standardized multi-metric framework exists for broad comparative evaluation. We introduce Tokka-Bench, an open-source framework that evaluates tokenizers on five complementary metrics -- bytes per token, unique token coverage...

Ben Gubler · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Validity Without Ground Truth: What Stated-Preference Economics Offers the Evaluation of Language Models

Many of the questions now put to large language models have no correct answer to score against: what a policy is worth, which option a user should choose, how to weigh competing values. Stated-preference economics has faced this problem for decades. It judges survey responses without knowing the true value, through a f...

Daniel Robert Kling Alexander, Catherine Louise Kling · 0 citations
#artificial intelligence Preprint Open access Oct 2026

RunningTab: Direct Workspace Interaction with Environment-Side Tabs

Much knowledge work produces new deliverables from files a workspace already holds, and LLM agents are beginning to take such work over. Through direct corpus interaction, an agent can search and read any of those files from a terminal with no indexing, and producing a deliverable from many of them in this way is what...

Jinheon Baek, Soyeong Jeong, Yumin Choi et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Reasoning-Token Spikes Under Prompted Untruthful Responding in Large Language Models

Monitoring the chain-of-thought of reasoning artificial intelligence (AI) models remains a key approach to detecting deception and other forms of misbehavior in such models. However, semantic chain-of-thought monitoring depends on reasoning traces being legible and sufficiently faithful to the underlying computations t...

Maverick Morales, Tom\'a\v{s} Dominik, Vermut Gao et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

GAGR-Lab: Evaluating Joint Spatial-Geometric and Analytic Function Reasoning

Joint spatial-geometric and analytic function reasoning requires translating a perceived spatial configuration into a symbolic function whose executed curve satisfies geometric constraints. We present GAGR-Lab, a framework for measuring this capability through Cartesian game scenes, explicit function semantics, and aut...

Jingyao Zhang, Yun Li, Lu Han · 0 citations
#artificial intelligence Preprint Oct 2026

SkillSandbox: Skill Verification via Dynamic Scenario Synthesis

Self-evolving agents distill task-solving experience into skills for future reuse, but these skills can encode incorrect procedures or non-transferable knowledge. It is therefore critical to verify each skill's reusability: whether its guidance remains useful beyond the experience from which it was distilled. Such veri...

Serin Kim, Kwangwook Seo, Dok-Yung Song et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

LiveMACE: Process-Aware Evaluation of LLM Agent Capabilities in Evolving Markets

Evaluating agents by outcomes alone can obscure the capabilities that produce them. This problem is especially pronounced in evolving environments, where outcomes reflect a closed-loop interaction between agent behavior and changing external conditions. We introduce LiveMACEBench, a process-aware benchmark that uses li...

Jun Zhao, Leiming Fu, Yanbo Wen et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

From Expert-Guided Proof Search to Automated Open-Problem Solving

Large language models are increasingly contributing to mathematical research, where progress often depends on efficient proof search, incremental improvements and careful verification. We describe Bolzano, a multi-agent open-source system that uses parallel prover agents with a verifier agent and maintains a human-read...

Adri\'an Z\'ame\v{c}n\'ik, Mat\v{e}j Kripner, Martin Kouteck\'y et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

TopoGraphRAG-Bench: Evaluating Multimodal GraphRAG on Layout-Grounded Evidence Reasoning

Real-world documents distribute evidence across text, tables, figures, and captions within complex page layouts. Answering complex questions over such documents therefore requires more than retrieving relevant passages: systems must recover the evidence topology that connects heterogeneous evidence units. Existing Grap...

Ruochi Li, Jianzhe Lin, Haoxuan Zhang et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Trajectory Abstraction for the Science of Language Agent Behavior

Scientific studies of language agents need behavioral variables that support hypotheses across tasks and models. We formulate this research problem as learning and testing a hierarchy of trajectory abstractions. A concrete recursive procedure first measures role- and phase-indexed events, proposes temporally constraine...

Tianqiang Yan · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Few Bits, One Law: Toward W2A4KV2

Extreme low-bit LLM compression is most challenging when weights, activations, and KV caches are quantized together: their distributions differ, and quantization errors interact throughout the network. We introduce CanonQ, a unified quantization-aware training framework that addresses these challenges by separating sou...

Kai Yi, Tarek Elgamal, Sruthikesh Surineni et al. · 0 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.