Skip to content

Category

natural language processing

6,613 papers

#artificial intelligence Preprint Open access Oct 2026

Recursive Agent Optimization

We introduce Recursive Agent Optimization (RAO), a reinforcement learning approach for training recursive agents: agents that can spawn and delegate sub-tasks to new instantiations of themselves recursively. Recursive agents implement an inference-time scaling algorithm that naturally allows agents to scale to longer c...

Apurva Gandhi, Satyaki Chakraborty, Xiangjun Wang et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Who Guards the Benchmarks? Automated Auditing of LLM Agent Benchmarks

As benchmarks grow in complexity, many apparent agent failures are not failures of the agent at all---they are failures of the benchmark itself: broken specifications, implicit assumptions, and rigid evaluation scripts that penalize valid alternative approaches. We propose employing frontier LLMs as systematic auditors...

Xinming Tu (Minta), Tianze Wang (Minta), Yingzhou (Minta) et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

How Do AI Agents Spend Your Money? Analyzing and Predicting Token Consumption in Agentic Coding Tasks

The wide adoption of AI agents in complex human workflows is driving rapid growth in LLM token consumption. When agents are deployed on tasks that require a significant amount of tokens, three questions naturally arise: (1) Where do AI agents spend the tokens? (2) Which models are more token-efficient? and (3) Can agen...

Longju Bai, Zhemin Huang, Xingyao Wang et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Rhetorical Questions in LLM Representations: A Linear Probing Study

Rhetorical questions are asked not to seek information but to persuade or signal stance. How large language models internally represent them remains unclear. We analyze rhetorical questions in LLM representations using linear probes on two social-media datasets with different discourse contexts, and find that rhetorica...

Louie Hong Yao, Vishesh Anand, Yuan Zhuang et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Is a Picture Worth a Thousand Words? Adaptive Multimodal Fact-Checking with Visual Evidence Necessity

Automated fact-checking is a crucial task that supports a responsible information ecosystem. While recent research has progressed from text-only to multimodal fact-checking, a prevailing assumption is that incorporating visual evidence universally improves verification accuracy. In this work, we challenge this assumpti...

Jaeyoon Jung, Yejun Yoon, Kunwoo Park · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Many Preferences, Few Policies: Compact Portfolios for Multi-Objective LLM Alignment

Aligning large language models (LLMs) requires balancing competing objectives such as helpfulness, harmlessness, and conciseness. The appropriate balance varies across users and applications, yet training, evaluating, and deploying many policies across different reward weights is costly. We study how to identify a smal...

Cheol Woo Kim, Jai Moondra, Roozbeh Nahavandi et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Sensory-Aware Sequential Recommendation via Review-Distilled Representations

Sequential recommenders learn behavioral patterns from item identifiers, while the experiential properties that users describe in reviews, such as how products look, feel, smell, taste, or sound, rarely enter item representations in a controlled, auditable form. We present ASER (Attribute-based Sensory-Enhanced Repre...

Yeo Chan Yoon, Chanjun Park, Kyuhan Koh · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Morality is Contextual: Learning Interpretable Moral Contexts from Human Data with Probabilistic Clustering and Large Language Models

A key question in current AI alignment research is how to make AI algorithms learn moral values. Because human morality is highly context-dependent, actions are judged not only by their outcomes but by the context in which they occur. We present COMETH (Contextual Organization of Moral Evaluation from Textual Human inp...

Geoffroy Morlat, Marceau Nahon, Augustin Chartouny et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

A Unified BERT-CNN-BiLSTM Framework for Simultaneous Headline Classification and Sentiment Analysis of Bangla News

In our daily lives, newspapers are an essential information source that impacts how the public talks about present-day issues. However, effectively navigating the vast amount of news content from different newspapers and online news portals can be challenging. Newspaper headlines with sentiment analysis tell us what th...

Mirza Raquib, Munazer Montasir Akash, Tawhid Ahmed et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Evaluating the Retrieval Robustness of Large Language Models

Retrieval-augmented generation (RAG) generally enhances large language models' (LLMs) ability to solve knowledge-intensive tasks. But RAG could also lead to performance degradation due to imperfect retrieval and the model's limited ability to leverage retrieved content. In this work, we evaluate the robustness of LLMs...

Shuyang Cao, Karthik Radhakrishnan, David Rosenberg et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

A Multi-Timescale Recursive Self-Improvement Engine for Open-Ended Persona Growth

Role-playing AI personas today do not grow: they hold a fixed character, so the relationship a user builds with them has nothing to accumulate on. We introduce AutoPersonas, a multi-timescale engine that applies recursive self-improvement (RSI) to persona growth: rather than improving its intelligence, the persona recu...

Mengchen Li · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Counterfactual Evidence Audits Predict LLM-Agent Susceptibility to Ranked Context

LLM agents increasingly decide from evidence assembled by upstream systems: retrievers choose documents, recommenders choose posts, and memory systems choose prior events. Existing evaluations usually hold this evidence fixed, missing failures in which individually ordinary items form a systematically one-sided context...

Rana Muhammad Usman · 0 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.