Skip to content

Category

natural language processing

6,613 papers

#computer vision Preprint Open access Oct 2026

World Embedding Benchmark

Physical fidelity has received increasing attention in world models and video generation, yet how video representations encode physical information remains less understood. We introduce the World Embedding Benchmark, comprising 8,000 controlled simulation cases from 80 families spanning fluid mechanics, solid mechanics...

Yiqi Liu, Ruifeng Yuan, Yang Wang et al. · 0 citations
#natural language process... Preprint Open access Oct 2026

Passing the Test You Trained On: Re-evaluating Prompt-Injection Detectors for LLM Agents

LLM agents increasingly screen tool outputs with small prompt-injection detectors, and teams choose among detectors by their scores on public benchmarks. We ask whether those scores predict how a detector behaves inside an agent. We replay the ground-truth tool calls of two agent benchmarks, AgentDojo and tau-bench, wi...

Zhuowen Liu · 0 citations
#natural language process... Preprint Open access Oct 2026

Benchmarking Literature Retrieval for a Model Organism: A Dictyostelium Case Study

Biological literature retrieval systems are often developed and evaluated using broad biomedical corpora and general-purpose search tasks. However, many curated knowledge bases operate in narrower model-organism domains, where the literature is sparse and terminology is organism-specific. We introduce a retrieval bench...

Yun Wang, Gad Shaulsky, Toma\v{z} Curk et al. · 0 citations
#natural language process... Preprint Open access Oct 2026

SecJev: Bringing Security Expertise to System One Decision Models

Security workflows need models that turn complex observations and explicit policies into decisions. System One models introduced by Jev return typed predictions and probabilities; security specialization supplies the domain expertise behind those predictions. We introduce SecJev, to our knowledge the first family of Je...

Zheng Chen, Fei Yu, Haohao Huang et al. · 0 citations
#computer vision Preprint Oct 2026

ReSCUE: Re-translation with Sentence Commitment for Unsegmented Long-Form Simultaneous Sign Language Translation

Simultaneous Sign Language Translation (SLT) is critical for real-time communication, yet existing methods remain largely confined to sentence-level, offline settings that assume pre-segmented inputs. These assumptions hinder deployment in realistic scenarios involving continuous, unsegmented video streams. We present...

Si-Han Ren, Gao-Zheng Li, Yuan-Shang Quan et al. · 0 citations
#natural language process... Preprint Open access Oct 2026

Asterism: Exploring and Synthesizing Scattered Observations into Literature-Grounded Hypotheses and Theories

A theory draws many independent observations into one framework with novel hypotheses. A researcher building such a theory must synthesize observations scattered across many papers, each describing related concepts but often in different terms. Which concepts matter most also depends on their preferences and research q...

Joseph Chee Chang, Michael D'Arcy, Amy X. Zhang et al. · 0 citations
#natural language process... Preprint Open access Oct 2026

Social bot detection in the age of ChatGPT: Challenges and opportunities

We present a comprehensive overview of the challenges and opportunities in social bot detection in the context of the rise of sophisticated AI-based chatbots. By examining the state of the art in social bot detection techniques and the more salient real-world application to date, we identify gaps and emerging trends in...

Emilio Ferrara · 0 citations
#natural language process... Preprint Oct 2026

SEDIMA: Cross-Run Hierarchical Insight Memory for Evolutionary Search Agents

Large language model (LLM)-driven evolutionary search is a powerful paradigm for automated program and algorithm discovery, yet existing systems are largely memoryless: each run explores from scratch, so agents repeatedly rediscover the same improvements and re-encounter the same dead ends. We introduce SEDIMA, a persi...

Amirhossein Abaskohi, Mahdi Mostajabdaveh, Zi-Rui Zhou · 0 citations
#natural language process... Preprint Open access Oct 2026

Language Models that Play Chess and Explain Their Moves

Modern chess engines are silent experts: they play at a superhuman level, but do not offer explanations for their play. On the other hand, language models (LMs) can generate plausible-sounding explanations, but their weak playing strength limits the utility of their explanations. We introduce Queen, a 4B-parameter ches...

Adithya Bhaskar, Jeffrey Cheng, Danqi Chen · 0 citations
#natural language process... Preprint Open access Oct 2026

Writerslogic at the CLEF 2026 SimpleText Track: Multi-Candidate LLM Simplification and Stacked Complexity Spotting

We describe the Writerslogic team's participation in the CLEF 2026 SimpleText shared task, addressing Task 1 (text simplification) and Task 2 (complexity spotting). For Task 1, we develop a multi-candidate generation pipeline using GPT-4o-mini that produces five simplification candidates per sentence at varying tempera...

David L. Condrey · 0 citations
#natural language process... Preprint Oct 2026

Writerslogic at PAN 2026: Process over Content for Robust Detection under Domain Shift

We describe the Writerslogic systems for three PAN at CLEF 2026 shared tasks (Reasoning Trajectory Detection, Voight-Kampff Generative AI Detection, and Multi-Author Writing Style Analysis), unified by a shared analytical framework: feature robustness under distribution shift is governed by support overlap between trai...

David Lee Condrey · 0 citations
#natural language process... Preprint Oct 2026

Author Representation Strategies for Zero-Shot Authorship Attribution: A Comparative Study of LLM-Based and Embedding-Based Approaches

Authorship Attribution (AA) requires capturing fine-grained stylistic characteristics, making it particularly challenging in zero-shot (ZS) settings where no task-specific supervision is available. In this work, we investigate the effect of author representations on ZS AA by evaluating a label-only prompting baseline t...

Nudrat Habib, Tosin P. Adewumi, Sana Al-azzawi et al. · 0 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.