Skip to content

Category

natural language processing

6,613 papers

#natural language process... Preprint Oct 2026

Multimodal Dual-Encoder Retrieval for Automated ICD Coding

Accurate International Classification of Diseases (ICD) coding is crucial for large-scale clinical research, documentation, and billing. There are three primary problems with current ICD prediction methods: (1) They are unable to comprehend multimodal patient data because they rely on either structured EHR data or unst...

Abhinav Bohra, Anuj Bohra · 0 citations
#natural language process... Preprint Open access Oct 2026

DimSteer: Steering LLM Authoring with Automatically Discovered Stylistic Controls

Large language model writing interfaces often make users steer outputs by repeatedly articulating desired changes in natural language. Yet writers may recognize useful stylistic directions only after seeing alternatives, making revision recall-heavy. We present DimSteer, an authoring interface that samples prompt-local...

Ajit Mallavarapu, Ziwei Gu · 0 citations
#computer vision Preprint Open access Oct 2026

PlotGround: Grounding Plot Digitization in Real Scientific Figures and Their Source Data

Scientific figures often encode quantitative results that are not readily available in machine-readable form, making accurate plot digitization important for verifying and reusing published findings. Yet it remains unclear how accurately current models recover plotted values from real scientific figures, as existing be...

Yaohui Zhang, Binxu Li, Haoyi Duan et al. · 0 citations
#natural language process... Preprint Open access Oct 2026

T-Search: An Open Agentic Retriever and Playground for Hard Multi-Step Search

We present T-Search, an open-weight agentic retriever for hard multi-step search. Given a question and a search tool over a fixed corpus, it runs a bounded multi-round search and returns a ranked list of evidence chunks with short justifications, leaving answer generation to a downstream model, so backend and generator...

Olga Tsymboi, Ramil Latypov, Aleksandr Medvedev et al. · 0 citations
#natural language process... Preprint Oct 2026

Improving Diversity in LLM Short Story Generation

Large language models (LLMs) can generate accurate responses, but these are void of diversity. We attempt to address this for the task of creative short story generation. Drawing on established writing conventions and known LLM limitations, we target variation in genre, tone, style, and named entities. To promote diver...

Zahra Solati Dehkordi, Vasileios Lampos · 0 citations
#natural language process... Preprint Open access Oct 2026

Domain adaptation of Russian ModernBERT for long legal documents

We investigate whether continued pretraining on Russian legislative documents improves a Russian ModernBERT encoder on legal text. The adapted model, RuModernBERT-ruLaw, was trained on a corpus reported to contain 304,382 legislative documents and 194,425,905 corpus tokens. Corpus token counts are distinguished from po...

I. Litvak, D. Gvozdetsky, F. Lashkin et al. · 0 citations
#natural language process... Preprint Open access Oct 2026

SAFE-MR: Evidence Sufficiency Learning for Selective Multimodal Rumor Detection

Multimodal rumor detectors increasingly rely on retrieved evidence, yet relevant evidence is not necessarily sufficient for verification. Missing provenance, duplicated reports, and unresolved contradictions can produce confident predictions without adequate support. We introduce SAFE-MR, a framework that separates cla...

Shiwen Ni · 0 citations
#natural language process... Preprint Open access Oct 2026

MedPrune: Topology-Efficient Multimodal Multi-Agent Communication Evolution for Medical VQA Tasks

While medical multimodal large language models (Med-MLLMs) advance medical visual question answering (VQA), existing clinical workflow-inspired multi-agent frameworks suffer from interaction patterns and excessive computational overhead caused by redundant communication topologies. In this paper, we propose MedPrune, a...

Jiuheng Wan, Runze Li, Chen Chen et al. · 0 citations
#natural language process... Preprint Oct 2026

Programmatic Search Agents: Extending Agentic Search Beyond Query Reformulation

Search agents adapt their queries, yet fixed search interfaces leave candidate processing and evidence presentation outside the agent's direct control. Our trajectory analysis shows that supporting passages can be retrieved yet never delivered to the agent; a same-page oracle intervention shows that changing the return...

Jia-Ming Qian, Hui-Yan Yang, Man-Di Liu et al. · 0 citations
#natural language process... Preprint Oct 2026

Reward Stealing Attack on Large Language Models

Adversarial attacks on Large Language Models (LLMs) aim to induce harmful content. However, existing methods suffer from high computational costs or strict model-pairing dependencies, limiting their scalability and transferability. We propose Reward Stealing Attack (ReSA), an adversarial attack framework that targets t...

Jia-Ming Qian, Peng-Yang Zhou, Jia-He Xu et al. · 0 citations
#natural language process... Preprint Oct 2026

Wikidata Search Traces: A Dataset for Training Knowledge Graph Search Agents

Wikidata is one of the largest open knowledge bases, yet answering a complex question over it still requires a SPARQL query that names the right entities and properties and chains their relations. Language models offer a natural-language alternative but answer largely from memory, which is least reliable for less promi...

Mohamed Chenene, C. Rosas Hinostroza, A. Stasenko et al. · 0 citations
#natural language process... Preprint Oct 2026

LoGRA: Scaling LLM Reinforcement Learning with Low-Rank Gradient Sketches

Reinforcement learning has greatly advanced the capabilities of large language models, but its memory demands remain a barrier to broader adoption. We introduce LoGRA, an approach to RL post-training that reduces memory by retaining useful learning signals in low-rank gradient sketches. These compact representations su...

Shao-Kun Zhang, Yi-Fan Zhang, Jian Hu et al. · 0 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.