Skip to content

Category

natural language processing

6,613 papers

#machine learning Preprint Oct 2026

Learning to Learn a Language

We present the Prior-Fitted Language Model (PFLM), a 300M-parameter byte-level transformer pretrained only on samples from a synthetic non-linguistic prior. Given a prefix of real text, it learns to predict the language in context with frozen weights, having never seen a word of any real language. Every training sequen...

Lennart Carstens-Behrens, Holger Fröhlich · 0 citations
#machine learning Preprint Open access Oct 2026

Lend Me Your Eyes: Instruction-Aware Text Embeddings via Attention Relay

Text embedding models trained with contrastive learning learn to follow task instructions from instruction-paired data, while instruction-tuned LLMs already know how to follow them. We show that this instruction-following ability can carry over from an LLM to a Transformer-based embedder without any training. We propos...

Yiyuan Luo, Vaggos Chatziafratis · 0 citations
#machine learning Preprint Oct 2026

SALUS: Automated Auditing of NL-to-SQL Benchmarks through Weak Supervision of Multi-Agent Output

Natural language to SQL (NL-to-SQL) benchmarks are foundational to progress in data analysis research, yet recent work has shown that widely-used benchmarks contain significant annotation errors. These errors silently corrupt evaluation metrics, penalize correct model output, and distort the field's understanding of st...

Shi-Yuan Zhou, Ashwin Gerard Colaco, Sainyam Galhotra et al. · 0 citations
#machine learning Preprint Open access Oct 2026

Unmentioned Checklist Findings Change How Reinforcement Learning Appears to Improve Chest Radiograph Report Checking

Automated checks of radiology reports may rely on AI-generated checklists that leave findings unmentioned. We used reinforcement learning to train a vision-language model to fill in a 12-finding checklist from a chest radiograph without seeing the sentence under test; a separate checking model judged the sentence from...

Ali Vosoughi, Akhil Kasturi, Chenliang Xu et al. · 0 citations
#machine learning Preprint Oct 2026

Mind the Gaps: From Failure Attribution to Closed-Form Repair of Code Language Models

Code language models must be maintained like the software around them: when a library evolves, a model keeps writing the interface that it saw during training. Repairing the model itself lets one correction reach all downstream uses. Existing repair methods attribute a failure to neurons, select the highest-ranked ones...

Jian Gu, Hong-Yu Zhang, Chun-Yang Chen et al. · 0 citations
#machine learning Preprint Oct 2026

FORGE: Verification-Gated Behavioral Repair for Generative Language Models

Generative large language models (LLMs) inherit undesirable behaviors from pre-training, including demographic bias and toxic generation, that often emerge only after deployment and affect a small subset of inputs. A repair should eliminate the identified defect, preserve the model's overall functionality and, ideally,...

Hsin-Ling Hsu, Min-Yue Chen, Nai-Chia Chen et al. · 0 citations
#machine learning Preprint Oct 2026

vMF Sentence LDA: A Spherical Topic Model over Sentence Embeddings

Latent Dirichlet Allocation (LDA) and models derived from it remain widely used topic models. LDA observes each document as a bag-of-words and models each topic by a categorical distribution over the vocabulary, so that it uses neither the internal structure of the document nor the similarity in meaning between words....

Ryotaro Kobayashi, Yuri Murayama, Kiyoshi Izumi · 0 citations
#machine learning Preprint Oct 2026

SepRQ : Self-Supervised Speech Mixture Representation Learning via Mask-Free, Multi-Scale Source Separation

Self-supervised learning (SSL) is standard for speech representation learning, but mainstream models are designed around single-speaker audio, limiting their usefulness in multi-speakers scenarios. We present SepRQ, an open-source SSL framework that replaces masked prediction with a pseudo-source-separation objective o...

Séverin Baroudi, Hervé Bredin, R. Marxer · 0 citations
#machine learning Preprint Open access Oct 2026

Does Neural Complexity Improve Health Misinformation Detection? A Leakage-Controlled Cross-Corpus Benchmark

Increasing architectural complexity is often assumed to improve health misinformation detection, yet reported gains are difficult to interpret when studies use different corpora, preprocessing pipelines, data splits, and leakage controls. This study provides a controlled cross-corpus benchmark of five compact neural ar...

Mkululi SIKOSANA · 0 citations
#machine learning Preprint Open access Oct 2026

SIFT: Robust Meta-Faithfulness Verification of Chain-of-Thought Reasoning Under Distribution Shift

Chain-of-Thought (CoT) faithfulness detectors are widely used to audit reasoning models, yet a detector is itself a predictor whose verdicts are treated as stable properties. We ask whether a detector is faithful to itself under distribution shift. We formalize meta-faithfulness as an invariance principle: a valid dete...

Noor Islam S. Mohammad, Md. Basim Al Zabir Shammo, Hasan Siddiki et al. · 0 citations
#machine learning Preprint Oct 2026

Homogeneous Semantic Alignment and Hierarchical Expert Routing for Radiology Report Generation

Radiology report generation (RRG) aims to convert medical images into diagnostic texts to assist in clinical decision-making and alleviate the workload of physicians. Although existing methods have made extensive progress in cross-modal interaction and the incorporation of external priors, the distribution shift of und...

Er-Jian Zhang, Jia-Yuan Ma, Lie-Jun Wang et al. · 0 citations
#machine learning Preprint Open access Oct 2026

DV-Lens: Revealing the Functional Organization of Language Model Parameters

Understanding parameter functions helps elucidate the internal mechanisms of large language models (LLMs). However, how to connect parameters from different modules to verifiable output effects and further characterize the relationship between their functional organization and model capability remains to be explored. T...

Chenhang Cui, Jian Yu, Shuyi Miao et al. · 0 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.