Skip to content

Category

natural language processing

6,613 papers

#machine learning Preprint Oct 2026

Saying, Not Knowing: Aggressively GGUF-Quantized Small Language Models Still Write Rare Words They Can No Longer Define

Post-training quantization to the GGUF format's mixed-precision K-quants is commonly how open-weight language models reach consumer hardware, yet its effect on fine-grained lexical competence is uncharacterized. We audit 27 quantized artifacts across 13 families and four architecture backbones, 0.35B-14B parameters, ev...

Saurabh Kumar Singh, Yogeshwar Singh Dadwhal, Malhar Vedak · 0 citations
#machine learning Preprint Oct 2026

Conformal Prediction with Paraphrase-Aware Scoring for LLM Uncertainty Quantification

Uncertainty quantification (UQ) for large language models (LLMs) aims to provide reliable measures of predictive confidence, yet current methods are often unstable under meaning-preserving perturbations. Semantically equivalent paraphrases can induce substantial variability in predictive confidence, even for methods wi...

Jiayi Xin, Evan Qiang, Zi-Han Zhu et al. · 0 citations
#machine learning Preprint Oct 2026

What Gradients Add to Text Leakage in Split Language Models, Counted per Token and per Document

Split learning lets a client train a language model on a server without sending its text. The client runs the first layers itself and sends the server only their output, a vector of numbers for each token. During training, the server sends gradients back. We show that an observer at the split can rebuild most of the cl...

G. Politis, E. Pappas · 0 citations
#machine learning Preprint Open access Oct 2026

Representation-Aligned Auxiliary Supervision for Language Model Adaptation

Language models exhibit strong reasoning capabilities, yet adapting them to structured domains remains challenging and can yield inconsistent outcomes. We identify representation compatibility, the extent to which a model effectively processes a representation for a structured task, as a key factor in adaptation. We st...

Kyuyoung Kim, Peiyao Sheng, Ashwin Hebbar et al. · 0 citations
#machine learning Preprint Open access Oct 2026

Periscope: Extending Frozen Language Models Beyond Their Context Window

A language model reads long text in one quadratic forward pass, stops at the context window, and loses accuracy with length before reaching it. We ask whether the read can be factorized when deciding over a finite set: which document is relevant, which option is supported, which passage is the evidence. Periscope, a tr...

Mohamed Eltahir, Anas Obayd, Raed Rashid et al. · 0 citations
#machine learning Preprint Open access Oct 2026

OncoNoteBERT: A Foundation Representation Model for Natural Language Processing of Real-World Outpatient Oncology Notes

Real-world outpatient oncology notes contain specialised terminology, tumour staging expressions, treatment names, toxicity descriptions, and institution-specific de-identification markers that may not be represented efficiently by general biomedical or adjacent clinical language models. We developed and evaluated onco...

Wuraola Oyewusi, Eliana Vasquez Osorio, Goran Nenadic et al. · 0 citations
#machine learning Review Oct 2026

Same Output, Different Gold: Measuring How Reference Choice Moves a Multilingual Benchmark Score

A benchmark score compares a system output against a reference, and methodological attention falls almost entirely on the first term. We measure the second. The retained annotation record of a six-language benchmark for personally identifiable information contains two independent annotator labellings, the aggregate shi...

Parth Kulshreshtha, Shivali Dalmia, Abhishek Mukherji · 0 citations
#artificial intelligence Preprint Open access Oct 2026

The Llama 3 Herd of Models

Modern artificial intelligence (AI) systems are powered by foundation models. This paper presents a new set of foundation models, called Llama 3. It is a herd of language models that natively support multilinguality, coding, reasoning, and tool usage. Our largest model is a dense Transformer with 405B parameters and a...

Aaron Grattafiori (Jack), Abhimanyu Dubey (Jack), Abhinav Jauhri (Jack) et al. · 0 citations
#machine learning Preprint Open access Oct 2026

Aligning Multimodal Patient Evidence with Biomedical Knowledge Graphs for Clinical LLMs

Clinical questions often depend on linking a patient's multimodal evidence to external biomedical knowledge, yet existing predictive systems rarely represent such links explicitly, so they can neither be traced to their evidence sources nor removed to measure their contributions. We present MM-KG (Multimodal Knowledge...

Jiawen Du, Arshan Ali Khan, Chenhao Zhang et al. · 0 citations
#machine learning Preprint Open access Oct 2026

How Sparse Probability Maps Shape Mixture-of-Experts Routing

Mixture-of-experts (MoE) routers typically apply softmax to the router scores and keep the top-K experts, making every token use exactly K experts. Sparsity-inducing probability maps such as sparsemax, alpha-entmax and normmax can adaptively assign exact zeros to selected experts, and therefore appear to offer token-de...

Tom\'as Brogueira, Marcos Treviso, Miguel Couceiro · 0 citations
#machine learning Preprint Oct 2026

What Matters for Latent Reasoning with Flow Matching

Latent reasoning lets a large language model (LLM) think in a continuous space and verbalize only the answer. We argue that an effective latent thought must meet five requirements: it should be useful, helping produce the correct answer rather than merely changing it, diverse, so that resampling yields different reason...

Yassine Ouali, Adrian Bulat, G. Tzimiropoulos · 0 citations
#machine learning Preprint Open access Oct 2026

HeuFouFT: Task-Guided Metaheuristic Coordinate Search for Fourier Fine-Tuning

We introduce Heuristic-Guided Fourier Fine-Tuning (HeuFouFT), a task-guided framework for selecting trainable frequency coordinates in Fourier fine-tuning. Existing uniform and Gaussian band-pass schemes allocate a limited spectral budget through fixed, task-agnostic rules. HeuFouFT instead searches for coordinates usi...

Ruiheng Wang, Yubo Hou, Yakun Zhu et al. · 0 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.