Skip to content

Category

small language model

2,863 papers

#artificial intelligence Preprint Sep 2026

MM-FinEval: A Multi-Task Multimodal Benchmark for Real-World Financial Forecasting

Financial forecasting from earnings conference calls requires models to reason over complex corporate disclosures, market expectations, and subtle communication signals. However, existing financial benchmarks are often limited to unimodal inputs or single-task settings, making it difficult to evaluate whether multimoda...

Dong Shu, Yan-Guang Liu, Huo-Pu Zhang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Better Deck or Different Judge? Evaluating Agentic Harness Gains in Corporate and Investment Banking

Corporate and investment banking teams use presentations to support credit decisions and advise clients on financing and transactions. Producing these decks requires reconciling financial data, tracing sources and turning analysis into a recommendation. We retrospectively study the development of an agentic harness com...

Ludovic Gibert, Matis Despujols, André Rochet · 0 citations
#artificial intelligence Preprint Sep 2026

Inferring Causal Relations between Two Sequences of Events with Language Models

It is shown in this study that it is possible to leverage the predictive power of Large Language Models (LLMs) to infer causal relations between only two sequences of events, which provides better results than standard causal discovery algorithms on several time series data, even though these data were converted into s...

Nishchal Prasad, Éric Gaussier, Emilie Devijver et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Targeted Retrieval, Compact Representations: How CoT Reasoning Improves Long-Context Counting

Large language models (LLMs) have been rapidly improving in long-context tasks, powered by Chain-of-Thought (CoT) reasoning. However, the internal mechanisms underlying this improvement remain unclear. We investigate these mechanisms through a needle-in-a-haystack (NIAH) counting task, where an LLM is asked to count th...

Liang Twist Shan, Tian-Yu Hu, Hao Yan et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Concept-Grounded Attention: A Controlled Evaluation of Graph-Injected Attention, Temporal Versioning, and Epistemic Status

Knowledge-intensive language-model systems typically represent external knowledge as text chunks or static graphs, with limited support for concept evolution, point-in-time reasoning, and distinctions between validated and inferred knowledge. We introduce the Concept Lifecycle Model (CLM), which represents concepts as...

S. Duggal, Pradyumna Swarnalatha Ramanna, Alexandros Vassiliades · 0 citations
#data science Oct 2026

Scientometric Analysis of Research Structure and Trends in Neutrino and Dark Matter Detection: Scientific Mapping and Keyword Co-occurrence Analysis in Web of Science

Purpose: Neutrinos and dark matter are two fascinating and mysterious topics in modern physics, with numerous experiments conducted worldwide to study their properties, interactions, and their impact on our understanding of the universe. Neutrino detection is primarily achieved through two methods: charged current inte...

Zeinab Sadat Tabatabaei Lotfi, Mohammad Mahdi Ettefaghi, ًReza Moazemmi · 0 citations
#machine learning Preprint Sep 2026

Explore Broadly, Reason Sharply: Push Small Models toward the Frontier via Sampling

Power-sharpened sampling is an inference-time alternative to reinforcement-learning (RL) post-training for enhancing reasoning in large language models (LLMs). High-probability sequences are amplified under the base model without parameter updates or external rewards, avoiding the costly optimization and jagged general...

Panagiotis Theodoropoulos, Nan Jiang, Xin-Tong Duan et al. · 0 citations
#machine learning Preprint Sep 2026

Storage Is Not Strategy: State-Conditioned Support Control for LLM Unlearning

Many localized large language model (LLM) unlearning methods select a small parameter subset from a localization signal and keep it fixed during optimization. The parameters most associated with a target, however, need not be the best ones to update, and candidate interventions can change value as optimization proceeds...

Tian-Hao Qian, Zi-Ming Hong, Chong-Yang Gao et al. · 0 citations
#natural language process... Preprint Sep 2026

LongHarness Bench: Stress-Testing Language Model Harnesses for Long-Context Reasoning

Language-model (LM) harnesses enable LMs to operate effectively over long contexts using additional compute. However, existing long-context evaluations are insufficient for distinguishing modern harnesses, reflected by saturated accuracy across harnesses and largely similar evaluation costs. In this paper, we introduce...

Quang Hieu Pham, Thuy Duong Nguyen, Jocelyn Qiaochu Chen et al. · 0 citations
#machine learning Preprint Sep 2026

Delta-Matching: Closing the Final Gap of Native 8-bit Training for LLMs

Reliable FP8 attention remains a barrier to fully native 8-bit large language model training. We derive how forward-backward inconsistencies produce stale delta and empirically show how it distorts training dynamics. Our stale-delta hybrid runs show a modest loss gap at 569M parameters but substantial loss increases an...

Hao-Zhan Tang, Hao Kang, Han Cai et al. · 0 citations
#machine learning Preprint Sep 2026

Why Adaptive Optimizers Underestimate Rare Tokens

In the softmax output layer, a rare token receives a small positive logit gradient on most steps and a much larger negative gradient on the few steps when it is the target. SGD simply adds these contributions. Coordinate-wise adaptive methods such as Adam, RMSProp, and sign descent instead divide each update by a runni...

Sangsidhya Kar · 0 citations
#natural language process... Preprint Sep 2026

MERGE: Multi-LLM Ensemble for Retrieval via Generative Enrichment

Large Language Models (LLMs) are increasingly used to enrich user queries in information retrieval (IR) so that a standard retriever such as BM25 can bridge vocabulary gaps with the target corpus. Any single LLM, however, is limited by its training data and architectural biases, and its enrichment behavior depends on h...

Tzu-I Ho, Yung-Yu Shih, Shang-Yu Su et al. · 0 citations

From tech blogs

See all →

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.