Skip to content

What Do Tabular Foundation Models Compute In Context? In-Situ Representation Refinement through Attention-Gated Updates

Sep 2026 · 0 citations · 21 references
Computer Science

TL;DR

In-situ representation refinement is developed: support labels guide changes to the episode's representations, improving the information available to later queries, in a competitive, memory-efficient model.

Abstract

A tabular foundation model must discover which distinctions matter for each new table without updating its parameters. We develop in-situ representation refinement: support labels guide changes to the episode's representations, improving the information available to later queries. A regularized leave-one-out objective yields a support correction and its query extension. The leading term separates attention-based reading from state-dependent scaling, motivating RefineICL: an attention-gated, FFN-free contextual stack with selected low-rank feature interaction and typed memory. A direct intervention tests the role of evolving support states: removing one intermediate support update while preserving the block's query output increases final query cross-entropy in all 72 tested episodes. RefineICL-L24 reaches 0.93836 OVR-AUC and 0.87173 accuracy on AMLB29. A benchmark-informed continuation reaches 1644.8 Elo on the 38-dataset TabArena snapshot, 31.4 Elo above TabPFN-3 under the same evaluation. It also improves all four reported metrics over TabPFN-v3 on both TabZilla views. In a matched 100K-update depth grid, an expanded FFN gives no consistent validation benefit and uses 60.2% more peak inference memory at L8. These results connect learning within a forward pass to representation refinement and show how this view guides a competitive, memory-efficient model.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

Towards Evolving Context Parameterization for Large Language Models

This work proposed PLUME, a training-free method that constructs a global update representation, activates memory evidence to form a local parameter view, and adaptively integrates their predictions during decoding, which demonstrated its effectiveness in sequential evolution settings.

Xiao Shi, Zhe-Rui Li, Yi-Ming Jiang et al. · 0 citations

The Evolution of Attention in Large Language Models: Mechanisms, Trade-offs, and Emerging Trends

Self-attention gives LLMs fine-grained, query-dependent access to context, but dense token interactions incur quadratic prefill cost and a key--value cache growing with context length. Research thus spans explicit-memory compression, sparse access, recurrent state construction, structured state dynamics, and heterogene...

Zhen-Tao Tan, Jing-Yi Shen, Yan-Bo Li et al. · 0 citations
#artificial intelligence Preprint Sep 2026

ALOE: Semantically Addressed Low-Rank Operators for Knowledge Editing

Knowledge editing changes what a model knows by modifying parameters so that a requested fact updates while unrelated behavior is preserved. This is usually treated as a write problem, but editing also involves an address problem: deciding which hidden states should receive the new residual. An update that activates to...

Zeyan Li, Hu Xu, Jian-Feng Xu · 0 citations
#artificial intelligence Preprint Sep 2026

From Parameters to Answers: How LLMs Retrieve and Use Their Internal Knowledge

This work compares country-continent questions with noun, adjective, and code answers while keeping several fitted measurements distinct across Qwen, Llama, and Gemma to separate early readability, natural strength, causal steering, and later content dependence.

Wen-Lin Wei, Yuan Fang, Ren-He Jiang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

RefCon: Iterative Refinement and Contrastive Memory Extraction for Context-Evolving Agent

Long-horizon agent interactions generate useful but noisy experience, and retraining models to absorb it is expensive. Context-evolving agents therefore need memory extraction methods that improve with more test-time compute without relying on gold labels. We propose RefCon, which combines sequential self-refinement wi...

Ubaidillah Ariq Prathama, Bo Liu, Yeonsun Hong et al. · 0 citations

Related blog posts

Microsoft Research Blog Sep 30, 2026

Forecasting space weather risks on power grids

Extreme space-weather events can damage power systems on Earth and degrade GPS accuracy and satellite operations. A new machine learning system can predict where damage is likely to occur 30-60 minutes before a storm arrives. The post Forecasting space weather risks on power grids appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.