Skip to content

Category

computer vision

3,022 papers

#machine learning Preprint Open access Sep 2026

From internal representations to model improvement through prediction errors

With limited annotation budgets, choosing which images to label determines how much a model improves. Data-selection methods that use features from a separately trained model, or scene descriptions written by vision-language models, have been successful, but those signals do not directly capture changes in the model be...

Yushi Nakaya, Kenichi Higuchi, Shuichi Ishida · 0 citations
#machine learning Preprint Sep 2026

Scaffold Then Internalize: Representation Injection for Diffusion Transformers

Recent representation alignment (REPA) methods accelerate diffusion transformer training by aligning projections of the transformer's hidden states with representations from pretrained visual encoders. In this work, we explore a reverse and complementary direction to REPA: rather than projecting diffusion representatio...

Han-Si Fu, Jia-Cheng Chen, Bao-Quan Zhao et al. · 0 citations
#machine learning Preprint Open access Sep 2026

AnswerMap: Faithful Spatial Interpretability of VLMs from Answer Posteriors

When a VLM answers a visual query, current interpretability tools rely on text rationales, which use a mismatched modality, or on internal read-outs, which originate too early to reflect the final output and require white-box access to the model. We introduce AnswerMap, a training-free, task-agnostic, black-box visual...

Mohamed Eltahir, Fardows Adam, Duaa M. Tahir et al. · 0 citations
#machine learning Preprint Sep 2026

Beyond Selection: Token Parameterization for Extreme Visual Token Compression

Braco is a lightweight four-step coder that combines transform-basis truncation, input-independent basis-coordinate embeddings, budget-dependent orthogonal re-parameterization, and learned spatial residual tokens from lightweight pooling that forms the favorable empirical accuracy-efficiency frontier under compression.

Rui-Lian Zhong, Yu Li, Zhe-Yu Yan et al. · 0 citations
#machine learning Preprint Open access Sep 2026

G$^3$-LoRA: Organizing Reward-Weighted Video Data with Gradient-Guided Grouped LoRA

Post-training foundation video models on heterogeneous reward-weighted data usually assume that all data categories induce compatible updates. This assumption is fragile when categories correspond to different skills, domains, or evaluation dimensions. We study this problem in text-to-video post-training, where VBench2...

Jia Song (The Hong Kong University of Science and Technology), Wenhow Li (The Hong Kong University of Science and Technology), Lichen Bai (The Hong Kong University of Science and Technology) et al. · 0 citations
#machine learning Preprint Open access Sep 2026

DF-CBM: Region-Aware Concept Bottleneck Models for Deepfake Detection

Deepfake detection methods have become increasingly effective yet most provide limited insight into the evidence behind their predictions. However, in forensic settings users also need to know which manipulation cues support the decision and where they appear. Existing explainability methods only partially address this...

Georgios Tsoumplekas, Vazgken Vanian, Alexandros Doumanoglou et al. · 0 citations
#machine learning Preprint Sep 2026

From Perception to Integration: Revisiting the Internal Dynamics of Reasoning in Vision-Language Models

Vision-language models (VLMs) can answer simple visual questions, but often struggle when one question requires several visual judgments. We study this gap with controlled tasks for feature binding, numerosity, spatial relations, and amodal completion, together with a Composite task that combines them. Matched counterf...

Rong-Yu Xu, Prayag Tiwari, Shao-Lei Zhang · 0 citations
#machine learning Preprint Sep 2026

Beyond Reconstruction Loss in Post-Training Quantization: Balanced Fitting for Large Vision-Language Models

Experiments on multiple LVLMs show that the Balanced Fitting method consistently outperforms prior PTQ approaches under both weight-only and weight-activation quantization, while lower reconstruction loss does not reliably translate into better downstream performance.

Min-Chan Kang, Kyeonghye Park, Seungyeon Sa et al. · 0 citations
#machine learning Preprint Sep 2026

Do Emotion Concepts Generalize Across Sources, Modalities, and Architectures in Vision-Language Models?

CMES (Cross-Modal Emotion Stimuli), a multi-source collection of emotion-conditioned stories, real facial expressions, synthetic portraits, and synthetic emotion-evoking scenes, suggests that emotion representations can share relational structure and causal effects across sources, modalities, and architectures, even wh...

Bo-Hao Xing, Xin Liu, Kai-Shen Yuan et al. · 0 citations
#machine learning Preprint Open access Sep 2026

Natural State-Prediction Accuracy can Hide Weak Controlled Responsiveness in VLA Readouts

Accurately decoding object states from the internal representations of vision-language-action (VLA) models does not establish that the predictions respond faithfully to changes in the target physical state. In natural observations, object state, robot configuration, occlusion, and task progress vary together, allowing...

Hyungjoon Kim, Wonbin Son, Mi Young Lee et al. · 0 citations
#machine learning Preprint Sep 2026

Rethinking Latent Visual Reasoning: Grounding Latent Reasoning in Visual Evidence

ReaLVR is proposed, which brings visual-evidence supervision to the model's own free-running latent trajectories, and is the first to scale visual reasoning in latent space, showing that the framework continues to deliver robust improvements at frontier model scales up to 235B.

Xi Xiao, Tian-Chen Zhao, Youngeun Kim et al. · 0 citations

From tech blogs

See all →
Microsoft Research Blog Oct 6, 2026

What AI gets wrong and what failure teaches us

Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity.  The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.