Skip to content

Category

computer vision

3,022 papers

#machine learning Conference Open access Sep 2026

STEPS: Scene Text Editing with Preserved Style Using Diffusion and Contrastive Style Encoding

We introduce Scene Text Editing with Preserved Style (STEPS), a novel diffusion model architecture for quality text replacement in images. Scene Text Editing (STE), also known as Visual Text Editing, consists of changing the textual content in an image while conserving the original style, e.g. font, colors, orientation...

Nicolas Thiébaut, Nameer Hirschkind, Xiao Yu et al. · 0 citations
#machine learning Preprint Sep 2026

Correcting WHERE, Preserving HOW: Compositional Generalization for Vision-Language-Action Models via Referential Guidance

While Vision-Language-Action (VLA) models enable flexible action generation, their generalization across diverse environmental elements, including manipulated objects, destinations, and backgrounds, is limited by the lack of diversity in robotic training data. Trained end-to-end on such data, VLAs tend to exploit visua...

Yan-Yan Zhang, Di-Sheng Liu, Xin-Peng Li et al. · 0 citations
#machine learning Preprint Sep 2026

GazeFlow: From Human Gaze Behavior to Generative Egocentric Gaze Prediction

Egocentric gaze prediction enables many downstream applications but remains challenging, as human gaze is inherently stochastic. This stochasticity is constrained by structured temporal dynamics alternating between fixations and saccades, top-down influences from tasks, and bottom-up visual saliency. Based on this obse...

Sheng Zhao, Wei-Kai Lin, Yu-Hao Zhu · 0 citations
#machine learning Preprint Open access Oct 2026

Beyond Layers: Position-Resolved Gradient Conflict and Position-Aware Modulation for Unified Multimodal Models

Unified multimodal models (UMMs) train image understanding and autoregressive image generation on shared parameters, and the two objectives are known to interfere. Existing diagnoses and remedies operate at the resolution of layers or experts, measuring conflict per layer and resolving it by separating parameters. We a...

Shuyang Jiang, Fucheng Deng, Yuchuan Luo et al. · 0 citations
#machine learning Preprint Sep 2026

It Takes Little to Rewrite Perception: Targeted Semantic Substitution in Vision-Language Models at $\epsilon \leq 4/255$

Vision Language Models (VLMs) are widely deployed in safety-critical scenarios, and understanding to which extent they can be controlled by adversarial perturbation is a prerequisite for evaluating their trustworthiness. Existing representation-alignment attacks, which make a VLM perceive a target image, achieve limite...

Binchi Zhang, Atrisha Sarkar, Apurva Narayan · 0 citations
#machine learning Preprint Open access Oct 2026

Ranking-Aware Prompt Optimization for Multimodal Clinical Diagnosis

Multimodal large language models (MLLMs) are rapidly advancing clinical diagnosis, yet their adaptation pipelines remain anchored to accuracy-based objectives. Clinical data are heavily class-imbalanced: a constant-majority predictor can score above 90% accuracy while being clinically useless. We therefore evaluate and...

Tian Xia, Minghao Liu, Yiqing Liang et al. · 0 citations
#machine learning Preprint Open access Oct 2026

Prototype-Rule Neurosymbolic Regularization for Rank-Constrained Tensor Neural Networks under Label Scarcity

Rank-constrained tensor neural networks reduce the parameterization of high-order inputs, but they do not explicitly constrain class geometry in the learned representation. This study investigates whether a differentiable prototype-rule can provide a complementary inductive bias for Rank-R tensor learning under limited...

Eftychios Protopapadakis, Konstantinos Makantasis, Konstantinos M. Giannoutakis · 0 citations
#machine learning Preprint Open access Oct 2026

Reliability-Aware Checkpoint Selection for Domain Generalization

Checkpoint selection in domain generalization often relies on source-validation accuracy, yet the selected checkpoint need not provide reliable probabilities on unseen target domains. Source-target distribution shifts can alter accuracy rankings, while accuracy alone does not measure predictive probability quality. We...

Jinshi Liu, Jiahao Li, Pan Liu et al. · 0 citations
#machine learning Preprint Sep 2026

Revisiting On-policy Adversarial Black-Box Distillation: Calibrating Groupwise Reward Geometry for Effective Advantage Construction

Black-box distillation is a practical route for transferring capabilities from API-accessible large language models that expose only text outputs into smaller student models. Recent on-policy adversarial methods such as GAD improve over SeqKD by forming an adversarial loop between a critic and a student, where the crit...

Xiao Cui, Mo Zhu, Yu-Lei Qin et al. · 1 citation
#machine learning Preprint Open access Oct 2026

Mutual Equilibrium: Multimodal Representation Learning through Reciprocal Feedback

This work proposes a mutual feedback architecture, MEQ, that refines the two inputs, of possibly different modalities, into a pair of coupled embeddings such that each embedding reflects the information of the other. The core idea is to incorporate continuous interchange of information between the two inputs. This idea...

Ho-min Park, Byungkon Kang · 0 citations
#machine learning Preprint Sep 2026

Anchoring Adversarial Trajectories to Data Manifolds: A Bilevel Transfer Optimization Framework

Manifold Anchored Bilevel Transfer (MABT), a unified framework that anchors adversarial trajectories to the shared semantic subspace, is proposed, and a Hessian-free solver with linear-time complexity is developed to handle the resulting hierarchy.

Yao-Hua Liu, Yi-Fan Guo, Jia-Xin Gao · 0 citations
#machine learning Preprint Open access Oct 2026

World-as-Graph: Relational World Modeling Through Latent Space Graphs

World models aim to learn representations of real-world environments and predict their future evolution. Recent object-centric world models have made expressive progress by representing visual scenes as sets of object-level latent states, but object-object relations are often captured only implicitly, which limits expl...

Yaqi Yang, Shuo Huang, Yujin Huang et al. · 0 citations

From tech blogs

See all →
Microsoft Research Blog Oct 6, 2026

What AI gets wrong and what failure teaches us

Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity.  The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.