Skip to content

Category

computer vision

3,022 papers

#computer vision Preprint Sep 2026

Simultaneous Translation between Sign Languages

This work presents, to their knowledge, the first simultaneous sign-to-sign (S2S) translation system, with two wait-k regimes: test-time wait-k inference applied directly to a full-sentence model, and a trained wait-k model via stochastic multi-path supervision.

Ze-Tian Wu, Bo-Wen Xie, Stefan Lee et al. · 0 citations
#computer vision Preprint Sep 2026

ColNanoVDR: Document-Free Query Distillation for Multi-Vector Visual Document Retrieval via Optimal Transport

Multi-vector retrievers built on vision-language models lead visual document retrieval (VDR), but they run a multi-billion-parameter query encoder on every search. Distilling this encoder into a small student that queries the teacher's existing index would remove the bottleneck. The standard recipe, however, matches th...

Zhu Liu, Zi-Yi Wang, Yao Zhang et al. · 0 citations
#computer vision Preprint Sep 2026

InfoEdit: Probing Global Layout Reasoning in Infographic Editing

Multimodal foundation models edit natural photographs at production quality, yet the same models struggle with structured visual content such as infographics. Unlike photographs, infographics encode information through logical relations; editing one element often requires surrounding elements to be adapted. We refer to...

Cheng Yang, Chu-Fan Shi, Hui-Juan Wang et al. · 0 citations
#computer vision Preprint Sep 2026

Attribution Gaps in Zero-Training LLM+OVOD Pipelines: A Fine-Grained Analysis of the CAAP--SNAP Discrepancy

The upshot is that a large share of the apparent CAAP--SNAP gap traces back to closed-category annotation limits rather than a real grounding failure, which matters for how the authors detect hallucination, analyze failure modes, and design evaluation for grounded multimodal systems meant to work in the open world.

Yu-Feng Yen · 0 citations
#computer vision Preprint Open access Sep 2026

MDL-Calibrated Significance-Gain Pair Encoding: Replication-Aware Automatic Stopping for Subword Tokenization

Byte-Pair Encoding (BPE) constructs subword vocabularies through greedy pair merging, but conventional BPE requires the number of merges or target vocabulary size to be specified externally. Significance-Gain Pair Encoding (SG-BPE) replaces frequency-only selection with a statistical criterion based on how strongly an...

Azam Nouri · 0 citations
#computer vision Preprint Open access Sep 2026

Who Is Left of Whom? Tracing Spatial Evidence and Role Binding in Relative-Position Reasoning

High instance-level accuracy can mask inconsistencies in spatial reasoning when objects exchange positions or their roles are reversed in the query. The internal representations supporting relative-position reasoning remain poorly understood. We investigate two complementary components of this process: tracking object...

Yingjin Song, Denis Paperno, Albert Gatt · 0 citations
#computer vision Preprint Sep 2026

SignFLIP: A Unified Model for Sign Language Translation and Generation via Stage-wise Alignment at Scale

Sign language translation and generation share the goal of bidirectional alignment between text and sign representations. However, existing approaches either treat them as isolated tasks or are only verified on limited datasets, limiting effective modeling between modalities. In this paper, we propose SignFLIP, a unifi...

Zhaoyi An, Si-Han Tan, Youngbae Hwang et al. · 0 citations
#computer vision Preprint Sep 2026

InfiMed2: A Generalist Medical Multimodal Foundation Model from Contextual Evidence and Stability-Aware Supervision

This work introduces InfiMed2, a family of 4B and 27B generalist medical multimodal foundation models built around stage-aware data design, and curates a 55.68B-token corpus that combines broad clinical knowledge with context-rich biomedical visual evidence through source-specific processing.

Guang-Hao Zhu, Ze-Yu Liu, Zhitian Hou et al. · 0 citations
#computer vision Preprint Sep 2026

The Text Beside the Image: Detection, Utility and Leakage for Trustworthy Multimodal Medical Data and Beyond

Medical images are released with the reports that describe them, and protecting the image does not protect the report. This paper measures the text component of such releases. We measure identifier detection, downstream utility and residual identity leakage on the same documents, with the pseudonymisation policy as the...

Andreas K. Maier, Monica Hinrichs-Mayer, Franziska Weber et al. · 0 citations
#computer vision Preprint Sep 2026

From Knowing to Abstaining: Bridging the Representation-Action Gap in Vision-Language Models

The ability of vision-language models (VLMs) to abstain from unanswerable questions is as important as their ability to answer answerable ones accurately. Recently, several benchmarks have emerged to evaluate and improve VLM abstention, but they have substantial limitations. First, samples often contain shortcut cues i...

Jia-Luo He, Huang-Xun Chen · 0 citations
#machine learning Preprint Open access Sep 2026

Tokenizer-Generator Coupling in Medical Image Generation

Latent medical image generators usually treat the tokenizer as fixed preprocessing. We test whether this separation holds in a controlled ChestMNIST study at $64\times64$ that crosses discrete tokenizers, generator families, and sampler settings under a shared latent grid, with continuous-latent reference cells. The ra...

Liam Chalcroft · 0 citations
#machine learning Preprint Open access Sep 2026

Beyond Flat Labels: Level-Restricted Contrastive Learning for Hierarchical Fine-Grained Vision Classification

Multimodal contrastive learning has enabled zero-shot visual classification by aligning images with textual categories. However, in hierarchically structured label spaces, existing methods often produce predictions that are inconsistent across taxonomic levels. For example, a model may predict a fine-grained category w...

Zhiyuan Tao, Srikumar Sastry, Matthew J Thompson et al. · 0 citations

From tech blogs

See all →
Microsoft Research Blog Oct 6, 2026

What AI gets wrong and what failure teaches us

Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity.  The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.