This work presents, to their knowledge, the first simultaneous sign-to-sign (S2S) translation system, with two wait-k regimes: test-time wait-k inference applied directly to a full-sentence model, and a trained wait-k model via stochastic multi-path supervision.
Ze-Tian Wu, Bo-Wen Xie, Stefan Lee et al.· 0 citations
Multi-vector retrievers built on vision-language models lead visual document retrieval (VDR), but they run a multi-billion-parameter query encoder on every search. Distilling this encoder into a small student that queries the teacher's existing index would remove the bottleneck. The standard recipe, however, matches th...
Zhu Liu, Zi-Yi Wang, Yao Zhang et al.· 0 citations
Multimodal foundation models edit natural photographs at production quality, yet the same models struggle with structured visual content such as infographics. Unlike photographs, infographics encode information through logical relations; editing one element often requires surrounding elements to be adapted. We refer to...
Cheng Yang, Chu-Fan Shi, Hui-Juan Wang et al.· 0 citations
The upshot is that a large share of the apparent CAAP--SNAP gap traces back to closed-category annotation limits rather than a real grounding failure, which matters for how the authors detect hallucination, analyze failure modes, and design evaluation for grounded multimodal systems meant to work in the open world.
Yu-Feng Yen· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
Byte-Pair Encoding (BPE) constructs subword vocabularies through greedy pair merging, but conventional BPE requires the number of merges or target vocabulary size to be specified externally. Significance-Gain Pair Encoding (SG-BPE) replaces frequency-only selection with a statistical criterion based on how strongly an...
High instance-level accuracy can mask inconsistencies in spatial reasoning when objects exchange positions or their roles are reversed in the query. The internal representations supporting relative-position reasoning remain poorly understood. We investigate two complementary components of this process: tracking object...
Yingjin Song, Denis Paperno, Albert Gatt· 0 citations
Sign language translation and generation share the goal of bidirectional alignment between text and sign representations. However, existing approaches either treat them as isolated tasks or are only verified on limited datasets, limiting effective modeling between modalities. In this paper, we propose SignFLIP, a unifi...
Zhaoyi An, Si-Han Tan, Youngbae Hwang et al.· 0 citations
This work introduces InfiMed2, a family of 4B and 27B generalist medical multimodal foundation models built around stage-aware data design, and curates a 55.68B-token corpus that combines broad clinical knowledge with context-rich biomedical visual evidence through source-specific processing.
Guang-Hao Zhu, Ze-Yu Liu, Zhitian Hou et al.· 0 citations
Medical images are released with the reports that describe them, and protecting the image does not protect the report. This paper measures the text component of such releases. We measure identifier detection, downstream utility and residual identity leakage on the same documents, with the pseudonymisation policy as the...
Andreas K. Maier, Monica Hinrichs-Mayer, Franziska Weber et al.· 0 citations
The ability of vision-language models (VLMs) to abstain from unanswerable questions is as important as their ability to answer answerable ones accurately. Recently, several benchmarks have emerged to evaluate and improve VLM abstention, but they have substantial limitations. First, samples often contain shortcut cues i...
Latent medical image generators usually treat the tokenizer as fixed preprocessing. We test whether this separation holds in a controlled ChestMNIST study at $64\times64$ that crosses discrete tokenizers, generator families, and sampler settings under a shared latent grid, with continuous-latent reference cells. The ra...
Multimodal contrastive learning has enabled zero-shot visual classification by aligning images with textual categories. However, in hierarchically structured label spaces, existing methods often produce predictions that are inconsistent across taxonomic levels. For example, a model may predict a fine-grained category w...
Zhiyuan Tao, Srikumar Sastry, Matthew J Thompson et al.· 0 citations
Exploring how generative AI could make machine vision more accessible to businesses. The post GenEye in a Box: Making Machine Vision Something You Can Just Ask For appeared first on GPT-Lab.
Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity. The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.
Computer scientist, entrepreneur, and philanthropist will collaborate with the MIT Schwarzman College of Computing to advance AI and scientific discovery.