Long-context reasoning in large language models incurs substantial computation and memory costs. Visual text compression (VTC) reduces input length by rendering text as images, but fixed-resolution rendering creates a compression-performance trade-off: low DPI saves tokens at the expense of legibility, whereas high DPI...
Fang-Zhi Zhong, Xue-Rui Qiu, Yu-Qi Pan et al.· 0 citations
Symbolic state supervision raises the 20M model's frame accuracy from 31.1% to 67.3% at the same training-data budget, suggesting that learning representations of state changes can complement scaling.
Wei-Hang Guo, Xiao-Yu Wu, Yi-Fei Wang et al.· 0 citations
While Multimodal Large Language Models (MLLMs) are increasingly deployed in safety-critical domains, their reliability is threatened by multimodal implicit risks. Unlike explicit threats, these hazards emerge when individually benign text and neutral visual entities logically converge to induce unsafe outputs. Current...
Ruo-Chen Zhang, Yao Huang, Yi-Tong Sun et al.· 0 citations
FM-ReID is proposed, an end-to-end framework that formulates local representation learning as selective competitive token routing as an effective way to augment holistic foundation-model representations, supporting competitive token routing as an effective way to augment holistic foundation-model representations.
Zhi-Qiang Li, Xiao-Wei Zhou, Ze-Yuan Sun et al.· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
A mechanistic analysis of the VLM model's internal attention patterns shows that a simple describe-then-decide prompting strategy increases vision attention by 30-40% during generation, and task-specific fine-tuning improves dermatology classification but reduces cross-domain medical question-answering performance in t...
Janet Wang, Yun-Bei Zhang, Xiao Wang et al.· 0 citations
When evaluated on unseen species, the best adapted VLMs outperformed specialist models trained on the same data in identifying merge errors and LoRA on a few thousand labels brought open models level with specialist models.
Yi-Cong Li, Jun-Jie Wang, Leander Lauenburg et al.· 0 citations
Cross-modal associations are systematic pairings of features across modalities, such as the association of'bouba'with round shapes and'kiki'with sharp shapes. Prior work has compared humans and vision-language models (VLMs) on such associations, but often using different stimuli or tasks between humans and models. Here...
Su-Min Hong, Katsumi Ibaraki, Renee Shi et al.· 0 citations
STAIRCASE POLICY is introduced, a streaming inference and training framework that turns a flow-matching VLA into a JEPA-style WAM and partitions a large action chunk into sub-chunks at staggered denoising stages, enabling long-horizon execution without repeated full policy inference.
Guoheng Sun, Chen Chen, Jin Wang et al.· 0 citations
This work introduces Online VIL (Online Versatile Incremental Learning), a novel scenario where class concepts and visual domains evolve simultaneously online without explicit boundaries, and proposes a novel framework TopFlow, Topology preservation with Flow matching representation that contains two complementary mech...
Jae-Ho Lee, Minji Park, Jun-Yeong Moon et al.· 0 citations
FineART is presented, a densely annotated bimanual manipulation dataset comprising 40,543 episodes and 533,913 subtasks across 151 tasks and FineART-VLA is introduced, a vision-language-action policy that predicts its own next subtask to guide its actions.
Jade Choghari, Pepijn Kooijmans, Mansi Agarwal et al.· 0 citations
Large foundation models have been introduced with the promise of efficient adaptation to downstream tasks. Yet, under limited supervision, MLLMs, an important class of large foundation models, remain challenging to adapt to various downstream tasks. Adaptation typically relies either on MLLM parameter fine-tuning or on...
Konstantinos D. Polyzos, Eleni Oikonomou, Tara Javidi· 0 citations
Standard video generators do not natively compact historical context into reusable memory tokens. As generation continues, the growing history makes it increasingly difficult to retain information from earlier frames due to long-context degradation. Key-frame-based approaches address this challenge by retaining selecte...
Xiao-Yu Wu, Wei-Hang Guo, Yi-Fei Wang et al.· 0 citations
Exploring how generative AI could make machine vision more accessible to businesses. The post GenEye in a Box: Making Machine Vision Something You Can Just Ask For appeared first on GPT-Lab.
Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity. The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.
Computer scientist, entrepreneur, and philanthropist will collaborate with the MIT Schwarzman College of Computing to advance AI and scientific discovery.