This report presents PolyOCR, a family of unified OCR foundation models of varying scales that combines a shared instruction-following framework with a large-scale data engine that converts heterogeneous visual resources into quality-verified OCR supervision.
Guang-Zhan Huang, Yong-Shuo Zhang, Bing-Tao Fu et al.· 0 citations
Recent multimodal image generation models can take multiple images and textual instructions as input, enabling reference-based generation guided not only by text but also by visual instructions such as layouts, arrows, and pose cues. However, existing benchmarks do not evaluate the joint setting in which multiple refer...
Yuta Oshima, Masakazu Yoshimura, Masahiro Suzuki et al.· 0 citations
Federated knowledge must remain usable by recipients with different modalities, private architectures, and tasks. We present FedSocket, which makes recipient execution a design requirement of the exchanged model. A shared Q combines recipient-computable inputs, task-owned outputs, and ownership-aware aggregation, conne...
TReVS is proposed, a training-free framework that combines textual relevance with vision-encoder saliency for pre-LLM pruning and leverages high-variance attention heads to remove task-irrelevant tokens at shallow-to-intermediate layers of the LLM.
Jing Wang, Zhi-Ping Wu, Dong-Dong Ren et al.· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
The analyses show how VLM evaluation behaves as model panels grow and evolve, providing guidance for designing future benchmarks that are more efficient, robust to model turnover, and explicit about the limits of evaluation-side pruning.
Modern vision-language models (VLMs) have shown promising results in long-video understanding due to the rich semantic information they can capture. However, most methods focus on coarse captioning of extracted image frames that are computationally inefficient and require models with large context windows. While past w...
Arka Mukherjee, Kaleen Shrestha, Larissa Zhu et al.· 0 citations
Coding agents can now write, run, and debug programs with little human help. Robot tasks, however, are usually specified by a sentence that leaves out how to grasp, in what order to make contact, and what the result should look like, and an agent given only the sentence must find these details by trial and error. We in...
Yi-Fan Kang, Zihan Wang, Zhi-Wen Fan et al.· 0 citations
Multi-step visual retrieval-augmented generation (RAG) answers complex questions by repeatedly retrieving visual evidence, updating an intermediate state, and deciding whether to continue searching or answer. Yet retrieving relevant evidence does not ensure its effective use throughout the reasoning trajectory. As mult...
Yalun Wu, Bingzhou Wang, Boyang Wang et al.· 0 citations
Image translation is a fundamental capability of multimodal models for multilingual applications, requiring visual understanding and meaning preservation across languages. However, existing benchmarks have limited language coverage and often lack explicit image-specific evaluation criteria, making it difficult to compr...
Affordance perception aims to localize actionable regions supporting embodied interaction, yet 2D and 3D affordance grounding have evolved as separate problems, with different task definitions, supervision formats, datasets, and evaluation protocols. This fragmentation limits the learning of transferable object-afforda...
Yu-Hao Liu, Yi-Ming Zhong, Han-Qing Wang et al.· 0 citations
P predictive visual latents are established as a foundation for effective WAM learning from task-specific demonstrations and for transferring future-modeling knowledge acquired from broader in-the-wild videos.
Yang Zhang, Jiang-Yuan Zhao, Chen-You Fan et al.· 0 citations
Exploring how generative AI could make machine vision more accessible to businesses. The post GenEye in a Box: Making Machine Vision Something You Can Just Ask For appeared first on GPT-Lab.
Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity. The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.
Computer scientist, entrepreneur, and philanthropist will collaborate with the MIT Schwarzman College of Computing to advance AI and scientific discovery.