This work proposes $\delta$-Vision, which replaces repeated Transformer evolution of visual tokens with lightweight low-rank adapters that construct layer-wise visual memories while preserving all visual tokens for text retrieval, and achieves higher accuracy than visual token pruning baselines at comparable or lower c...
Jing-Di Lei, Junxian Li, Di Zhang et al.· 0 citations
This work proposes Action Upcycling, a training-free algorithm that reuses actions the policy would otherwise discard, without accessing model internals or drawing extra samples, and finds that discarded actions stay close to their replanned versions as long as the action velocity remains smooth.
Taesung Kwon, Jangho Park, Sunwoo Park et al.· 0 citations
EviSplat is introduced, which preserves individual observation features as evidence for later text queries within class-agnostic 3D instances that represent objects, object parts, or background regions and learns a distribution describing which visual appearances its observations support.
Sung-Ho Moon, Kota Shimomura, Junwoo Park et al.· 0 citations
Human dexterity is guided by two eyes watching two hands: binocular vision supplies the metric 3D structure that fine-grained manipulation consumes. Egocentric stereo is therefore the natural perceptual interface for robots, AR, and VR-yet metric 3D hand reconstruction from this very signal still has neither an end-to-...
Hong-Yu Ma, Hai-Rong Qu, Shi-Qi Zhao et al.· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
This work introduces CoHuB (Collaborative Multi-Humanoid Benchmark), a simulation benchmark for multi-humanoid collaboration under egocentric visual observations and provides a foundation for developing and evaluating multi-humanoid collaboration policies.
Hyunji Park, Jebeom Chae, Minwoo Park et al.· 0 citations
These results show that reliable scene-text recognition requires VLMs to balance visual character evidence with contextual information, preserving clear text while using context mainly when the visual evidence is uncertain.
This work revisits model completeness: how well the concepts can reproduce the model's outputs, and shows that model incompleteness of the concepts can be bounded by the autoencoder's reconstruction error.
Vojtech Kur, A. Kukucka, T. Brázdil et al.· 0 citations
Real-world driving is inherently multi-agent, yet most existing driving world models generate observations from a single ego vehicle. Independently extending them to multiple vehicles does not ensure that different agents observe a consistent shared world. We present CoDrive, a cross-vehicle, multi-view driving video g...
Yu Meng, Bai-Ning Zhao, Jun-Tao Wu et al.· 0 citations
Triangular Resampling is introduced, a post-training method for mitigating long-horizon error accumulation in motion diffusion models that addresses the mismatch between ground-truth-derived training windows and model-generated inference states.
Kun-Hang Li, Yi-Yi Cai, Xiang-Yue Zhang et al.· 0 citations
World models predict future observations from current experience and actions, yet prediction can depend on observations seen far in the past. Episodic memory preserves past observations for later recall; however, as memory accumulates, it raises a fundamental question: which memories are useful for the current predicti...
Beomsu Kim, Chieh-Hsin Lai, Bac Nguyen et al.· 0 citations
Fine-grained visual understanding requires models to recognize small details within complex images. Multimodal on-policy self-distillation (OPSD) addresses this challenge by using a teacher conditioned on evidence-centered crops to supervise a student conditioned on original images along student-generated trajectories....
Nanxing Hu, Qiwei Yan, Jinchao Zhang et al.· 0 citations
This paper proposes a ``summarize before grounding''framework (named ``SumGround'') for long-video temporal grounding, and introduces query-guided latent summaries, which is represented as KV states of query-guided prompts, to compress redundant visual tokens into compact query-relevant chunk summaries.
Nan-Xing Hu, Xiao-Yue Duan, Qi-Wei Yan et al.· 0 citations
Exploring how generative AI could make machine vision more accessible to businesses. The post GenEye in a Box: Making Machine Vision Something You Can Just Ask For appeared first on GPT-Lab.
Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity. The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.
Computer scientist, entrepreneur, and philanthropist will collaborate with the MIT Schwarzman College of Computing to advance AI and scientific discovery.