Skip to content

Category

computer vision

3,022 papers

#artificial intelligence Preprint Sep 2026

Just MLPs: Efficient Visual State Reconstruction for Multimodal Language Models

This work proposes $\delta$-Vision, which replaces repeated Transformer evolution of visual tokens with lightweight low-rank adapters that construct layer-wise visual memories while preserving all visual tokens for text retrieval, and achieves higher accuracy than visual token pruning baselines at comparable or lower c...

Jing-Di Lei, Junxian Li, Di Zhang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Don't Throw Away the Tail: Action Upcycling for Policy Acceleration

This work proposes Action Upcycling, a training-free algorithm that reuses actions the policy would otherwise discard, without accessing model internals or drawing extra samples, and finds that discarded actions stay close to their replanned versions as long as the action velocity remains smooth.

Taesung Kwon, Jangho Park, Sunwoo Park et al. · 0 citations
#artificial intelligence Preprint Sep 2026

EviSplat: Preserving Multi-View Evidence in 3D Gaussian Splatting for Open-Vocabulary Segmentation

EviSplat is introduced, which preserves individual observation features as evidence for later text queries within class-agnostic 3D instances that represent objects, object parts, or background regions and learns a distribution describing which visual appearances its observations support.

Sung-Ho Moon, Kota Shimomura, Junwoo Park et al. · 0 citations
#artificial intelligence Preprint Sep 2026

ESTHER: Egocentric Stereo Hand Estimation and Reconstruction in the Wild

Human dexterity is guided by two eyes watching two hands: binocular vision supplies the metric 3D structure that fine-grained manipulation consumes. Egocentric stereo is therefore the natural perceptual interface for robots, AR, and VR-yet metric 3D hand reconstruction from this very signal still has neither an end-to-...

Hong-Yu Ma, Hai-Rong Qu, Shi-Qi Zhao et al. · 0 citations
#artificial intelligence Preprint Sep 2026

CoHuB: A Simulation Benchmark for Multi-Humanoid Collaboration

This work introduces CoHuB (Collaborative Multi-Humanoid Benchmark), a simulation benchmark for multi-humanoid collaboration under egocentric visual observations and provides a foundation for developing and evaluating multi-humanoid collaboration policies.

Hyunji Park, Jebeom Chae, Minwoo Park et al. · 0 citations
#artificial intelligence Preprint Sep 2026

CoDrive: Cross-Vehicle World-Consistent Video Generation with Precise Trajectory Control for Cooperative Driving

Real-world driving is inherently multi-agent, yet most existing driving world models generate observations from a single ego vehicle. Independently extending them to multiple vehicles does not ensure that different agents observe a consistent shared world. We present CoDrive, a cross-vehicle, multi-view driving video g...

Yu Meng, Bai-Ning Zhao, Jun-Tao Wu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Triangular Resampling for Long-Horizon Motion Generation

Triangular Resampling is introduced, a post-training method for mitigating long-horizon error accumulation in motion diffusion models that addresses the mismatch between ground-truth-derived training windows and model-generated inference states.

Kun-Hang Li, Yi-Yi Cai, Xiang-Yue Zhang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Learning What to Recall: Adaptive Multi-Cue Episodic Memory for World Models

World models predict future observations from current experience and actions, yet prediction can depend on observations seen far in the past. Episodic memory preserves past observations for later recall; however, as memory accumulates, it raises a fundamental question: which memories are useful for the current predicti...

Beomsu Kim, Chieh-Hsin Lai, Bac Nguyen et al. · 0 citations
#artificial intelligence Preprint Open access Sep 2026

Evidence-Aligned Multimodal On-Policy Self-Distillation for Fine-Grained Visual Understanding

Fine-grained visual understanding requires models to recognize small details within complex images. Multimodal on-policy self-distillation (OPSD) addresses this challenge by using a teacher conditioned on evidence-centered crops to supervise a student conditioned on original images along student-generated trajectories....

Nanxing Hu, Qiwei Yan, Jinchao Zhang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Summarize Before Grounding: Query-Guided Chunk Condensation for Long-Video Temporal Grounding

This paper proposes a ``summarize before grounding''framework (named ``SumGround'') for long-video temporal grounding, and introduces query-guided latent summaries, which is represented as KV states of query-guided prompts, to compress redundant visual tokens into compact query-relevant chunk summaries.

Nan-Xing Hu, Xiao-Yue Duan, Qi-Wei Yan et al. · 0 citations

From tech blogs

See all →
Microsoft Research Blog Oct 6, 2026

What AI gets wrong and what failure teaches us

Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity.  The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.