Skip to content

Category

computer vision

3,022 papers

#artificial intelligence Preprint Open access Oct 2026

DeskForge: Dense Supervision from Desktop Environments for Computer-Use Agents

Computer-use agents need to reliably ground action targets in complex desktop scenes, where multiple applications, overlapping windows, and visually similar controls compete for attention. Existing training data rarely pair such scenes with dense annotations or vary them in a controlled way. We introduce DeskForge, a c...

A. Said Gurbuz, Ahmed Nassar, Sunghwan Hong et al. · 0 citations
#artificial intelligence Review Oct 2026

EditHero: A Benchmark for Long-Horizon Part-Level 3D Editing and Vibe Modeling

3D editing methods are usually tested on a single edit, yet an asset is built through a long sequence of revisions, each of which must implement the requested change while leaving everything else unchanged. We introduce EditHero, to our knowledge the first benchmark for long-horizon, part-level 3D editing, with natural...

Rui-Han Yu, Yu-Ju Tsai, Mu-Yao Niu et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Toward Controlling Biology with Language:Offline Learning of Prompt-Conditioned Interventions for Cells, Organoids, and Biobots

Artificial intelligence increasingly serves as a natural-language interface to complex technical systems, letting people accomplish sophisticated tasks by describing what they want rather than specifying how to do it. Extending this interface to living systems is harder: unlike code or images, a biological intervention...

Nam H. Le, Douglas Blackiston, Michael Levin et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Low-Cost Video--Time Priors as a Strong Baseline for EEG--fNIRS Emotion Regression on Familiar Videos

Continuous emotion regression estimates moment-to-moment valence and arousal while a viewer watches a video. In familiar-video deployment, responses fron training participant-specific estimate, and prior-dominating fixed fusion tests whether physiology adds residual correction. In five-fold subject-held-out evaluation...

Minghao Kong, Jiurun Chen, Ying Gao et al. · 0 citations
#artificial intelligence Preprint Oct 2026

World Action Modeling with Progressive Visual Planning

World action models (WAMs) have emerged as a promising paradigm for robotic control by jointly predicting future visual dynamics and actions from an initial observation and instruction. However, existing WAMs struggle with long-horizon prediction, as generating dense video rollouts is highly inefficient. Some recent WA...

Fei Zhang, Zhao-Chong An, Duncan P. Frost et al. · 0 citations
#computer vision Preprint Sep 2026

Paying for Too Many Tokens? Valid and Cost-Efficient Multimodal LLM Annotation with Simple Heuristics

Vision-Language Models (VLMs) enable video annotation at scale, but costs accumulate quickly: processing a typical 60-second short-form video at one frame per second requires millions of tokens. To reduce costs, researchers rely on heuristics such as sampling a subset of frames, compressing videos into image grids, or...

Zhi-Xiu Zhu, Kristina Gligoric · 0 citations
#computer vision Preprint Sep 2026

DramaAgent: Agentic Storytelling Video Generation

Recent diffusion and autoregressive models have substantially improved text-to-video generation, yet producing coherent long-form story videos with consistent characters and aligned audio remains challenging. Existing methods often suffer from narrative drift, unstable character identity, weak cross-scene continuity, a...

Ting Huang, Biao Wu, Rong-Hao Chen et al. · 0 citations
#computer vision Preprint Open access Oct 2026

Controllable Multi-label Video Safety Detection via Adaptive Tversky Policy Optimization

The rapid growth of video-based social media has increased users' exposure to harmful content, creating a need for reliable automated video safety detection. Although recent Vision-Language Models (VLMs) show strong video understanding capabilities, existing harmful video detection systems face two key limitations: the...

Guangyu Yang, Jingbiao Mei, Mingsheng Sun et al. · 0 citations
#machine learning Preprint Open access Oct 2026

More with Less: a Large Scale Remote Sensing VLM with a Simple Recipe

Remote sensing vision-language models are increasingly expected to support open-ended reasoning over Earth Observation data and a variety of tasks. Most recent progress in this area has been driven by remote-sensing-specific architectural designs, often introducing new encoders, alignment modules, or task-specific fusi...

Stefan Maria Ailuro (INSAIT, Sofia University "St. Kliment Ohridski"), Mario Markov (INSAIT et al. · 0 citations
#machine learning Preprint Open access Oct 2026

Compressing History into Memory: Distilling Transformers into Recurrent Transformers

Transformers are AI's workhorse but their computational cost becomes prohibitive when processing long sequences. We target long-horizon streaming vision and robotics applications, where it is particularly impractical to store and maintain a history of observations. Recurrent Transformers address this limitation by main...

Philippe Weinzaepfel, Christian Wolf, Mert B\"ulent Sariyildiz et al. · 0 citations
#machine learning Preprint Open access Oct 2026

Do Vision Transformers Need All-to-All Attention? Global Communication Through Elastic Learned Cores

Vision Transformers (ViTs) learn rich visual-semantic representations through all-to-all self-attention among patch tokens. However, this design implicitly assumes that direct pairwise patch interactions are necessary for effective representation learning. In this work, we challenge this assumption and show that repres...

Alan Z. Song, Yinjie Chen, Mu Nan et al. · 0 citations
#machine learning Preprint Open access Oct 2026

Augmented Equivariant Mesh Networks for Anatomical Segmentation

Anatomical mesh segmentation requires models that operate directly on irregular surface geometry while remaining robust to changes in coordinate pose across meshes of varying resolution. Existing task-specific mesh and point-cloud methods are not equivariant, and can degrade sharply under test-time perturbation, for ex...

Daniel Saragih · 0 citations

From tech blogs

See all →
Microsoft Research Blog Oct 6, 2026

What AI gets wrong and what failure teaches us

Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity.  The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.