Skip to content

Category

computer vision

3,022 papers

#artificial intelligence Preprint Oct 2026

Joint Class-Time Learning for Video Classification with Multi-Instance Partial-Label Learning

Multi-instance partial-label learning (MIPL) addresses inexact supervision in both the instance and label spaces, which can be applied to video classification. However, bag-level labels do not explicitly supervise the correspondence between candidate classes and temporal evidence. We propose {\ours}, which couples labe...

Ling-Yu Shen, Wei Tang, Fakhri Karray et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Vision Transformer Ensembles for Panoramic Street Segmentation

Semantic segmentation of street panoramas can support detailed descriptions of urban environments, yet small datasets and unequal training costs make model selection difficult. This paper presents the system used for a first place submission to the PalmCity challenge in the leaderboard snapshot dated 5 October 2026. Ni...

Yunus Serhat B{\i}\c{c}ak\c{c}{\i} · 0 citations
#artificial intelligence Preprint Oct 2026

ROT: Rotating Hidden States towards Contextual Vectors for Hallucination Mitigation in LVLMs

Large Vision-Language Models (LVLMs) frequently suffer from object hallucination. Existing training-free interventions primarily manipulate attention weights, which indirectly affect the deep semantics reaching the final predictive layers. In this work, we shift our focus to the hidden state vectors extracted after sel...

Yi-Jin Du, Xiang-Cheng Zhan, Shuo Yang · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Local2Mesh: Spatially Localized Contour-to-Mesh for Left Ventricular Reconstruction from Sparse 2D Cardiac MRI

Three-dimensional (3D) left ventricular (LV) reconstruction from sparse cardiac magnetic resonance (CMR) imaging remains challenging due to inter-slice misalignment and insufficient local spatial information between slices. Global aggregation of contour features may obscure local contour-to-surface relationships. We pr...

Haoyu Wu, Ling Lin, Pascal Lef\`evre et al. · 0 citations
#artificial intelligence Preprint Oct 2026

Representation Disentanglement for Fair Chest X-Ray Diagnosis

Deep learning has advanced chest X-ray (CXR) diagnosis, yet demographic biases in learned representations may contribute to performance disparities across intersectional groups. We propose a single-encoder framework combining dual-level decorrelation with prototype-guided cross-group contrastive learning to reduce demo...

Yu-Jie Sun, Rui-Zhe Li, Xiao-Wu Sun · 0 citations
#artificial intelligence Preprint Oct 2026

Scalable Minimal-Change Learning for Controllable Image Editing

Image editing should change only the attributes specified by an instruction while preserving everything else, yet current methods often make unintended changes. We treat this minimal-change principle as an optimization objective for instruction-based editing. Latent L1 regularization is a poor proxy for output locality...

Shuo Chen, Feng-Ming Huang, Yu Yao et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Investigating Query-Insensitive Behavior in Spatio-Temporal Video Grounding

Spatio-temporal video grounding (STVG) aims to localize objects or events described by natural language queries in both space and time. Existing STVG models are typically trained and evaluated under the assumption that each query is relevant to the input video. In this work, we challenge this assumption by studying the...

Eryk Ko{\l}odziejczyk, Alberto Presta, Karol Szurkowski et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

JLD: Perceptual Distance Through A Jacobian Lens

Image compression, restoration, and generation all require a way to measure how different two images look to a person. Pixel error ignores how people see, while the most accurate perceptual distances are typically fitted to human judgments, tying them to a fixed data and resolution. For example, when image resolution i...

Shreshth Saini, Balu Adsumilli, Alan C. Bovik · 0 citations
#artificial intelligence Preprint Oct 2026

ReMem: Streaming Video Understanding With Long Context Retention

Despite their impressive performance on a wide range of video understanding tasks, current Vision Language Models (VLMs) are predominantly designed for offline scenarios and struggle to handle online streaming videos that demand low latency response. Several studies have explored memory and token compression strategies...

Yi-Heng Li, Xu He, Shao-Bo Wang et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

UltraDub: Towards Authentic Dubbing by Unifying Visually-Steered Flow Learning and Trajectory Guidance

Visual voice cloning requires intelligible, speaker-consistent speech synchronized with visible articulation. However, sequential multimodal conditioning can disrupt previously established temporal and speaker cues, while imbalanced inference guidance can improve linguistic accuracy at the expense of lip synchronizatio...

Gaoxiang Cong, Liang Li, Jianwei Wen et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

On Hyperparameter Tuning on the Test Set

"Don't tune hyperparameters on the test set" is often stated in machine learning textbooks. Violating it is considered a cardinal sin that produces misleadingly optimistic results, corrupts benchmark integrity, and thus can even be interpreted as scientific fraud. Yet evidence suggests that test set hyperparameter tuni...

Matteo Fregonara, Tom Viering, Jan van Gemert · 0 citations
#artificial intelligence Preprint Open access Oct 2026

TasteRoute: Personalized Routing for Video Generation

Rapid progress in video generation has led to a plethora of models that differ substantially in capability and generation cost. This raises a natural question: can each request be efficiently routed to an appropriate model? We find that even when the consensus of the other annotators is used as an oracle, it agrees wit...

Zhi Rui Tam, Chao-Chung Wu, Sin-Han Yang et al. · 0 citations

From tech blogs

See all →
Microsoft Research Blog Oct 6, 2026

What AI gets wrong and what failure teaches us

Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity.  The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.