Skip to content

Category

computer vision

3,022 papers

#artificial intelligence Preprint Open access Oct 2026

Dependable AI-Assisted Engineering: A Formal Framework for AI Participation and Assurance in Safety-Critical Workflows

Generative AI can produce engineering artefacts, but generation alone does not determine whether or how those artefacts should enter safety-critical workflows. This paper develops a formal framework for assigning AI participation and assurance at the level of individual workflow units. Each unit has a participation and...

Puxue Tan · 0 citations
#artificial intelligence Preprint Open access Oct 2026

MAGEFormer: Learning Metric-Consistent Representations for Anisotropic CT Segmentation

Vision Transformers (ViTs) have shown strong performance in volumetric segmentation, but their effectiveness on clinical CT is limited by an isotropic Euclidean lattice assumption. This conflicts with anisotropic CT acquisition, leading to two key issues: (1) a metric mismatch between voxel indices and physical anatomy...

Jiaying Li, Paolo Remagnino · 0 citations
#artificial intelligence Preprint Oct 2026

Selective Backpropagation for Efficient Few-Shot Class-Incremental Learning

Few-Shot Class-Incremental Learning (FSCIL) requires models to continuously learn new classes from limited samples while retaining prior knowledge, under strict constraints on compute and memory. Existing approaches lie along a difficult trade-off: simple fine-tuning is computationally efficient but suffers from catast...

Eeham Khan, Abdulmoumen Al-Atrash, Ali Ayub · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Streaming Multi-Track Timeline Control for 3D Human Motion Generation

Text-driven human motion generation has advanced substantially, yet most methods assume instructions are available before synthesis. Interactive applications require responding to new instructions while continuing ongoing actions, such as answering a phone while walking. Existing approaches address streaming generation...

Yangsong Zhang, Anujith Muraleedharan, Rikhat Akizhanov et al. · 0 citations
#artificial intelligence Preprint Oct 2026

GOTT: Object-centric Dexterous Manipulation with a Reusable Cross-Embodiment Primitive

Foundation models and large-scale human data provide rich sources of manipulation intent, but translating this intent into multi-fingered robot behavior remains difficult. Dexterous hands still lack a reusable low-level primitive that reliably establishes contact across tasks and embodiments. We propose GOTT, a reach-a...

Yu-Lin Liu, Lai Wei, Yen-Jen Wang et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

WAMJET: A Harness for World Action Model Acceleration

World Action Models (WAMs) leverage pretrained video foundation models for robot manipulation, but their large backbones and video-action co-prediction are expensive. Although existing acceleration techniques offer many ways to reduce this cost, selecting and composing them requires substantial engineering for each mod...

Le Chen, Lixin Liu, Jan Schneider et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Do Motion Tokenizers for Co-Speech Gesture Generation Encode Gesture Semantics?

Discrete motion tokenizers encode motion as atomic units and are widely used for co-speech gesture generation. It remains unclear which motion properties, especially those relevant to gesture semantics, are recoverable from these codebooks. We probe a reconstruction-trained codebook using 19 co-speech gesture descripto...

Varsha Suresh, Divij Jain, Jia Liu et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Impact of Data Augmentation on Confidence Calibration in Melanoma Classification

Accurately quantifying the predictive uncertainty or improving model calibration plays an important role in medical image classification, in particular in melanoma diagnosis, where accurate uncertainty quantification can have significant implications for patient care. One of the methods for calibration improvement is d...

Morgan May, Simon Caton, Pierpaolo Dondio · 0 citations
#artificial intelligence Preprint Open access Oct 2026

On Impact of Loss Function on the Performance of Neural Networks in Melanoma Diagnosis

Melanoma is the deadliest type of skin cancer, whose early diagnosis is crucial for patients' survival. Image classification using deep learning models has shown promising results for melanoma diagnosis. However, the performance of these models on the melanoma datasets such as SIIM-ISIC melanoma classification dataset...

Morgan May, Pierpaolo Dondio, Simon Caton · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Visual Grounding Safety in Vision-Language Models

Vision-language models (VLMs) are increasingly trained to generate structured outputs like points and bounding boxes that downstream interfaces, agents, and robots can act on, yet safety alignment of this output channel has not been systematically analyzed. We study visual grounding safety by repurposing three safety b...

Erfan Shayegani, Kundan Krishna, Yue Dong et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

DREAM: Dynamic Resolution Assignment For Multimodal Multi-agent Debate

Multi-agent debate (MAD) has emerged as an effective paradigm to improve the reasoning capabilities of large language models (LLMs) and is increasingly being extended to multimodal settings. However, existing multimodal MAD frameworks typically expose agents to the same fixed visual input, ignoring substantial variatio...

Khanh-Binh Nguyen, Van Dai Do, Tien Anh Nguyen et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

CASE: Cost-Aware Stopping for Efficient Long-Video Agents

Long-video agents can actively gather question-relevant evidence, but they typically leave a central decision implicit: when has the agent seen enough to answer? We propose CASE, a plug-in termination framework that frames this decision as policy-conditioned sequential stopping. At each causal checkpoint, CASE combines...

Yiming Du, Chenghao Liu, Zhiyuan Liu et al. · 0 citations

From tech blogs

See all →
Microsoft Research Blog Oct 6, 2026

What AI gets wrong and what failure teaches us

Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity.  The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.