Skip to content

Category

computer vision

2,913 papers

#computer vision Preprint Open access Oct 2026

CADReasoner: Iterative Program Editing for CAD Reverse Engineering

Computer-Aided Design (CAD) powers modern engineering, yet producing high-quality parts still demands substantial expert effort. Many AI systems tackle CAD reverse engineering, but most are single-pass and miss fine geometric details. In contrast, human engineers compare the input shape with the reconstruction and iter...

Soslan Kabisov, Vsevolod Kirichuk, Andrey Volkov et al. · 0 citations
#computer vision Preprint Open access Oct 2026

EduGage: A Multimodal Dataset and Benchmark for Sensor-Based Momentary Assessment of Engagement in Self-Guided Video Learning

Engagement, which links to attentional, emotional, and cognitive dimensions, plays an important role in learning. In online and video-based learning environments, learners often need to regulate their own interactions with instructional materials. Measuring and reflecting on engagement can therefore support both learne...

Zikang Leng, Edan Eyal, Yingtian Shi et al. · 0 citations
#computer vision Preprint Open access Oct 2026

MemoCare: An Interactive Multimodal Mobile System for Automated Cognitive Screening

MemoCare is an interactive mobile system for automated multimodal cognitive screening. A React Native application combines spoken responses, temporal and spatial orientation, touchscreen actions, and visuoconstruction in complete English and Vietnamese workflows. Speech is transcribed by Google Speech-to-Text and score...

Duy-Cat Can, Mau Minh Phuc Le, Tuan-Khoa Hoang et al. · 0 citations
#computer vision Preprint Open access Oct 2026

PalmSpace: Towards a Versatile On-Palm Interaction Space through Unified Touch Modeling

As smart glasses and lightweight MR devices become increasingly practical, input remains a key challenge. The bare palm is an always-available, tactile, and proprioceptively accessible surface, but it has neither an explicit coordinate system nor embedded touch sensing. Prior on-palm systems typically expose isolated t...

Chentao Li, Mingze Gao, Runze Sun et al. · 0 citations
#computer vision Preprint Open access Oct 2026

Inverting Multi-Vector Visual Document Indices

Prevailing multi-vector visual document retrievers store each page as about a thousand patch vectors, often in vector databases run by a third party. Since no one can read a page from its vectors, this index is easily treated as less sensitive than the page. However, because the index keeps one vector per patch in rast...

Zhuchenyang Liu, Yao Zhang, Yu Xiao · 0 citations
#computer vision Preprint Oct 2026

Large-scale Repository Engineering via Agent-Native Reusable Code Primitives

Large language models equipped with development environments have moved code generation toward repository-scale construction, yet building complete repositories remains difficult because interacting modules, interfaces, configurations, tests, and dependencies must work together. We introduce Code Primitives, agent-nati...

Hai-Bo Jin, Peng Kuang, Xu-Cheng Yu et al. · 0 citations
#machine learning Preprint Open access Oct 2026

Modeling Robotics Dataset Construction as an Artifact-Based Build Process

Robotic systems generate large volumes of multimodal sensor data, but converting ROS bag recordings into machine learning datasets is often handled by ad hoc sequential scripts, creating engineering overhead and slow iteration cycles. We model dataset construction as an artifact-based build process over a dependency gr...

Leon Pohl, Lukas Beer, George Sebastian et al. · 0 citations
#machine learning Preprint Open access Oct 2026

Empirical Evidence for Simply Connected Decision Regions in Image Classifiers

The topology of a classifier's decision regions determines how inputs with the same predicted label can be connected and deformed without changing that prediction. Prior empirical work constructed paths between same-label images within a single region, but did not examine whether loops bound surfaces within that region...

Arjhun Swaminathan, Mete Akg\"un · 0 citations
#machine learning Preprint Open access Oct 2026

Component-Adaptive and Lesion-Level Supervision for Improved Small Structure Segmentation in Brain MRI

Small lesions in brain MRI are hard to segment because they occupy a tiny fraction of the volume and are dominated by background and larger lesions during voxel-wise optimization, so a model can reach a high Dice similarity coefficient (DSC) while missing many of them. We propose CATMIL, a training objective that adds...

Minh Sao Khue Luu, Evgeniy N. Pavlovskiy, Bair N. Tuchinov · 0 citations
#machine learning Preprint Open access Oct 2026

A Variational Latent-Space Framework for Uncertainty-Aware Spectral Image Emulation

Synthetic spectral image generation is essential for remote sensing simulation and mission design, yet physically based radiative transfer models (RTMs) remain computationally expensive. Existing learning-based emulators reduce this cost, but are mostly deterministic parameter-to-spectrum regressors with limited spatia...

Chedly Ben Azizi, Claire Guilloteau, Gilles Roussel et al. · 0 citations
#machine learning Preprint Open access Oct 2026

MAdam: Metric-Aware Multi-Objective Adam

Multi-objective optimization (MOO) underlies many machine learning problems, yet MOO solvers across the loss-balancing, gradient-balancing, and Pareto-based families almost universally hand their reconciled directions to Adam~\citep{kingma2015adam}. We show this coupling introduces two systematic gaps between the solve...

Fengbei Liu, Rachit Saluja, Sunwoo Kwak et al. · 0 citations
#machine learning Preprint Open access Oct 2026

Rubix: Global Correspondence-Free Point Set Alignment through Assignment Geometry

Procrustes-Wasserstein alignment jointly estimates a matching and rotation without supplied correspondences, but alternating minimization can stop at suboptimal solutions. Rubix solves the equally weighted planar problem globally under squared Euclidean loss. Each matching $\sigma$ of two centered $n$-point sets define...

Subhransu S. Bhattacharjee, Dylan Campbell, Rahul Shome · 0 citations

From tech blogs

See all →
Microsoft Research Blog Oct 6, 2026

What AI gets wrong and what failure teaches us

Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity.  The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.