Skip to content

Category

computer vision

3,022 papers

#machine learning Preprint Open access Oct 2026

AEGIS: Anchor-Enforced Gradient Isolation for Knowledge-Preserving Vision-Language-Action Fine-Tuning

Fine-tuning pre-trained Vision-Language Models (VLMs) for robotic manipulation introduces a fundamental stability-plasticity dilemma: continuous flow-matching action experts backpropagate concentrated, low-rank regression gradients into transformer backbones trained on high-dimensional cross-entropy objectives. This cr...

Guransh Singh · 0 citations
#machine learning Preprint Open access Oct 2026

Embedding Prediction Helps Image Generation

In diffusion transformers, a class label or a text prompt is embedded once, and the same condition is reused at every denoising step. We ask whether predicted embeddings can serve as this condition instead. Next-Embedding Predictive Autoregression (NEPA) trains a Transformer to predict the next continuous embedding in...

Sihan Xu, Ji Xie, Zilin Wang et al. · 0 citations
#machine learning Preprint Open access Oct 2026

Weather-Aware Domain Adaptation for Street-View Weather Recognition

Adverse conditions such as rain, snow, fog, and dust remain challenging for camera-based perception in autonomous driving. We study multi-class weather recognition from street-view images under domain shift, where most available training data come from non-street-view sources that differ markedly from real driving scen...

Hossein Maghsoumi, George Atia, Yaser P. Fallah · 0 citations
#machine learning Preprint Open access Oct 2026

Comparing a gradient boosting algorithm to the GOES FDC for wildfire detection

Wildfires pose severe risks to human life, ecosystems, and property. This study presents a machine learning approach for wildfire detection from GOES ABI imagery. A CatBoost model was trained on a large dataset with thousands of ABI images and over 300,000 matching VIIRS fire detections. An evaluation on a separate dat...

Asaf Vanunu, Boaz Nadler, Arnon Karnieli · 0 citations
#machine learning Preprint Open access Oct 2026

PhaseAT: Fourier Phase Adversarial Training for Medical Image Domain Generalization

Reliable clinical deployment of deep medical image models is hindered by distribution shifts across scanners, sites, and acquisition protocols. Existing domain generalization (DG) methods often focus on style or intensity diversification, but they can still leave networks dependent on domain-specific texture correlatio...

Ahmed Sharshar, Asif Hanif, Naveen Kumar Kummari et al. · 0 citations
#machine learning Preprint Oct 2026

End-to-End Learning vs. Modular Architectures: Comparative Insights into Autonomous Driving Systems

Autonomous driving systems have become a central focus of intelligent transportation research, with End-to-End Learning and Modular Architectures offering two prominent design paradigms for their implementation. E2E Learning uses deep learning algorithms to map raw sensory inputs directly to driving actuators, providin...

Kartik B. Kapse · 0 citations
#machine learning Preprint Oct 2026

Do MLLM Judges Judge the Edit? Auditing Bias in Image Editing Evaluation with Verified Quality Preservation

Multimodal large language models (MLLMs) are increasingly used as automated judges for instruction-based image editing and as reward signals for model training. However, systematically auditing whether these judges are influenced by cues irrelevant to editing quality is challenging because visual interventions may them...

Yuan Huang, Zirui Song, Xiu-Ying Chen · 0 citations
#machine learning Preprint Oct 2026

Smoother Flow Matching via Contrastive Trajectory Repulsion

Trajectory crossing remains a critical bottleneck in Flow Matching (FM), and previous works typically view these crossings from a theoretical optimization perspective causing velocity averaging. They attempt to address it indirectly by post-hoc distillation or endpoint coupling, without explicitly regulating the interm...

Zi-Qi Jiang, Zhen-Qi He, Long Chen · 0 citations
#machine learning Preprint Open access Oct 2026

Open Vocabulary Word Recognition From Transcribed Bangla Texts

An optical character recognition (OCR) can scan a paper and extract text using technology, making people's jobs easier. While various OCR systems are available in the software industry, finding a reliable equivalent solution for Bangla takes much work. When it comes to handwritten texts, the situation is much more unus...

Faias Satter, Sk. Md. Masudul Ahsan · 0 citations
#machine learning Review Open access Oct 2026

A Survey on End-to-End Autonomous Driving Training From the Perspectives of Data, Strategy, and Platform

A Data-Strategy-Platform taxonomy is introduced that conceptualizes training as an interdependent system and surveys recent advances across data-centric pipelines, learning paradigms, and training infrastructures, and analyze their interplay in shaping model performance, robustness, and deployability.

Cheng-Kai Xu, Yi-Ming Cui, Jia-Qi Liu et al. · 5 citations
#machine learning Preprint Oct 2026

EyeTAG: Eye Trajectory-Aware Gaze Estimation

Gaze estimation under natural head-eye motion underpins applications from driver monitoring to human-computer interaction. Single-frame methods predict each frame independently, so consecutive outputs fluctuate as jitter. Multi-frame methods reduce this, but they learn motion implicitly inside appearance features, so t...

Jungmin Lee, Niamat Ullah, Yoseob Han · 0 citations
#machine learning Preprint Oct 2026

SmoothOperator: Enhancing Representations for Fine-grained Open-set Recognition via Modulated Label Smoothing

A plug-in, SmoothOperator (SmoothOP), which sets the smoothing coefficient of each sample from its prominence, an embedding-space signal measuring how clearly the sample's own class stands out against its strongest competing class, integrates into four existing spherical representation learning methods at minimal train...

Thiru Thillai Nadarasar Bahavan, Yu Xia, Sachith Seneviratne et al. · 0 citations

From tech blogs

See all →
Microsoft Research Blog Oct 6, 2026

What AI gets wrong and what failure teaches us

Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity.  The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.