Skip to content

Category

computer vision

3,022 papers

#machine learning Preprint Open access Oct 2026

Less Is More: A Leakage-Controlled Study of Dermoscopic Preprocessing for Joint Skin Lesion Classification and Segmentation with YOLO26

Handcrafted preprocessing is widely employed in automated dermoscopic analysis to suppress imaging artifacts and enhance lesion visibility. Nevertheless, its actual contribution to modern real-time models remains unclear, particularly when evaluation protocols do not adequately control correlations among images of the...

Truong Viet Vu, Nguyen Chi Hai, Nguyen Phuc Nguyen et al. · 0 citations
#machine learning Preprint Open access Oct 2026

Mu-DisCoCat: A Variational Pipeline for Compositional Generalization on Quantum Processors

Achieving compositional concept generalization (CoCoGen), the ability to understand novel situations by recombining learned primitives, remains a fundamental challenge in artificial intelligence. Compositional semantic models such as Compositional Distributional Semantics (DisCoCat) offer solutions by generalising vect...

Mina Abbaszadeh, Matilda Karabina Moore, Raem Haq et al. · 0 citations
#machine learning Preprint Oct 2026

Optimization Encoders: Rethinking Second-Order Meta-Learning for Neural Fields

Conditional neural fields represent signals continuously, but their effectiveness depends on how the conditional latent representations are inferred from observed data. In meta-learning, this encoding occurs through gradient updates induced by the decoder, tying representation learning directly to decoder design. We fo...

R. L. M. van Herten, Soufiane Ben Haddou, Rachit Saluja et al. · 0 citations
#machine learning Preprint Oct 2026

A Broader Look at Model Merging: Rethinking Implicit Regularization Induced by Task Arithmetic

Model merging aims to build a multi-task model cheaply by combining the weights of individual task-specific models. To perform well across multiple tasks, most existing merging methods use an additional dataset to find the coefficients for the best linear combination of task-specific weight updates. However, we identif...

Sin-Han Yang, Shih-Cheng Huang, Chieh-Yen Lin et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Label-Efficient Deep Learning for ECG Delineation: A Multi-Dataset Benchmark against Widely Used Delineation Tools

Electrocardiogram (ECG) delineation, the identification of waveform boundaries, is a foundational step that translates raw ECG signals into clinically interpretable measurements. Deep learning has advanced this task but remains dependent on costly expert annotations. Label-efficient strategies such as self-supervised p...

Jeonghwa Lim, Minje Park, Yeongyeon Na et al. · 0 citations
#machine learning Preprint Oct 2026

$\alpha$Transfer: Coefficient Transfer for Efficient Model Merging

Model merging offers a promising solution for combining multiple fine-tuned checkpoints into a single model through parameter arithmetic. However, finding optimal merging coefficients requires an extensive search that becomes prohibitively expensive as models scale in both size and number, due to high memory requiremen...

Shih-Cheng Huang, Zhi Rui Tam, Chieh-Yen Lin et al. · 0 citations
#machine learning Preprint Open access Oct 2026

Adversarially Trained Linear Transformers Are Optimal Robust In-Context Learners for Gaussian Mixtures

Adversarial training is one of the most reliable defenses against adversarial attacks, but its high computational cost must generally be paid anew for each task. Robust foundation models offer a promising alternative: adversarially pretrain a model once and then transfer its robustness to downstream tasks through light...

Soichiro Kumano · 0 citations
#machine learning Preprint Open access Oct 2026

Decoupling What from Where: How Should a Small GUI Grounding Model Receive the Action Type?

A GUI agent decides which action to take and where to take it; we ask how a small grounding model should receive the action type. Fine-tuning Qwen2-VL-2B with LoRA on Android in the Wild, we compare a flat baseline with five ways of supplying the type under matched data, compute, and decoding: an auxiliary loss, a hard...

Aadi Chauhan, Arthur Ilyasov · 0 citations
#machine learning Preprint Open access Oct 2026

Learnable Spectral Activations

Implicit neural representations (INRs) are shaped by the spectral structure induced by their input encodings and activation functions. Existing methods improve fitting primarily by modifying which frequencies are available to the network, through coordinate encodings or periodic nonlinearities. However, frequency acces...

Tamir Shor, Or Litany, Alex Bronstein · 0 citations
#machine learning Preprint Open access Oct 2026

ATLAS-AL: Adaptive Trust-Region for Latent Adversarial Searches via Active Learning

Security evaluation of learning-based systems requires more than just testing the system against a fixed collection of attacks. It requires adaptive mechanisms that can efficiently discover \textit{sets} of inputs that induce model failure. We introduce ATLAS (Adaptive Trust-Regions for Latent Adversarial Searches), wh...

Marsalis Gibson, Claire Tomlin, Shankar Sastry · 0 citations
#machine learning Preprint Open access Oct 2026

Sample-Optimal Estimation of the Fr\'echet Inception Distance

The Fr\'echet Inception Distance (FID) is widely used to evaluate generative models, but its empirical plug-in estimator suffers from finite-sample bias [BSAG18, CF20]. We study the sample complexity $n$ of estimating FID to error $\epsilon$ between $d$-dimensional Gaussians with bounded mean distance and covariances,...

Ziyun Chen, Jerry Li, Kevin Tian et al. · 0 citations
#machine learning Preprint Open access Oct 2026

Should We Skip Diffusion?

Diffusion models learn semantic representations while generating images. In the Decoupled Diffusion Transformer (DDT), a condition encoder provides features that guide a velocity decoder in denoising. To enable effective denoising at all noise levels, these features must capture both high-level abstract structures and...

Yiping Ji, James Martens, Simon Lucey · 0 citations

From tech blogs

See all →
Microsoft Research Blog Oct 6, 2026

What AI gets wrong and what failure teaches us

Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity.  The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.