Skip to content

Category

data science

2,431 papers

#artificial intelligence Preprint Open access Oct 2026

On the Limits of LLM Adaptability: Impact of Model-Internalized Priors on Annotation Task Performance

Large Language Models (LLMs) are increasingly used for zero-shot annotation and LLM-as-a-judge tasks, yet their reliability hinges on how model-internalized priors interact with user-provided instructions. We investigate three dimensions of this interaction: (1) how an LLM's familiarity with data and task definitions r...

Etienne Casanova, Rafal Kocielnik, R. Michael Alvarez · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Escaping the Capacity Ceiling: Routing on the Stiefel Manifold for Bilinear SPD Layers

Deep networks on the symmetric positive-definite (SPD) manifold promise expressive representations by encoding data geometry as an inductive bias, but stacking BiMap layers with the standard ReEig nonlinearity often adds no capacity: on real, preconditioned EEG data, ReEig rarely activates, so the stack behaves as a si...

Isabella C. Maia, Salem Said, Pedro L. C. Rodrigues et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

AREX: Affine-Residual Exponential Integrator for Few-Step Sampling in Flow Matching

We introduce AREX, a training-free sampler for pretrained flow matching models that uses the target mean and covariance to capture an analytically tractable part of the sampling dynamics. We show that the velocity field of the moment-matched Gaussian target is the $L^2$-optimal affine approximation to the marginal velo...

Shizheng Lin, Soon Hoe Lim, N. Benjamin Erichson · 0 citations
#artificial intelligence Preprint Oct 2026

Nearly Optimal Fixed-Confidence Best-Arm Identification with 1-Bit Feedback

We study fixed-confidence best-arm identification under strict 1-bit feedback constraints. At each round, the learner selects an arm and a query set, and receives only a single bit indicating whether the sampled reward belongs to that set. We consider a distribution-free finite-variance setting with arm-wise localizati...

Khang A. Luong, Sơn Thái Đinh, Ho-Ang Ta et al. · 0 citations
#artificial intelligence Preprint Oct 2026

Generalization Bounds for Flow-matching Generative Models for Intrinsically Low-dimensional Data

Despite the remarkable empirical success of flow-matching models, their statistical generalization guarantees remain underdeveloped. Existing analyses often impose restrictive assumptions on the estimated velocity field and yield convergence rates that fail to reflect the intrinsic low-dimensional structure common in r...

Saptarshi Chakraborty, Quentin Berthet, Peter L. Bartlett · 6 citations · ⚡1
#artificial intelligence Preprint Open access Oct 2026

Learning Style, Forgetting Semantics: A Case Study of SFT and RFT on Classification Tasks

Why does supervised fine-tuning (SFT) lead to more forgetting than reinforcement fine-tuning (RFT), even when all teacher demonstrations are semantically correct? We study this question on classification tasks where tokens within each semantic class express the same semantic answer in different styles. The tasks share...

Haodong Liang, Yanhao Jin, Krishnakumar Balasubramanian et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Counterfactual Predictions in Scientific Emulators Without Controlled Experiments

Many scientific questions require reasoning about what was never observed: What if the conditions, interventions, or history had been different? Models can predict accurately on observed data yet fail on such what-if queries when correlated inputs are varied independently. A common remedy is to add controlled simulatio...

Ding-Ling Yao, Kahaan Gandhi, Valentin Duruisseaux et al. · 0 citations
#data science Open access Oct 2026

Conditional data value and decision loss in credit screening

Replication materials accompanying “The marginal value of alternative data in credit screening: Evidence from Chinese digital lending” by Yuan Chen and Jiawei Xu. The package contains instructions for obtaining the source data, environment specifications, data processing and modelling scripts, parameter settings, rando...

Chen Yuan, Xu Jiawei · 0 citations
#data science Open access Oct 2026

Development and in Vitro Evaluation of Ferti-Max for Extended Liquid Storage of Broiler Breeder Rooster Semen

This study describes the formulation and in vitro evaluation of FERTI-MAX, a proprietary extender developed through iterative, osmolality-guided optimization to approximate the osmotic and pH environment of chicken seminal plasma. Successive prototypes were narrowed from an initial hypertonic range of approximately 520...

Debashis Dutta, Apratim Maity, Subhashish Batabyal et al. · 0 citations
#data science Open access Oct 2026

Development and in Vitro Evaluation of Ferti-Max for Extended Liquid Storage of Broiler Breeder Rooster Semen

This study describes the formulation and in vitro evaluation of FERTI-MAX, a proprietary extender developed through iterative, osmolality-guided optimization to approximate the osmotic and pH environment of chicken seminal plasma. Successive prototypes were narrowed from an initial hypertonic range of approximately 520...

Debashis Dutta, Apratim Maity, Subhashish Batabyal et al. · 0 citations
#data science Open access Oct 2026

Influence of Meteorological Conditions and Temperature-Humidity Index on Fertility, Hatchability and Embryonic Mortality in Kadaknath Chickens

The present study was conducted from May 2023 to August 2024 at the Rashtriya Krishi Vikas Yojana (RKVY) backyard poultry farm, Faculty of Veterinary and Animal Sciences, Barkachha, Mirzapur, Uttar Pradesh, India, to assess the influence of meteorological conditions on fertility, hatchability, and embryonic mortality i...

anuradha kumari, Utkarsh Kumar Tripathi, Kumar Anshuman et al. · 0 citations
#data science Open access Oct 2026

Influence of Meteorological Conditions and Temperature-Humidity Index on Fertility, Hatchability and Embryonic Mortality in Kadaknath Chickens

The present study was conducted from May 2023 to August 2024 at the Rashtriya Krishi Vikas Yojana (RKVY) backyard poultry farm, Faculty of Veterinary and Animal Sciences, Barkachha, Mirzapur, Uttar Pradesh, India, to assess the influence of meteorological conditions on fertility, hatchability, and embryonic mortality i...

anuradha kumari, Utkarsh Kumar Tripathi, Kumar Anshuman et al. · 0 citations

From tech blogs

See all →
Microsoft Research Blog Oct 6, 2026

What AI gets wrong and what failure teaches us

Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity.  The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.