Skip to content
Open access

Selection Bias Correction in Retail Intelligence

Jul 2026 · International Conference on Data Technologies and Applications · pp. 291-298 · 0 citations · 22 references
Computer Science

Abstract

Retail intelligence often relies on monitoring popular, high-velocity products, potentially biasing economic indicators by ignoring the"long tail"of niche items. This simulation study investigates selection bias in inflation estimation and compares correction methods across diverse data-generating processes. Through 400 Monte Carlo replications spanning four scenarios--aligned step functions, smooth gradients, misaligned breaks, and polynomial relationships--we test the robustness of Inverse Probability Weighting (IPW) with five specifications against stratification with varying strata counts. Our findings reveal fundamental limits of weighting methods in retail long-tail contexts: stratification achieves superior performance in three of four scenarios, maintaining sub-0.04pp median error even when boundaries deliberately misalign with population breaks (116x advantage over IPW). However, IPW with spline propensity models wins under smooth polynomial relationships (median error 0.007pp vs. 0.013pp), demonstrating context-dependency. Critically, even an oracle IPW specification with perfect structural knowledge achieves 6.06pp error compared to stratification's 0.008pp in step-function scenarios. This reflects violation of the Positivity Assumption--a fundamental causal inference requirement--rather than IPW methodological inferiority. When selection probabilities differ dramatically (90% vs. 1%), weighting methods operate outside their theoretical design envelope. These results demonstrate that stratification provides a safer engineering choice in retail long-tail distributions with severe positivity violations.

Read PDF

Similar papers

Jul 2026

Tabular Foundation Models for Discrete Choice Estimation

Tabular foundation models (TFMs) generate predictions on structured data via in-context learning, without task-specific estimation. We ask whether TFMs can be effectively applied to discrete choice, a central demand estimation framework in marketing and operations, and find that directly applying TFMs yields limited performance. The gap is structural: TFMs assume row-independent observations, whereas discrete choice is inherently set-valued and subject to persistent consumer preference heterogeneity. We propose a reformulation that encodes both choice-set dependence and individual heterogeneity within a row-based learning framework. Evaluated on a yogurt scanner panel, individual-level heterogeneity encoding is the dominant driver of predictive accuracy. The best reformulation outperforms hierarchical Bayesian estimation on both holdout log-likelihood and hit rate, running 16 times faster, a practical advantage for large-scale demand estimation. The advantage is largest in the medium-data regime (10--40 purchase occasions per consumer), where parametric Bayesian shrinkage most distorts estimates for atypical consumers. Fine-tuning on population choice data provides additional gains for consumers with shallow purchase histories, where in-context learning has limited individual-specific signal to condition on. These results establish a principled approach for applying foundation models to consumer choice problems more broadly.

Liu Liu, Danlu Zhang · 0 citations
Preprint Aug 2026

Characterizing Bias in Post-Bandit Inference under Index Algorithms

Bandit algorithms generate data for downstream inference, but adaptive sampling biases post-bandit sample means. We analyze this bias for stable index algorithms, including UCB1 and its generalizations, and derive sharp leading-order expressions for the sample-mean bias and expected $Z$-statistic. Our characterization reveals the algorithmic origin of bias through a key index-function-dependent quantity, which we term effective exploration rate. For example, under UCB1, the effective exploration rate is of order $\sqrt{\log T}$, and the standardized bias of any arm (that is not uniquely optimal) decays at the extremely slow rate $1/\sqrt{\log T}$. We also show how the choice of the index function affects both regret and bias, which reveals a regret-bias trade-off: more exploratory algorithm reduces bias but increases regret. Our sharp characterization for bias uses a novel empirical fluid approximation of the algorithm's sampling dynamics, which may be of independent interest.

Lisu Wang, Yilun Chen, Jiaqi Lu · 0 citations
Open access Sep 2026

The Fairness Illusion? A Cross-Dataset Audit of Accuracy and Demographic Bias in Credit Scoring Based on Machine Learning

Machine learning has transformed consumer credit scoring, delivering substantial gains in predictive accuracy over traditional scorecards—but whether those gains come at a cost to fairness has remained contested. The dominant assumption in the literature is that more complex, accurate models amplify bias by encoding historical patterns of disadvantage more effectively. This paper challenges that assumption with direct empirical evidence. We evaluate four model families—logistic regression, random forest, XGBoost, and a multilayer perceptron—across two real-world datasets: the UCI Credit Card Default Dataset and the 2024 US Home Mortgage Disclosure Act national loan-level data, comprising over six million mortgage applications. Using repeated cross-validation, we report predictive performance alongside two primary fairness metrics—demographic parity difference and equalized odds difference—supplemented by false positive rate difference and calibration difference, with confidence intervals across 15 estimation folds. On the UCI data, where demographic disparities are modest, model choice has negligible effect on fairness outcomes. On the HMDA mortgage data, where racial disparities are large and legally consequential, the expected accuracy–fairness tradeoff does not hold; more accurate models produce significantly fairer outcomes on equalized odds within the models, data, and fairness criteria examined here, with logistic regression occupying the worst position simultaneously on all dimensions. Persistent demographic parity disparity among the more complex models is consistent with feature-level bias that no model architecture can resolve. The findings have direct implications for the less-discriminatory-alternatives framework under US fair lending law and for the high-risk classification of credit scoring AI under the EU AI Act.

Unknown authors · 0 citations
Open access Jul 2026

An Empirical Study of Robust Portfolios Based on Statistical Noise Reduction and Regularization Constraints

The results suggest that including a covariance noise or normalization term in the traditional expected value and variance-based approach to portfolio management helps reduce the impact of estimation error and market volatility on portfolio performance.

Zichun Fu · 0 citations
Preprint Jul 2026

Beyond recency and magnitude: Learning-aware uncertainty estimation for safety stock under evolving forecast-error distributions

We develop a non-parametric approach to refine forecast-error histories for safety stock estimation without relying on error magnitude or recency alone. Using forecast-error data from a fast-moving consumer goods environment, we show that one year of errors is insufficient to characterize service-level risk, while pooling several years is problematic because error distributions change over time. Non-parametric annual comparisons indicate a progressive reduction in bias and dispersion, consistent with learning in the forecasting process. We propose LOWDII (Leave-One-Out Wasserstein Distributional Influence Index), a diagnostic that evaluates the distributional influence of each historical forecast error on the empirical uncertainty distribution used for safety stock calibration. LOWDII identifies observations whose influence is disproportionate to their representativeness, separating transient distortions from persistent tail behaviour without imposing parametric assumptions. We evaluate the method in a discrete-event simulation using subsequent-year demand realizations as validation. The results show that LOWDII achieves the target service level while reducing average stock by 5.1% to 22.6% relative to benchmark methods. From a managerial perspective, LOWDII helps firms translate improvements in the forecasting process into safety stock decisions, avoiding the projection of obsolete historical errors into the future and capturing the economic benefit of lower uncertainty requirements earlier.

Luis Fernández-Palacios, M. Ceballos, Yolanda Muñoz-Ocaña · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.