Skip to content
Preprint

Large Behavior Model: A Promptable Digital Twin of the Retail Customer

Jul 2026 · 0 citations · 35 references
Computer Science

TL;DR

The results demonstrate that behavioral knowledge encoded in transaction histories can be effectively learned by language models, providing a scalable foundation for customer digital twins and behavior simulation.

Abstract

Customer behavior modeling underpins recommendation, marketing, and decision support, yet existing approaches either optimize predictive accuracy without explaining decisions or simulate users without grounding them in real behavioral data. We present the Large Behavioral Model (LBM) that learns customer decision making directly from large-scale retail transactions through a unified Person-Environment formulation. Customer state is represented by a behavioral profile derived from historical purchases, while product context is incorporated through retrieval-augmented generation. The model is trained using continued pre-training on verbalized behavioral data, supervised fine-tuning for decision generation, and reinforcement learning with verifiable rewards for evidence-based calibration. We evaluate the proposed framework on purchase prediction, hard-negative discrimination, basket completion, promotion response, and cross-domain voucher redemption. The model consistently outperforms frontier general-purpose language models on in-domain retail tasks while demonstrating strong zero-shot and fine-tuned transfer across retailers and decision domains. Ablation studies show that continued pre-training is the primary driver of behavioral generalization, retrieval is most effective when applied during both training and inference, and reinforcement learning improves reliance on explicit behavioral evidence over generic language-model priors. These results demonstrate that behavioral knowledge encoded in transaction histories can be effectively learned by language models, providing a scalable foundation for customer digital twins and behavior simulation.

View source

Similar papers

Jul 2026

Beyond Action Imitation: Learning a Decision-Aware User Simulator for Online Advertising

DASH is a decision-aware user simulator that jointly generates thinking traces and predicts behavioral actions from heterogeneous cross-domain histories and tailors a rubric-based reward model that evaluates thinking traces along form, content, and logic for RL training.

Zi-Hang Chen, Jiaer Zheng, Xiangyang Xu et al. · 0 citations
Conference Jul 2026

Agent-Driven Customer Retention using Reinforcement Learning and Behavioral Analytics

Customer retention continues to be a significant issue for contemporary firms because of increasing competition and evolving customer dynamics. Existing churn prediction models tend to concentrate on detecting churners but lack mechanisms for formulating adaptive decisions for retention purposes. In this study, we propose a customer intelligence approach to identify customer personas through RFM (recency, frequency, monetary value) feature engineering, natural language processing, and K-Means Clustering techniques to generate customer personas. We then use reinforcement learning to select specific retention actions based on customer personas. For this purpose, we design an agent-based system that takes as input a set of persona-dependent actions and simulates feedback using a feedback simulation environment to generate rewards and learn optimal retention strategies for customers. Our experimental evaluation conducted on an e-commerce behavioral data set shows promising results in developing agentic decisions for retention actions based on the generated customer persona. Our study contributes to the literature by showing how to use reinforcement learning for decision-making processes in customer retention problems.

Abhay Pratap Singh, Prakyath S Kumar, Advait Gujar et al. · 0 citations
Conference Jul 2026

Behavioral Satisfaction Inference: A Semi-Supervised Framework for Large-Scale Telecom Customer Intelligence

Customer satisfaction is a key factor influencing customer retention in the telecom industry. However, conventional survey-based measurement approaches are often constrained by low response rates, sampling bias, and limited population coverage, making it difficult to obtain a comprehensive view of subscriber sentiment. This paper presents Behavioral Satisfaction Inference, a semi-supervised learning framework designed to estimate customer satisfaction across an entire subscriber base without relying on explicit feedback from every customer. The proposed approach combines a relatively small set of labeled survey responses with a large volume of unlabeled behavioral data. Using a pseudo-labeling strategy, satisfaction signals are propagated across more than 2 million subscribers while addressing a highly imbalanced distribution in which fewer than 15% of customers are dissatisfied. The framework incorporates multiple data sources, including service usage patterns, billing behavior, and network quality indicators, to generate continuous satisfaction predictions at scale. Experimental results demonstrate that the semi-supervised framework outperforms supervised-only baselines, improving recall for dissatisfied customers by more than 18%. The final model achieves an 86.2% recall rate in identifying subscribers at risk of dissatisfaction. Deployed in a production telecom environment, the system supports proactive retention efforts by enabling broader and more timely identification of potentially dissatisfied customers. The findings highlight the effectiveness of semi-supervised learning for large-scale satisfaction estimation and offer an alternative to traditional survey-centric approaches.

Ammar Ayub · 0 citations
Jul 2026

Predicting Consumer In-Store Purchase Through Real-Time Video Analytics: An Advanced Computer Vision and Deep Learning Approach

Physical retailers have long lacked the real-time behavioral visibility that online platforms enjoy through clickstream data. This research addresses that gap by introducing a video analytics framework that transforms in-store security camera footage into a rich, structured behavioral record: an "offline clickstream." Using computer vision and deep learning techniques, including person re-identification, trajectory reconstruction, pose estimation, and vision-language models, the system extracts moment-by-moment signals of shopper intent: how customers move through the store, how they interact with products, and how their body language evolves during a visit. A transformer-based prediction model trained on these signals achieves dramatically better purchase prediction accuracy than conventional demographic or contextual benchmarks alone: improving predictive performance by up to 79% on key metrics. Beyond prediction, the framework supports five real-time targeting policies; simulations show that a persuadability-based policy yields a 13.1% profit lift over no targeting. For retailers and policymakers, this research offers a scalable, privacy-conscious blueprint for bridging the capability gap between physical and digital commerce, enabling timely, personalized interventions that improve customer experience and store profitability.

Rubing Li, Wen Wang, Kaiquan Xu et al. · 0 citations
Conference Open access 2026

E-commerce User Purchase Intention Prediction Based on Temporal Behavior Modeling

The paper deals with the problems of heterogeneous behavior fusion and long-sequence interest modeling, which naturally leads to the proposal of a temporal prediction model in the Taobao user behavior dataset context, which combines a multi-behavior aware attention mechanism with gated recurrent units (GRU), and uses separate behavior embeddings to encode browsing, favoring, adding to cart, and purchasing. The paper first describes how it dynamically focuses on relevant behavior segments using an attention mechanism and then naturally introduces multitask learning to learn better representations. On the basis of empirical analysis, it convincingly establishes that user behavior has a single peak at noon, a sharp spike on weekends, and that adding items to the cart is the most reliable purchase signal. The experiments also show that the proposed model outperforms the baseline in both F1 score (0.926) and recall (0.938). Ablation experiments clearly and convincingly showed that behavior type embedding is the fundamental building block of the model, and hence its omission seriously degrades performance. Therefore, the paper naturally and logically establishes the effectiveness of temporal modeling for purchase intention prediction, which has direct implications for e-commerce applications.

Jianyi Lyu · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.