Skip to content
Preprint

Preference-Driven Online Adaptation for Personalized Interaction Initiation in Proactive AI Assistants

Aug 2026 · 1 citation · 34 references
Computer Science

TL;DR

Evidence-driven Online Preference Adaptation (EOPA), which grounds a user's interaction-timing preferences in measurable contextual evidence through two evidence carriers: temporal preference anchors and evidence-bearing activity prototypes, is proposed.

Abstract

AI assistants are typically reactive, relying on users to initiate interactions. Proactive assistants go beyond this paradigm by autonomously initiating interactions based on users'activity contexts. However, appropriate interaction timing is user-specific and difficult to determine in advance, while online feedback offers valuable signals for personalization. Direct feedback-driven adaptation is therefore appealing, but remains challenging due to sparse interaction-worthy moments scattered across fine-grained user states. To address the issues, we propose Evidence-driven Online Preference Adaptation (EOPA), which grounds a user's interaction-timing preferences in measurable contextual evidence through two evidence carriers: temporal preference anchors and evidence-bearing activity prototypes. At each polling step, EOPA derives temporal and activity evidence from the carriers through user-prior-smoothed evidence estimation and uncertainty-guided evidence scaling, and adaptively fuses the evidence for interaction-or-silence decisions. When interaction is selected, an LLM uses high-quality historical responses as demonstrations to generate a context-aware response that better reflects user preferences. EOPA updates its evidence carriers and decision parameters from received online feedback without LLM-based reasoning or retraining. Extensive experiments on a ProPerSim-based benchmark show that EOPA improves the interaction-timing F1 score by 19.80 points over the strongest baseline in our experiments, substantially reduces inference latency for both silence and interaction steps, and lowers the average daily adaptation time from 11.41 to 0.39 seconds.

View source

Similar papers

#artificial intelligence Preprint Aug 2026

Hypotheses-Guided Self Distillation for Continual Personalization

HypReflect is introduced, a reliable, scalable framework for continual personalization that infers explicit, uncertainty-aware preference hypotheses from diverse user signals, reflectively refines them as new evidence accumulates, and incorporates the resulting user model through hypotheses-guided self-distillation.

Eunjeong Hwang, Kushan Mitra, Dan Zhang et al. · 0 citations
Preprint Aug 2026

Behavior2Trip: Towards Personalized Travel Planning via User Behavior Trajectory

A new task, Behavior-Aware Travel Planning, which infers user preferences directly from past behaviors and generates personalized travel plans and proposes B2T-Agent, a reinforcement learning-based agent that leverages user behavior trajectories, interacts with external tools for preference-aligned retrieval, and maint...

Zihao Cheng, Yingyu Shan, Hongru Wang et al. · 0 citations
Preprint Aug 2026

Intent Speaks Louder: Controllable User Simulation Beyond Response Imitation

This work introduces UserIDA (User Intent-Directive Alignment), which exposes interaction intent as an explicit per-turn directive, and establishes per-turn intent control as a complementary dimension to response fidelity in user simulation.

Bo Wang, Ruixing Zhang, Yunqi Liu et al. · 0 citations
Preprint Aug 2026

Learning from Online User Feedback for Shopping Agents

LOFA combines reinforcement learning over verifiable purchase outcomes with feedback-aware on-policy distillation, which identifies users' in-dialogue directives and converts them into dense token-level supervision, which captures both collaborative behavioral patterns and user-specific preferences.

Haobo Zhang, Ke-Long Mao, Su-Long Xu et al. · 1 citation
Book Open access Aug 2026

Shape Your Feed: An LLM-based Agentic System for Conversational Recommendation

Industrial recommendation systems predominantly adopt a passive ranking paradigm that infers user preferences from implicit behavioral signals (e.g., clicks, dwell time) rather than explicit, natural language inputs. As a result, users experience a persistent discrepancy between their explicit interests and what passiv...

Zi-Yun Xu, Bo-Sen Ding, Yue Zhang et al. · 1 citation
Jun 2026

Report on The 1st Workshop on Human-Centered Proactive and Personalized Agents for Interactive Information Access at CHIIR 2026

Interactive information access is increasingly moving beyond reactive query-response paradigms toward agentic systems that can personalize interaction, retain context, infer latent needs, recommend next steps, and initiate support. This shift creates new opportunities for adaptive and context-aware assistance, while al...

Kirandeep Kaur, Vinayak Gupta, Tanya G. Roosta et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.