Skip to content

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Book Open access Jul 2026

Cross-Domain Interest Representation Learning for Scenario- and Task-Aware Recommendation

Many internet companies operate multiple flagship applications, each of which can be regarded as a distinct business domain, covering areas such as video, reading, and gaming. Within each domain, diverse recommendation scenarios coexist, and users engage in various tasks with heterogeneous behaviors. In our industrial setting, we observe three key phenomena that existing methods rarely address: (i) users' Cross-Domain Interests (CDI) are weakly exploited, since behavior sequences are often pooled for efficiency, losing transferable cross-domain dependencies; (ii) multi-domain, scenario, and task variations are under-modeled, making it difficult to capture fine-grained complementarities; and (iii) multimodal features remain misaligned with ID features, especially when different domains emphasize different modalities. These gaps hinder cross-product collaboration. We propose STAR-CDR, a Scenario- and Task-Aware Cross-Domain Recommendation architecture. Following the modular decomposition of industrial recommenders, STAR-CDR introduces three innovations: (i) a CDI Extractor that enhances sequential modeling with domain indicators and other key features; (ii) a CDI Injector that recalibrates activations and enables domain-/scenario-/task-aware expert routing; and (iii) a CDI Multimodal Adapter that aligns image, text, and ID features under CDI-guided gating. Compared to the strongest baseline in our setting, STAR-CDR improves offline AUC 3.5% across five domains and yields +3.17% to +9.24% relative KPI gains in online A/B tests. The system now supports multiple domains (flagship applications), scenarios (recommendation scenarios), and tasks (user behaviors) at scale, serving over 280 million daily active users. We provide a GitHub repository. https://github.com/Rescomk/STAR-CDR with STAR-CDR's code, training scripts, and a synthetic data construction pipeline.

Bokai Lin, Naijun Gao, Yucen Gao et al. · 0 citations
Preprint Aug 2026

Track4Action: Distilling World-Centric 3D Tracker into Vision-Language-Action Policies

Action labels tell a vision-language-action (VLA) policy which robot commands to imitate, but not how those commands change the 3D world. The aligned demonstration clip contains this missing supervision because its $K$ frame transitions record the geometry, motion, visibility, and camera change produced during the corresponding $K$ actions. We introduce Track4Action, a framework that distills this realized transition from a frozen world-centric 3D tracker into a current-observation VLA policy. During training, Track4World encodes the clip $V_{t:t+K}$ into a pooled tracker feature. Learnable track queries infer this feature from current VLA hidden states, match it in a shared space, and condition a flow-matching action head through a feature-wise gate. The tracker feature only defines the alignment target, so neither the clip nor the tracker is used at deployment. Track4Action reaches 82.3% on zero-shot LIBERO-Plus, improving the alignment-free variant by 7.6 points and LaMP by 3.0 points. It obtains 80.44% and 81.48% on the clean and randomized RoboTwin 2.0 splits, and 67.5% average success across four physical bimanual tasks, 25.0 points above the alignment-free variant. The gains across simulation and physical tasks support action-aligned 3D tracker features as privileged supervision for tracker-free VLA deployment. Our project page is available at https://wing0night.github.io/track4action-project-page.

Chenyi Wang, Xinkai Wang, Bokai Lin et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.