Research on Optimization of Personalized Recommendation Strategies for Short Videos
Abstract
Personalized recommendation in short-video platforms requires accurate modeling of user interests from heterogeneous content signals while maintaining diversity, latency, and privacy constraints. To address data noise, recommendation homogenization, delayed interest capture, and ethical risks, this study proposes a multimodal recommendation optimization framework consisting of data governance, dynamic interest modeling, and compliance constraints. User behavior, video text, visual keyframes, audio features, and contextual information are collected through front-end logging and cleaned through anomaly filtering, missing-value processing, and differential privacy. CLIP, Swin Transformer, BERT, and MFCC features are integrated to form unified multimodal content representations. A three-stage recall–coarse-ranking–fine-ranking architecture is then developed, combining a dual-tower model, LightGBM, Wide&Deep, attention mechanism, dynamic interest decay, and PLE-based multi-task learning. Experiments on 1 million users and 5 million videos show that the optimized strategy improves click-through rate by 18. 3%, completion rate by 14.5%, content-type coverage by 22.1%, and average daily usage time by 15.7%, while reducing repeated recommendation and privacy-leakage risk. The framework supports multimodal signal processing, adaptive information delivery, and intelligent communication-service optimization.