Preprint
Aug 2026
CultureVidBench: Benchmarking Cultural Understanding in Text-to-Video Generation
Results show that although current models achieve strong semantic adherence and visual quality, they often fail to faithfully capture fine-grained cultural details, particularly for underrepresented regions, rituals, and multimodal cultural cues.
Xianjing Han, Yuhan Su, Yang Deng et al.
· 2 citations