Generating text-rich images from prompts requires both textual fidelity and the coherent integration of text into the surrounding image. An explicit layout can provide structured guidance about what text should appear and where, but a well-formed plan alone does not guarantee that the renderer will realize it faithfull...
Guan-Qiao Chen, Jing Tan, Dong-Xing Mao et al.· 0 citations
DAFS (Dynamic Attention-based Budget-aware Frame Selection), a training-free frame selector that improves over uniform sampling by up to 6.4 points on Video-MME and outperforms prior training-based selectors under matched frame budgets, while generalizing across selector and answerer backbones, and across tasks, withou...
Yilin Wang, Xiangxi Zheng, Dongxing Mao et al.· arXiv.org· 2 citations
SPIRAL (Self-improving Path Integration and Realignment), a self-supervised alignment framework that closes this gap using only the model's own text-path behavior as supervision, requiring no external teachers or additional annotations, generalize to out-of-domain benchmarks, confirming that effective VTC hinges on ali...
Tianyu Liang, Xiangxi Zheng, Yilin Wang et al.· 2 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.