Video and audio are perceived together, yet most generative models treat them in isolation. We examine methods that model the two modalities jointly, generate one from the other, or edit them in a coupled manner, organized around a single question: how is the output kept coherent across modalities in time and semantics...
Abhinav Sharma, S. Navuluru, Wang Wei et al.· 0 citations
This work proposes Co-E, a training-free system built around synchronized bidirectional graph-text working memory, which improves over comparable training-free open-backbone baselines and is competitive with larger or trained systems.
Hieu Man, Thien Huu Nguyen· arXiv.org· 0 citations
This work replaces the Jensen-Shannon Divergence routing with C_struct, a structural proxy that measures mass at Vertical-Slash compatible positions and reproduces JSD's routing decisions while eliminating both the pooled matmul and subsequent KL divergence overhead.
H. Nguyen, Chien Van Nguyen, Franck Dernoncourt et al.· 0 citations
This work proposes to synthesize alignment sequence pairs and fine-tune an encoder model with span alignment objective and introduces EXP - the first benchmark for explicit evaluation of label projection, thereby reducing confounders and non-determinism in method assessment.
Thang Le, Huy Huu Nguyen, A. Luu et al.· Annual Meeting of the Associ...· 0 citations
O CTOPUS is proposed, a framework that confers fixed-memory inference onto pretrained Transform-ers without the information loss of linearization and outperforms state-of-the-art linearized baselines on the GSM8K benchmark, demonstrating that learned sparse retention serves as an effective regular-izer for long-horizon...
C. Nguyen, Ryan A. Rossi, L. Van et al.· Annual Meeting of the Associ...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.