TetherMem is introduced, a training-free, query-aware spatiotemporal memory router for frozen video generators that separates subject and scene queries and modulates historical access with region- and age-conditioned priors: subject queries retain identity-bearing history, while scene queries reduce reliance on subject history and stale backgrounds.
Chen Li, Peng Zhang, Han-Yu Zhou et al.· 0 citations
FilmEval is introduced, a systematic evaluation framework that couples a difficulty-graded benchmark of 15 representative novels with an automated protocol of nine objective metrics spanning three dimensions: cinematic presentation, film consistency, and novel fidelity.
Jialong Zuo, Haotong Zuo, Shiwei Zhang et al.· 0 citations
PaDoc is proposed, a layout-grounded parser that treats the predicted layout as a branching structure over a shared page representation that is the fastest end-to-end parser at five concurrency levels and is the fastest end-to-end parser at five concurrency levels.
A Cognitive-structured Multimodal Agent that externalizes visual information into an Episodic Visual Memory and selectively reactivates relevant episodes during reasoning is proposed, enabling reinforcement learning to optimize abstraction and retrieval policies.