LogiShot: Logically Coherent Cross-Shot Video Generation
LogiShot is proposed, which incorporates information through two complementary paths that jointly encodes the context video and other conditioning signals, yielding dense multimodal cues that provide visual-semantic evidence for cross-shot generation and the model maintains a visual memory of the context video throughout generation to preserve visual consistency across shots.