Counterfactual Trace On-Policy Distillation (CT-OPD), which combines completed teacher responses with trajectory masks from the current student, consistently enhances multimodal understanding and reasoning capabilities, and improves both visual understanding and image generation.
Long Qian, Bing-Ke Zhu, Jia-Qi Wei et al.· 0 citations
OmniMoE-VL, a sparse VLM with a coupled visual-depth routed projector, enables question-dependent visual access while preserving the native visual-token sequence, and complements token-level expert routing in the vision and language stacks.
Long Qian, Bing-Ke Zhu, Jia-Qi Wei et al.· 0 citations
Fidelity visual compression is formulated as constructing a compact coreset for decoder messages, and Grounded Message Coreset Pruning is introduced which jointly allocates support across query-grounded, appearance, and coordinate-aware evidence, then transports discarded states into selected representatives at their o...
Long Qian, Jiaqi Wei, Bingke Zhu et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.