Counterfactual Trace On-Policy Distillation (CT-OPD), which combines completed teacher responses with trajectory masks from the current student, consistently enhances multimodal understanding and reasoning capabilities, and improves both visual understanding and image generation.
Long Qian, Bing-Ke Zhu, Jia-Qi Wei et al.· 0 citations
OmniMoE-VL, a sparse VLM with a coupled visual-depth routed projector, enables question-dependent visual access while preserving the native visual-token sequence, and complements token-level expert routing in the vision and language stacks.
Long Qian, Bing-Ke Zhu, Jia-Qi Wei et al.· 0 citations
Time-series language models provide a shared natural-language interface across temporal tasks, but plausible text does not guarantee reliable task outputs. Responses may appear reasonable while hallucinating the required object: numerical sequences can violate shape, scale, channel order, or temporal alignment, and tex...
Jia-Hui Chen, Bing-Ke Zhu, Hong-Yu Pan et al.· 0 citations
Fidelity visual compression is formulated as constructing a compact coreset for decoder messages, and Grounded Message Coreset Pruning is introduced which jointly allocates support across query-grounded, appearance, and coordinate-aware evidence, then transports discarded states into selected representatives at their o...
Long Qian, Jiaqi Wei, Bingke Zhu et al.· 0 citations
Results demonstrate that model-external structured state maintenance and closed-loop agentic decision making can effectively extend the local control capabilities of WAMs into embodied task execution that is plannable, verifiable, and recoverable.
Zhao-Peng Gu, Bing-Ke Zhu, Tianxin Lin et al.· 5 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.