Video multimodal large language models (MLLMs) keep climbing video question answering benchmarks, yet shuffling the frames, masking the segment that supports the answer, or occluding the target object barely changes their predictions. The accuracy rests on appearance and language priors, not on the temporal evidence th...
Zhaolu Kang, Shi-Yu Liu, Tai-Long Luo et al.· 0 citations
The conditional generation of 3D motion has emerged as a key research topic due to its wide applicability across robotics, AR/VR, gaming, and content creation. However, extending recent advances in text-driven human motion generation to the animal domain remains challenging due to two core limitations. First, animals e...
Ze-Yu Zhang, Zhi-Yuan Zhang, Si-Heng Wang et al.· 0 citations
This method generates candidate responses from the target policy, evaluates helpfulness, factuality, and conciseness with rubric-specialized evaluators, applies a process-critic correction, and retains only high-consensus desirable or undesirable examples.
PreDiff-LM preserves causal attention within the observed prompt while allowing full bidirectional attention within the masked target, position hybrid attention as a complementary mechanism for adapting pretrained causal backbones, while making explicit the remaining quality and inference-efficiency gaps to optimized A...