COMET: Contrastive Motion-Enhanced Temporal Reasoning for Video Multimodal Large Language Models
COMET is a temporally grounded framework that systematically strengthens video MLLMs through explicit temporal representation, appearance-motion fusion, and direction-aware optimization and achieves consistent overall improvements with a pronounced motion-temporal bias.