With limited annotation budgets, choosing which images to label determines how much a model improves. Data-selection methods that use features from a separately trained model, or scene descriptions written by vision-language models, have been successful, but those signals do not directly capture changes in the model be...
Recent representation alignment (REPA) methods accelerate diffusion transformer training by aligning projections of the transformer's hidden states with representations from pretrained visual encoders. In this work, we explore a reverse and complementary direction to REPA: rather than projecting diffusion representatio...
Han-Si Fu, Jia-Cheng Chen, Bao-Quan Zhao et al.· 0 citations
When a VLM answers a visual query, current interpretability tools rely on text rationales, which use a mismatched modality, or on internal read-outs, which originate too early to reflect the final output and require white-box access to the model. We introduce AnswerMap, a training-free, task-agnostic, black-box visual...
Mohamed Eltahir, Fardows Adam, Duaa M. Tahir et al.· 0 citations
Braco is a lightweight four-step coder that combines transform-basis truncation, input-independent basis-coordinate embeddings, budget-dependent orthogonal re-parameterization, and learned spatial residual tokens from lightweight pooling that forms the favorable empirical accuracy-efficiency frontier under compression.
Rui-Lian Zhong, Yu Li, Zhe-Yu Yan et al.· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
Post-training foundation video models on heterogeneous reward-weighted data usually assume that all data categories induce compatible updates. This assumption is fragile when categories correspond to different skills, domains, or evaluation dimensions. We study this problem in text-to-video post-training, where VBench2...
Jia Song (The Hong Kong University of Science and Technology), Wenhow Li (The Hong Kong University of Science and Technology), Lichen Bai (The Hong Kong University of Science and Technology) et al.· 0 citations
Deepfake detection methods have become increasingly effective yet most provide limited insight into the evidence behind their predictions. However, in forensic settings users also need to know which manipulation cues support the decision and where they appear. Existing explainability methods only partially address this...
Georgios Tsoumplekas, Vazgken Vanian, Alexandros Doumanoglou et al.· 0 citations
It is found that applying textual guidance too early can limit its ability to identify answer-relevant visual regions, whereas text-to-visual attention becomes more informative at intermediate decoder depths.
Min-Chan Kang, Kyeonghye Park, Seoyoung Cho et al.· 0 citations
Vision-language models (VLMs) can answer simple visual questions, but often struggle when one question requires several visual judgments. We study this gap with controlled tasks for feature binding, numerosity, spatial relations, and amodal completion, together with a Composite task that combines them. Matched counterf...
Experiments on multiple LVLMs show that the Balanced Fitting method consistently outperforms prior PTQ approaches under both weight-only and weight-activation quantization, while lower reconstruction loss does not reliably translate into better downstream performance.
Min-Chan Kang, Kyeonghye Park, Seungyeon Sa et al.· 0 citations
CMES (Cross-Modal Emotion Stimuli), a multi-source collection of emotion-conditioned stories, real facial expressions, synthetic portraits, and synthetic emotion-evoking scenes, suggests that emotion representations can share relational structure and causal effects across sources, modalities, and architectures, even wh...
Bo-Hao Xing, Xin Liu, Kai-Shen Yuan et al.· 0 citations
Accurately decoding object states from the internal representations of vision-language-action (VLA) models does not establish that the predictions respond faithfully to changes in the target physical state. In natural observations, object state, robot configuration, occlusion, and task progress vary together, allowing...
Hyungjoon Kim, Wonbin Son, Mi Young Lee et al.· 0 citations
ReaLVR is proposed, which brings visual-evidence supervision to the model's own free-running latent trajectories, and is the first to scale visual reasoning in latent space, showing that the framework continues to deliver robust improvements at frontier model scales up to 235B.
Xi Xiao, Tian-Chen Zhao, Youngeun Kim et al.· 0 citations
Exploring how generative AI could make machine vision more accessible to businesses. The post GenEye in a Box: Making Machine Vision Something You Can Just Ask For appeared first on GPT-Lab.
Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity. The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.
Computer scientist, entrepreneur, and philanthropist will collaborate with the MIT Schwarzman College of Computing to advance AI and scientific discovery.