Scalable driver identification requires embedding models that maintain discriminative performance as fleet size grows, yet existing triplet-loss formulations degrade rapidly with driver pool size and overfit to session-specific patterns under rigorous temporal evaluation. We introduce TEMPEST, a Temporal Convolutional...
Kyle Musgrove, Dylan B. Lewis, Sarah Powers et al.· 0 citations
Roadside traffic reasoning requires every free-form textual claim to be backed by visual evidence. Existing grounded multimodal large language models (MLLMs) frequently exhibit say-point mismatch, in which the textual answer contradicts the bounding boxes the model localizes. Evaluation metrics that score answers and b...
Runwei Guan, Rongsheng Hu, Shangshu Chen et al.· 0 citations
Continuous colormaps are widely used to visualize scalar fields, and their quality is typically evaluated using measures such as discriminative power and uniformity. Existing measures primarily characterize the intrinsic perceptual properties of the colormap itself, largely independent of the underlying data distributi...
Xi Duan, Yiwei Lin, Shiqing Xin et al.· 0 citations
High-fidelity diagram creation requires the complex orchestration of semantic topology, visual styling, and spatial layout, posing a significant challenge for automated systems. Existing methods also suffer from a representation gap: pixel-based models often lack precise control, while code-based synthesis limits intui...
Tianfu Wang, Leilei Ding, Ziyang Tao et al.· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
Turning a pretrained language model (LM) into a vision-language model (VLM) through multimodal fine-tuning often erodes its native language ability, a form of catastrophic forgetting that shows up even on text-only tasks. This loss is hard to undo with further fine-tuning, and existing remedies add adapters or alignmen...
Patrick Amadeus Irawan, Erland Hilman Fuadi, Shanu Kumar et al.· 0 citations
Current Large Language Models have achieved Olympiad-level logic, yet Vision-Language Models paradoxically falter on elementary spatial tasks like block counting. This capability mismatch reveals a critical ``spatial intelligence gap,'' where models fail to construct coherent 3D mental representations from 2D observati...
Shaoxiong Zhan, Yanlin Lai, Zheng Liu et al.· 0 citations
The rapid advancement of long-context vision language models (LCVLMs) has led to a significant expansion of their context windows. However, an extended context window does not guarantee the effective utilization of the context, posing a critical challenge for real-world applications. Current evaluations of such long-co...
Keyan Zhou, Zecheng Tang, Lingfeng Ming et al.· 0 citations
The proliferation of disinformation demands reliable and scalable fact-checking solutions. We present Dynamic Evidence-based FAct-checking with Multimodal Experts (DEFAME), a modular, zero-shot MLLM pipeline for open-domain, text-image claim verification. DEFAME operates in a six-stage process, dynamically selecting th...
Tobias Braun, Mark Rothermel, Marcus Rohrbach et al.· 0 citations
Multimodal Large Language Models have achieved remarkable progress in short-form audio-video understanding, yet long-form audio-video comprehension remains challenged by limited context windows and severe information redundancy. To address these bottlenecks, we propose AVOC, a framework for long-form audio-video unders...
Yijing Chen, Wenhui Tan, Xiaoyi Yu et al.· 0 citations
Vision-language models (VLMs) achieve strong zero-shot transferability but remain vulnerable to target-domain shifts at inference time. Test-time adaptation (TTA) offers a practical remedy, yet most existing VLM-TTA methods follow a prediction-side adaptation paradigm. They use test samples to adjust logits, prototypes...
Scientific figures often encode quantitative results that are not readily available in machine-readable form, making accurate plot digitization important for verifying and reusing published findings. Yet it remains unclear how accurately current models recover plotted values from real scientific figures, as existing be...
Yaohui Zhang, Binxu Li, Haoyi Duan et al.· 0 citations
Visual entailment (VE) asks whether an image supports, contradicts, or leaves undecided a textual hypothesis. Strong results come from fine-tuning large vision-language models on labelled data, while zero-shot and hybrid approaches remain far behind. A VE hypothesis often bundles several visual claims, yet existing zer...
Exploring how generative AI could make machine vision more accessible to businesses. The post GenEye in a Box: Making Machine Vision Something You Can Just Ask For appeared first on GPT-Lab.
Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity. The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.
Computer scientist, entrepreneur, and philanthropist will collaborate with the MIT Schwarzman College of Computing to advance AI and scientific discovery.