Recent advances in multimodal foundation models yield strong performance on static perception and reasoning benchmarks, yet such evaluations largely overlook a central aspect of intelligence: acting competently in dynamic environments over extended time horizons. We introduce PlaySuite, a large-scale benchmark for eval...
Dheeraj Varghese, Anna Vettoruzzo, Walter Simoncini et al.· 0 citations
We introduce Loop Flow Transformers (LiFT), a family of looped generative models that scales computation by repeatedly applying a shared Diffusion Transformer (DiT) core, with only light changes to the standard architecture. Rather than asking every recurrent step for the final prediction, LiFT trains each step with a...
Mohammad Mahdi Derakhshani, P. Curvo, G. Burghouts et al.· 0 citations
Generalized Category Discovery aims to recognize known categories while identifying novel ones within unlabeled data. Existing methods, typically based on self-supervision and contrastive learning, often struggle to capture fine-grained distinctions, relying on superficial visual cues rather than the intrinsic attribut...
Sarah Rastegar, Mina Ghadimi Atigh, Pascal Mettes et al.· 0 citations
This work proposes FineHOI, a zero-shot HOI framework that explicitly models interactions from dense patch-level features, and introduces an Adaptive Part-Level Attention module that decomposes humans and objects into semantically coherent parts via unsupervised clustering, and re-weights them based on their interactio...
Francesco Tonini, Lorenzo Vaquero, Mohammad Mahdi Derakhshani et al.· 0 citations
This paper proposes TWIST, a twin-expert stepwise tuning module that modifies the decoder of the language model using one frozen module pre-trained on image understanding tasks and another learnable one for visual grounding tasks, which allows the MLLM to retain previously learned knowledge and skills, while acquiring...
AritraBhowmik, MohammadMahdiDerakhshani, Dennis C. Koelma et al.· 0 citations
SnapViT: single-shot network approximation for pruned Vision Transformers is introduced, a new post-pretraining structured pruning method that enables elastic inference across a continuum of compute budgets, and a self-supervised importance scoring mechanism that maintains strong performance without requiring retrainin...
Walter Simoncini, Michael Dorkenwald, Tijmen Blankevoort et al.· 0 citations
It is shown in a variety of image classification settings and on several datasets, that quasibinary classifiers are considerably better in classification settings where regular binary and softmax classifiers suffer, including zero-label and multi-label classification.
Shuai Liao, E. Gavves, Changyong Oh et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.