Dense image captioning is critical for vision-language pretraining and for text-to-image generation, but scaling expert-quality annotations is prohibitively expensive. While synthesizing captions from strong vision-language models (VLMs) is a practical alternative, supervised distillation often yields limited output di...
Tzu-Heng Huang, Sirajul Salekin, Javier Movellan et al.· 0 citations
Vision Transformers (ViTs) often degrade under distribution shifts because they rely on spurious correlations, such as background cues, rather than semantically meaningful features. Existing regularization methods, typically relying on simple foreground-background masks, which fail to capture the fine-grained semantic...
This work pretrain decoder-only models with approximately GPT-2 small and medium sizes on the FineWeb-Edu dataset and introduces Deep Delta Learning (DDL), which applies the delta rule over network depth.
Yi-Fan Zhang, Yi-Feng Liu, Mengdi Wang et al.· arXiv.org· 8 citations· ⚡2
Out-of-domain (OOD) robustness is challenging to achieve in real-world computer vision, especially in unsupervised domain adaptation scenarios, where shifts in image background, style, and acquisition instruments often degrade model performance. Generic augmentations show inconsistent gains under such shifts, whereas d...
Ruoqi Wang, Haitao Wang, Shaojie Guo et al.· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
Vision-Language-Action models (VLAs) represent a significant frontier in embodied intelligence, aiming to bridge digital knowledge with physical-world interaction. Despite their remarkable performance, foundational VLAs are hindered by the prohibitive computational and data demands inherent to their large-scale archite...
Zhaoshu Yu, Bo Wang, Pengpeng Zeng et al.· 0 citations
A new model family (CortexMAE) trained using the masked autoencoder framework is introduced, and in this setting the CortexMAE family outperforms prior models by a large margin, and the first open evaluation suite (Brainmarks) for fMRI foundation models is released.
Connor Lane, Mihir Tripathy, L. Murali et al.· arXiv.org· 5 citations· ⚡1
Recent video generation models have achieved remarkable progress and are now deployed in film, social media production, and advertising. Beyond their creative potential, such models also hold promise as world simulators for robotics and embodied decision making. Despite strong advances, current approaches still struggl...
David Romero, Ariana Bermudez, Viacheslav Iablochnikov et al.· 0 citations
Vision Graph Neural Networks (ViGs) have demonstrated promising performance in image recognition tasks against Convolutional Neural Networks (CNNs) and Vision Transformers (ViTs). An essential part of the ViG framework is the node-neighbor feature aggregation method. Although various graph convolution methods, such as...
Hakan Emre Gedik, Andrew Martin, Mustafa Munir et al.· 0 citations
Spiking Neural Networks (SNNs) have recently received increasing attention in both computational neuroscience and artificial intelligence owing to their potential for energy-efficient computation and reduced memory requirements. Despite these advantages, improving adversarial robustness in SNNs (particularly for vision...
We propose Graph Consistency Regularization (GCR), a novel framework that injects relational graph structures, derived from model predictions, into the learning process to promote class-aware, semantically meaningful feature representations. Functioning as a form of self-prompting, GCR enables the model to refine its i...
Xi Ding, Lei Wang, Piotr Koniusz et al.· 0 citations
Experimental results show that PRI achieves up to 10x faster inversion than standard Dense Model Inversion (DMI) and 2x faster than SMI, while consistently outperforming SMI in accuracy and matching the performance of DMI.
As demands for resource efficiency and safety in modern neural networks intensify, substantial research effort has gone into model compression and adversarial robustness. Yet despite progress on each in isolation, a systematic understanding of how compressibility shapes robustness remains elusive. In this paper, we dev...
Melih Barsbey, Ant\^onio H. Ribeiro, Umut \c{S}im\c{s}ekli et al.· 0 citations
Exploring how generative AI could make machine vision more accessible to businesses. The post GenEye in a Box: Making Machine Vision Something You Can Just Ask For appeared first on GPT-Lab.
Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity. The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.
Computer scientist, entrepreneur, and philanthropist will collaborate with the MIT Schwarzman College of Computing to advance AI and scientific discovery.