Skip to content
Preprint

FineHOI: Part-Aware Dense Representations for Zero-Shot Human-Object Interaction Detection

Sep 2026 · 0 citations · 38 references
Computer Science

TL;DR

This work proposes FineHOI, a zero-shot HOI framework that explicitly models interactions from dense patch-level features, and introduces an Adaptive Part-Level Attention module that decomposes humans and objects into semantically coherent parts via unsupervised clustering, and re-weights them based on their interaction relevance.

Abstract

Human-Object Interaction (HOI) detection aims to localize humans and objects in images and classify their interactions. Zero-shot HOI focuses on recognizing interactions that are not observed during training, requiring models to generalize beyond seen verb-object compositions. Recent approaches leverage Vision-Language Models (VLMs), benefiting from rich semantic representations. However, they often rely on global or detector-centric features that compress interaction cues and hinder fine-grained spatial reasoning. To overcome this limitation, we propose FineHOI, a zero-shot HOI framework that explicitly models interactions from dense patch-level features. Our approach is motivated by the observation that human-object interactions are defined by localized spatial relationships, which are not preserved by global and detector-centric representations. To this end, we introduce an Adaptive Part-Level Attention module that decomposes humans and objects into semantically coherent parts via unsupervised clustering, and re-weights them based on their interaction relevance. These representations are then integrated through a Region-Aware Interaction Transformer that integrates part-aware and global features and produces the final HOI embedding. Extensive experiments demonstrate that FineHOI consistently outperforms existing zero-shot HOI methods, achieving particularly strong gains on unseen interactions. Code is available at https://github.com/francescotonini/fine-hoi.

View source

Similar papers

Preprint Oct 2026

Harnessing Multimodal Large Language Models for Training-Free Human-Object Interaction Detection

Human-object interaction (HOI) detection aims to localize human-object pairs and recognize their interactions. Traditional supervised methods perform strongly but rely on task-specific training. Recent multimodal large language models (MLLMs) offer a promising route to training-free HOI detection through their broad vi...

Zhao-Lin Cai, Hui-Yu Duan, Liu Yang et al. · 0 citations
Aug 2026

Text-guided zero-shot localization of unseen object categories

A text-guided Zero-Shot Localization framework for unseen object categories (ZSOL) for addressing the aforementioned challenges, which can be guided by prompt words to identify and localize unseen object categories in images by transferring localization knowledge learned from supervised base categories.

Jingjing Wang, Xing-Lin Piao, Zongzhi Gao et al. · 0 citations
Conference Aug 2026

MASF-Net: efficient linear attention guided few-shot fine-grained image recognition

Extensive experiments on three challenging fine-grained benchmarks demonstrate that MASFNet consistently outperforms state-of-the-art methods in both 5-way 1-shot and 5-way 5-shot settings.

Jinyu Wang, Bing-Xin Xu, Weiguo Pan et al. · 0 citations
Preprint Sep 2026

Exploiting Target Knowledge from MLLMs for Robust Few-Shot Segmentation

A novel framework that mines target knowledge using the strong reasoning capacity of Multimodal Large Language Models (MLLMs) and employs it to enhance FSS, which shows promising results and largely surpasses existing methods.

Yi-Jun Hu, Heng Fan, Li-Bo Zhang · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.