Jul 2026· Italian National Conference on Sensors· Vol 26· 0 citations· 51 references
Medicine
TL;DR
GazeHRNet is proposed, a head-centric reasoning framework for RGB-based gaze target detection that combines coarse spatial reasoning with fine-grained anisotropic heatmap prediction, enabling reliable target localization under cluttered scenes and varying head positions.
Abstract
Gaze target detection requires understanding where a person is looking by jointly reasoning about the gazer and the surrounding scene. While recent methods have benefited from powerful pretrained visual backbones, they often treat gaze prediction as a generic localization problem and overlook a key property of the task: the target should be interpreted in relation to the person’s head. This limits their ability to model direction, distance, and head-scene dependencies in a unified manner. We propose GazeHRNet, a head-centric reasoning framework for RGB-based gaze target detection. Instead of relying on absolute image coordinates or auxiliary geometric inputs, GazeHRNet represents the scene from the gazer’s perspective through Head-Centric Polar Encoding and organizes visual features by their spatial relevance to the head via Head-Aware Attention Routing. It further combines coarse spatial reasoning with fine-grained anisotropic heatmap prediction, enabling reliable target localization under cluttered scenes and varying head positions. Experiments on GazeFollow and VideoAttentionTarget show that GazeHRNet achieves 0.952 and 0.929 AUC with L2 distances of 0.102 and 0.103, respectively, using only RGB input and 3 M trainable parameters. Cross-dataset evaluation further demonstrates improved robustness and generalization across different scenes and subject distributions.
Gaze target estimation aims to infer the position of a person's gaze within a scene. Within mainstream design logic, multi-branch methods require extra supervision and annotations, while streamlined designs prioritize low-level visual saliency over true gaze intent. The former leads to a high annotation burden and hinders domain transfer, whereas the latter causes misalignment between predicted attention and actual gaze targets. To address this issue, we propose TextGaze, a unified cross-modal architecture that leverages a Large Vision-Language Model (LVLM) as scalable semantic guidance to balance the two design paradigms. The model extracts visual features from a frozen encoder and utilizes an LVLM to obtain gaze-aligned textual cues. We design a transformer-based fusion module with hierarchical text supervision to preserve task semantics. Lightweight decoding heads enable the joint prediction of gaze heatmaps and in-/out-of-frame status. We evaluate our method on four mainstream datasets, and the results show competitive performance across key metrics with robust cross-dataset generalisation without extra fine-tuning. Overall, we provide a streamlined alternative to traditional designs and highlight the potential of LVLMs as accessible auxiliary guidance for gaze estimation.
G3Ego, a graph-based framework for egocentric action understanding that uses gaze as a structural cue to identify action-relevant entities in the scene, achieves competitive performance compared with video-based approaches and consistently improves Macro-F1 under class-imbalanced evaluation, while avoiding reliance on computationally expensive video pretraining.
Marko Haralović, Akash Ramakrishnan, E. T. Martínez· 0 citations
General gaze following aims to infer the region attended to by a person within a natural scene and offers a computational perspective on visual attention and social scene understanding. Existing direction-guided methods usually reduce gaze direction to a single vector or spatial mask. This deterministic treatment can obscure directional ambiguity, discard coexisting candidate directions, and propagate early estimation errors to gaze target localization when head cues are weak or multiple targets are plausible. To address these limitations, we propose an uncertainty-aware framework based on Circular Direction Distribution Learning (CDDL) and Probabilistic Gaze Geometry Modeling (PGGM). CDDL represents gaze direction as a 72-bin circular probability distribution under von Mises soft supervision, thereby preserving neighboring directional hypotheses before target localization. Rather than predicting a target space distribution directly, PGGM aggregates the direction distribution into 18 groups and projects the retained hypotheses into cone-like spatial probability fields, allowing spatial tolerance to expand with distance while preserving a channel-wise direction-to-region representation. The full probability volume is fused with the RGB scene image at the input level and processed by a ResNet-50-FPN network for gaze heatmap prediction. Controlled experiments on GazeFollow demonstrate competitive localization and support the contribution of direction-distribution learning and channel-wise geometric projection. Zero-shot transfer, dataset-level uncertainty and failure analyses, efficiency measurements, and qualitative results further indicate interpretable behavior together with domain-dependent and computational trade-offs.
Yanzhao Li, Jin Li, Xiaona Zhang et al.· Journal of Eye Movement Rese...· 0 citations
Gaze Object Prediction (GOP) aims to localize and recognize the objects humans attend to, a task crucial for understanding human-centric interactions. However, existing methods are typically trained under a closed-vocabulary paradigm with a fixed label space and evaluated on scene-specific datasets, limiting their applicability to real-world scenarios where gaze targets often follow a long-tail distribution or belong to unseen categories. To address this gap, we introduce Diverse Scenes for Gaze object prediction (DiSG), a benchmark containing 86 in-the-wild categories that facilitates the evaluation of Open-Vocabulary GOP (OVGOP). Building on DiSG, we propose a framework that leverages text-driven object discovery to localize potential gaze candidates, with a gaze-guided selection module to pinpoint the intended target from the candidate objects. Furthermore, to better capture semantic knowledge across diverse in-the-wild categories, we introduce Gradient-Informed Selection Tuning (GIST) to selectively update parameters most relevant to a given class vocabulary. Extensive experiments demonstrate that our proposed model performs effectively in open-vocabulary settings and also outperforms existing methods in the conventional closed-vocabulary setting. The benchmark and code is available at https://github.com/sensniu/ovgop.
This survey presents a critical review of VLMs for egocentric video understanding, tracing the progression from conventional recognition architectures to multimodal foundation models and embodied systems, and examines how first-person perception and multimodal foundation models support wearable assistance, robot skill learning, human-to-robot transfer, and embodied decision making.
ANCHOR, a target-centric paradigm designed to decode gaze-anchored social intent by modeling the joint distribution of visual attention and latent implicit relations, is proposed, providing the first quantitative evidence that implicit social hierarchies can be robustly disentangled and learned directly from static gaze patterns.
Yuqi Hou, Zhuo Chen, Han Hu et al.· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.