SPOT is introduced, a structured prompting framework that augments open-source VLMs with spatial reasoning abilities for scene graph generation with minimal training, and achieves competitive or superior relation prediction compared to large proprietary models.
The Cross-Modal Dual-Path Generator (CM-DPG), a model for robust open-vocabulary SGG that enhances object-level semantic representations through joint visual-textual encoding and improves relational reasoning using complementary visual and geometric cues, is proposed.
Minghao Zou, Qingtian Zeng, Shangkun Liu et al.· 0 citations
Scene graphs provide structured visual knowledge that can enhance multimodal large language models (MLLMs) for tasks like visual question answering and image captioning. However, existing approaches inject scene graphs as deterministic, hard prompts, ignoring the inherent uncertainty in visual perception, leading to overconfident and sometimes hallucinated outputs. We propose Probabilistic Scene Graph Prompting (PSGP), a framework that models scene graph generation as a distribution over plausible graphs and encodes this uncertainty into soft, continuous prompt tokens that condition the MLLM. By propagating perceptual uncertainty from detection to language generation, PSGP produces more accurate, faithful, and better-calibrated responses, especially in ambiguous visual scenarios. Experiments on GQA, Visual Spatial Reasoning, and a new Ambiguous-GQA benchmark show that PSGP outperforms strong baselines: including LLaVA-1.5, BLIP-2, and SG-LLaVA, in accuracy, faithfulness, and calibration, while maintaining computational efficiency. Our work establishes a principled pathway toward uncertainty-aware, structured multimodal intelligence.
Dynamic scene graphs (DSGs) capture spatio-temporal interactions across videos as $\langle$subject, predicate, object$\rangle$ triplets, and underpin downstream tasks such as video captioning, video question answering, and action analysis. However, end-to-end dynamic scene graph generation (DSGG) methods are closed-set: they recognize only objects and predicates from a fixed training vocabulary and struggle with the long-tailed distribution of rare concepts, severely limiting their real-world applicability. Existing open-vocabulary models typically inherit pretrained large language models, resulting in multi-stage training and inference with substantial cost. We introduce OvDSGG, the first end-to-end framework for open-vocabulary DSGG. OvDSGG builds on top of an open-vocabulary Spatial Backbone and a Temporal Backbone; we further propose a Triplet Feature Extraction Module that bridges them, and a Visual-Language Alignment Module that preserves open-vocabulary recognition by learning an adaptive decision boundary in the joint visual-language feature space, without expensive knowledge distillation in existing methods. We further introduce a rigorous open-vocabulary DSGG benchmark adapted from Action Genome, with disjoint Base/Novel splits for both objects and predicates. OvDSGG significantly outperforms open-vocabulary baselines across all metrics, with zero-shot Recall@$K$ scores 10.0--20.4 percentage point higher than the next-best baseline, while on closed-set DSGG remaining competitive with state-of-the-art models. Code and benchmark are publicly available at https://github.com/jhelsby/OvDSGG/.
John Helsby, Yi Yang, B. Rosenhahn et al.· 0 citations
This work extends its previous work in scene graph generation to integrate image captioning, incorporating large language models to interpret scene graph structures and produce high-quality captions encompassing complex actions and relationships.
Anfel Amirat, N. Baha, Lamine Benrais· Signal, Image and Video Proc...· 0 citations
OVIP-SG is presented, a unified framework for instance-preserving semantic mapping, functional scene partitioning, and language-guided small, fine-grained object retrieval that outperforms ConceptGraphs under a unified evaluation protocol on Replica.
Tianjing Hao, Hai-Yu Lan, Ang Li et al.· 0 citations
KNA-SG, a framework for constructing open-vocabulary 3D scene graphs from RGB sequences with explicit keyframe–node associations, is proposed and Experimental results show that KNA-SG outperforms existing methods on open-vocabulary 3D semantic segmentation and 3D object grounding tasks.
Yang Xu, Wen-ku Shi, Jing Xing et al.· Technologies· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.