Skip to content

Boosting scene captioning with pairwise semantic spatial reasoning and contextualized large language model integration

Aug 2026 · Signal, Image and Video Processing · Vol 20 · 0 citations · 40 references
Computer Science

TL;DR

This work extends its previous work in scene graph generation to integrate image captioning, incorporating large language models to interpret scene graph structures and produce high-quality captions encompassing complex actions and relationships.

View source

Similar papers

Conference Jul 2026

3D vision-language question answering with explicit scene graphs and local topology priors

TA-LMM, a 3D visual question answering method built on explicit scene graphs and local topological priors, is proposed, suggesting that explicit local topological priors can improve scene consistency in 3D visual question answering.

Kaixin Wu, Kunlin Zhou, Boxin Li et al. · 0 citations
Preprint Aug 2026

Modeling Scientific Experiment Scenes: Dataset and Model

The Cross-Modal Dual-Path Generator (CM-DPG), a model for robust open-vocabulary SGG that enhances object-level semantic representations through joint visual-textual encoding and improves relational reasoning using complementary visual and geometric cues, is proposed.

Minghao Zou, Qingtian Zeng, Shangkun Liu et al. · 0 citations

SPOT: Structured Prompting with Object-centric Tokens for open-world scene graphs

SPOT is introduced, a structured prompting framework that augments open-source VLMs with spatial reasoning abilities for scene graph generation with minimal training, and achieves competitive or superior relation prediction compared to large proprietary models.

Unknown authors · 0 citations
Conference Jul 2026

Multimodal Context-Enriched Visual Representation Learning for Enhanced Vision–Language Image Captioning

Image captioning models can produce rapid Sentences, without visual relationships, or insert non-existing plausible objects. A common cause is to compress image evidence into visual symbols that carry a weak neighbourhood context. The multimodal context-enhanced visual representation learning framework (MCVRL) addresses this error mode by adding local neighbourhood descriptors, global scene tokens, prefix-conditioned visual doors and adaptive contextual corrections be-fore caption decoding. The encoder is trained with cross-entropy and contrast terms for image–text alignment. MSCOCO 2014’s Karpathy test classification results show that BLEU-4, METEOR, CIDEr and SPICE are more powerful captioning bases. The best configuration received a score of 1.352 of the CIDEr compared to 1.308 of BLIP-2 in the same evaluation protocol. The results of the ablation show that the most important contribution is the visual feature enriched by the context, followed by crossmodal gating and adaptive contextual attention. Qualitative examples show that objects with hallucinations are fewer and that spatial relationships are better recovered.

E. Divya, Johnson Kolluri, Kiran Siripuri · 0 citations
Open access Aug 2026

Image captioning using a transformer with topic–word semantic modeling and multimodal feature fusion

Despite the recent advances in Transformer-based image captioning models, reliance on implicit semantic representations and the lack of integration of topic-level and word-level semantic information with appearance and geometric features remain challenging. To address these limitations, we propose a semantic modeling framework consisting of two components: Topic Embedding and Estimation Predictor (TEEP) and Topic-Related Representative Words (TRRW). TEEP explicitly predicts image-level semantic topics, whereas TRRW generates the top-10 representative words corresponding to each predicted topic. The predicted topics and their corresponding representative words are jointly fused to establish a unified semantic representation. This representation provides a more informative and accurate context, thereby improving caption generation accuracy. Moreover, we propose an extended Transformer architecture that jointly fuses semantic representations generated by TEEP and TRRW with appearance and geometric features. By concurrently modeling semantic context, visual attributes, and spatial relationships, the proposed approach produces more descriptive and semantically consistent captions. Additionally, we propose incorporating the exponential moving average (EMA) into cross-entropy training. This strategy enhances training stability, thereby improving performance across evaluation metrics. Extensive experiments on the MS-COCO dataset demonstrate that our model shows competitive performance compared to several approaches across standard evaluation metrics.

Ali Abdullah Yahya, Majjed Al-Qatf, Ammar Hawbani et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.