Skip to content

Author

Junjie Chan

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Open access 2026

STFPD: Vision–Language Guided Tri-Modal Fusion for Trust-Worthy Pedestrian Detection

Accurate and trustworthy pedestrian detection is a foundational requirement for safe autonomous systems, particularly as environmental conditions become increasingly complex. However, conventional detection systems frequently struggle to maintain reliability under low-visibility conditions and are highly susceptible to false alarms caused by human-like interference, limiting their real-world trustworthiness. To overcome these perceptual bottlenecks, we propose STFPD, a novel interpretable semantic-target fusion strategy via vision-language self-supervision designed for multispectral pedestrian detection. Our framework introduces a tri-modal feature fusion methodology that integrates complementary evidence from RGB images, thermal signals, and text. Specifically, we utilize a self-supervised vision-language model to generate textual representations, explicitly modeling semantic context without requiring manual annotations. During the fusion phase, we exploit both parallel and cross-channel similarities among the three modalities, extracting effective representations through dynamic spatial sampling. Crucially, to ensure explainability and verifiable reasoning, we introduce a mask generation sub-network in the refinement phase, which enhances feature contrast and provides precise evidence localization. Extensive evaluations demonstrate that STFPD achieves outstanding performance and robust reliability. On the KAIST benchmark, STFPD achieves a competitive all-day miss rate of 4.45%, outperforming the Faster R-CNN baseline by 14.41% and the previous best model by 1.24%, while recording an exceptional night-time miss rate of only 3.33%. Furthermore, on the LLVIP dataset, it attains an Average Precision (AP) of 72.4% and an AP50 of 98.1%, yielding a substantial 6.1% AP improvement over existing methods. Qualitative 3D saliency visualizations further confirm that STFPD provides high-contrast target localization with near-zero false alarms, suggesting promising potential for deployment in safety-oriented intelligent transportation systems.

Junjie Chan, Yangfan Luo · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.