Aug 2026· Frontiers in Plant Science· Vol 17· 0 citations· 44 references
Medicine
TL;DR
BLAP provides a lightweight solution that balances accuracy, efficiency, and interpretability for crop disease diagnosis in resource-constrained settings and may be extended to parameter-efficient fine-tuning of visionlanguage models in other domain-specific applications.
Abstract
Introduction Applying general-purpose vision-language models (VLMs) to crop disease diagnosis presents three critical bottlenecks: reliance on large-scale annotated data, the high computational cost of full finetuning, and existing adaptation methods designed mainly for discriminative classification without sufficient visual-linguistic interaction for generative diagnosis. Methods We propose BLAP, an adaptive multi-scale visual prompt fine-tuning framework built upon BLIP-2. BLAP introduces an adaptive visual prompt fusion module (APFM) with learnable prompt vectors and a gating mechanism, together with a multi-scale pyramid feature fusion module (PFM). All BLIP-2 backbone parameters are frozen, and only 0.11% of the model parameters are optimized. Results On a few-shot dataset comprising 990 images from 11 crops and 33 disease categories, BLAP achieved 92.78% recognition accuracy, outperforming the BLIP-2+LoRA baseline by 21.67 percentage points. BLEU-4 and ROUGE-L scores reached 0.6507 and 0.7184, respectively, while inference latency increased by only 2.15%. Discussion BLAP provides a lightweight solution that balances accuracy, efficiency, and interpretability for crop disease diagnosis in resource-constrained settings. The proposed dynamic prompt fusion and multiscale pyramid adaptation strategy may also be extended to parameter-efficient fine-tuning of visionlanguage models in other domain-specific applications.
: Two major problems with bean production that result in yield loss are bean rust and angular leaf spot. Though the older procedures require specialist knowledge, timely disease detection promotes production. Combining Pyramid Vision Transformer (PVT) and Group Context Aware Depthwise Shuffle Network (GCADSN) demonstra...
S. Sumayya, Varikuti Mamatha, Gurram Harinath et al.· Proceedings of the 1st Inter...· 0 citations
Few-shot fine-grained image classification (FS-FGIC) aims to distinguish visually similar subcategories with only a handful of labeled examples, posing significant challenges due to subtle inter-class differences and large intra-class variations. Existing methods often fail to fully leverage complementary information f...
Jinyu Wang, Bing-Xin Xu, Weiguo Pan et al.· International Conference on...· 0 citations
The proposed AgriX-SENet framework effectively combines high classification performance with model interpretability, addressing a key limitation of existing CNN-based plant disease detection systems and making it a promising solution for scalable agricultural diagnostics.
S. Raj, Prashant Johri, Vishwadeepak Singh Baghela et al.· Frontiers in Plant Science· 2 citations
DC-FEN, a MobileNetV3-based design that models spatial-token relations and channel interactions in parallel and injects them through gated residual fusion is introduced and shows that adding intermediate transfer constraints does not guarantee a stronger student.
Xin Lei, Yonghuai Liu, Ardhendu Behera et al.· Agriculture· 0 citations
Industrial visual sensing systems play a critical role in automated quality inspection. However, deploying data-driven perception models on industrial visual sensors faces significant bottlenecks due to the extreme scarcity of annotated anomaly samples and the pronounced long-tailed distribution of defect types. These...