To address the challenges of low detection accuracy, high miss rates, and limited model lightweightness arising from dense fruit distribution, foliage occlusion, and small fruit size during citrus fruit localisation and recognition in complex orchard environments, this study proposes a lightweight small-object citrus fruit detection model based on an improved YOLO11 architecture, termed SOCD-YOLO (YOLO for Small Object Citrus Detection). Firstly, a DLTBlock is designed by integrating element-wise multiplication with a Triplet Attention mechanism to reconstruct the C3k2 module, thereby enhancing the nonlinear fusion of high-dimensional features. This design effectively suppresses interference from occluding foliage and complex backgrounds, and improving the robustness of fruit target recognition. Secondly, the traditional downsampling operation is replaced with the ADown module, which effectively alleviates information loss during feature propagation for small-object features, while simultaneously reducing model parameters and computing complexity, thus enhancing small-object recognition accuracy. Finally, a lightweight P2FP structure is constructed to further enhance the model’s capability in identifying small and densely distributed objects, while significantly reducing the missed detection rate. Experimental findings indicate that, on the CitDet dataset, the proposed model achievesimprovements of 4.3%, 6.3%, and 5.6% in Precision, Recall, and mAP, respectively, while reducing the parameter count and model size to 1.5 M and 3.5 MB. Moreover, the FPS reaches 103.5 frames/s. On the Tomato and PASCAL VOC 2007 datasets, overall performance is consistently improved. Compared to existing object detection models, SOCD-YOLO demonstrates enhanced performance in terms of citrus fruit detection accuracy and robustness, providing a valuable reference for artificial intelligence based real-time fruit detection and position measurement in densely occluded environments.
Yiran Zhao, Jianbo Lu· Measurement science and tech...· 0 citations
Semantic 3D Gaussians provide a compact representation for 3D semantic occupancy prediction by rendering semantic primitives into a voxel volume under voxel-wise supervision. Recent methods have improved the modeling ability and efficiency of this representation through more flexible primitive shapes, geometry-guided initialization, and progressive densification. However, these advances mainly determine how primitives are represented, initialized, or added, and do not explicitly address how to select the most useful Gaussians when their total number must be limited to control memory and computation. This imbalance creates an allocation bottleneck: redundant Gaussians remain in simple regions, while difficult regions receive insufficient semantic support. We propose the Semantic Gaussian Allocation Transformer (SAGFormer), which uses Gaussian attributes and local geometric-semantic features to score candidates and select a fixed final Gaussian set. Experiments on nuScenes-SurroundOcc and SSCBench-KITTI-360 show that SAGFormer improves occupancy prediction under the evaluated protocols and yields more semantically consistent and better-utilized Gaussian representations. Under similar final counts and raw coverage, it reduces semantic mixing, strengthens class-consistent voxel support, and produces fewer unused Gaussians. The results indicate that explicit capacity allocation is a useful complement to Gaussian refinement for semantic occupancy prediction.
Kanglin Ning, Yiran Zhao, Wenrui Li et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.