Skip to content

Author

Hyun-Soo Kang

3 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Review Open access Jul 2026

Bird’s-Eye-View Road Occupancy Prediction for Autonomous Driving: A Survey of Representations, Methods, and Benchmarks

Bird’s-eye-view (BEV) perception has become the dominant paradigm for camera-centric scene understanding in autonomous driving, as well as road occupancy prediction, which involves the dense estimation of which regions of space are occupied and by what has emerged as its most expressive form. Between 2020 and 2026, the field underwent three overlapping transitions: from two-dimensional BEV semantic map segmentation to dense three-dimensional voxel-based 3D semantic occupancy, catalyzed by the 2022 industrial adoption of “occupancy networks,” and, most recently, to efficient, generative, and four-dimensional forecasting formulations. This survey organizes the literature along six orthogonal axes output representation, view-transformation mechanism, input modality, supervision paradigm, temporal scope, and efficiency strategy and uses the representation lineage as a primary spine connecting the 2020 BEV-segmentation works to the 2026 Gaussian and 4D frontier. Alongside the ego-centric mainstream, we review the parallel multi-view and infrastructure-side lineage from multi-view pedestrian occupancy to roadside traffic occupancy, which shares the BEV occupancy-map output and contributes generalization tools the ego-centric thread has yet to absorb. We review the canonical methods at each stage, summarize the standard datasets (CARLA, GMVD, MultiviewX, WildTrack, nuScenes, SemanticKITTI, Occ3D, OpenOccupancy) and evaluation metrics (MODA, mIoU, RayIoU, RayPQ), and consolidate reported results on the Occ3D-nuScenes benchmark into a single comparison. We close by identifying open problems in label efficiency, robustness, temporal forecasting, and deployment. Our intent is to bridge the historically separate BEV-segmentation and 3D-occupancy literatures within a single taxonomy.

Abdelrahman S. Heikal, Mostafa Farouk Senussi, Ahmed Salem et al. · 0 citations
Open access Aug 2026

Robust multi-class lumpy skin disease diagnosis for practical livestock applications

Lumpy skin disease (LSD) is a highly contagious viral infection that severely impacts cattle health and livestock productivity worldwide, necessitating reliable and automated diagnostic solutions. Recent advances in deep learning (DL) have demonstrated significant potential for image-based disease classification; however, achieving robust performance under real-world variability and class imbalance remains a critical challenge. In this study, we propose ConvNeXtLSD, a DL-based multi-class classification framework for automated cattle LSD diagnosis. The proposed model is built upon a modified ConvNeXt backbone architecture, consisting of a patchify stem followed by four hierarchical feature extraction stages composed of ConvNeXt blocks and downsampling layers, enabling effective multi-scale representation learning. To enhance class separability, the standard classification head is redesigned using feature flattening and layer normalization, which improves feature distribution stability and discriminative capability. Extensive experiments on a five-class cattle LSD dataset containing 8,014 images demonstrate that ConvNeXtLSD achieves superior classification performance, with an accuracy of 98.29%, a precision of 95.82%, an F1-score of 96.58%, and a Cohen’s Kappa coefficient of 97.51%. Moreover, the proposed framework maintains practical computational efficiency with 15.372 GFLOPs and a real-time inference speed of 90.18 FPS, demonstrating an effective balance between predictive performance and computational cost for real-world livestock disease diagnosis.

Mostafa Farouk Senussi, Abdelrahman S. Heikal, Asmaa Gamal Abdelbasset et al. · 0 citations
Open access Aug 2026

SETAS-VAD: Semantically Enriched Text-Aligned Scoring for Weakly Supervised Video Anomaly Detection

Weakly supervised video anomaly detection (WS-VAD) localizes anomalous events in untrimmed videos using only video-level annotations. While CLIP-based methods have advanced this task through vision–language alignment, widely adopted approaches construct text prototypes from short category-name prompts of at most five words, leaving the CLIP text encoder not fully exploited. We propose SETAS-VAD, which addresses this gap through a Category Semantic Alignment (CSA) loss function: for each anomaly category, a large language model generates multi-sentence descriptions covering complementary semantic aspects, encoded once offline into frozen prototype vectors. An InfoNCE contrastive objective pulls attention-weighted anomaly features toward ground-truth category prototypes at zero additional inference overhead (prototype generation and encoding are performed once offline as a preprocessing step, not at test time). Under fully reproducible conditions on UCF-Crime and XD-Violence, SETAS-VAD achieves state-of-the-art temporal localization (30.45% mAP on XD-Violence, 12.16% on UCF-Crime), with per-threshold gains increasing at stricter IoU values, indicating improved boundary precision rather than coarse detection sensitivity.

Mohamed Mahmoud, Mostafa Farouk Senussi, Mahmoud Abdalla et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.