This framework establishes a reproducible baseline for low-latency edge-AI deployment in automotive cabin monitoring, autonomous robotics, and assistive healthcare, establishing a hardware-aware reference baseline for future Neural Processing Unit (NPU) and neuromorphic architectures.
Abstract
Neuromorphic vision sensors offer significant advantages for real-time embedded perception due to their microsecond latency, high temporal resolution, wide dynamic range, and low power consumption. However, deploying multi-stage event perception pipelines onto edge hardware remains fundamentally constrained by memory, throughput, and execution bottlenecks on resource-limited silicon. In this work, we introduce an end-to-end, real-time perception pipeline deployed on a Raspberry Pi 5 platform, establishing a hardware-aware reference baseline for future Neural Processing Unit (NPU) and neuromorphic architectures. Our system integrates a custom attention-enhanced YOLOv8-nano model for joint person and face detection, a parameter-free ByteTrack multi-object tracking framework, and a lightweight kinematic feature extraction engine for behavioral triage and velocity-based activity classification. To evaluate localized spatial trade-offs—particularly critical for in-cabin automotive application domains such as Driver and Occupant Monitoring Systems (DMS/OMS) we systematically quantify the sensor bias configurations and optical geometries (76° vs. 104° Fields of View) across detection accuracy, temporal consistency, face region extraction, and kinematic activity mapping. The detection backbones are fine-tuned on a custom indoor event dataset and exported using static ONNX execution graphs optimized for edge deployment. Operating strictly on sparse spatiotemporal event representations without intermediate RGB frame reconstruction, our pipeline maintains robust, privacy-responsive performance under extreme lighting and motion dynamics, sustaining 19–24 FPS end-to-end throughput. By providing complete quantitative and qualitative benchmarks, this framework establishes a reproducible baseline for low-latency edge-AI deployment in automotive cabin monitoring, autonomous robotics, and assistive healthcare. The project details, fine-tuned models, empirical results, and curated event datasets are publicly available at http://mali-farooq.github.io/NeuroVision
Real-time visual perception on resource-constrained embedded hardware must reconcile computational economy, low latency, and dependable sensing accuracy within tight power and cost envelopes. This paper reports on the Smart Navigation Device (SND), a wearable assistive perception system that performs object detection, distance ranging, sensor fusion, and speech feedback entirely on a Raspberry Pi 4B without any cloud dependency. At the core of the system is the Cascaded Detection-Ranging Fusion (CDRF) framework, a four-stage pipeline that couples the lightweight YOLO11n detector with ultrasonic time-of-flight ranging through confidence-guided detection acceptance, spatial-zone partitioning, dominant-object association, and adaptive suppression of redundant announcements. The hardware-software co-design keeps every processing stage — image capture, neural inference, ranging, fusion, and text-to-speech synthesis — local to the device, eliminating transmission latency, removing a major privacy exposure, and preserving operability where network connectivity is unreliable or absent. The framework was evaluated across 169 controlled trials spanning three obstacle categories — person, chair, and laptop — at distances from 0.30 m to 4.20 m. Detection rates of 90.4%, 92.9%, and 77.0% were obtained for the three classes respectively, with mean absolute ranging errors of 2.11 cm, 1.60 cm, and 1.79 cm. Agreement between the ultrasonic estimate and ground-truth distance was excellent (Pearson r = 0.9995, p < 10⁻²¹⁹), and Bland–Altman analysis revealed a small systematic bias of −1.10 cm (95% limits of agreement: −7.88 cm to 5.68 cm). A chi-square test indicated a statistically meaningful distance-dependent decline in laptop-class detection reliability. Benchmarked against previously reported wearable travel aids, the SND achieves comparable or better detection reliability than low-cost ultrasonic-only alternatives while additionally providing object identity, and does so at a fraction of the hardware cost and without any of the connectivity dependencies of cloud-assisted alternatives — positioning it as a reproducible, statistically grounded, and economically accessible baseline for future assistive-perception research.
S. R. Katke, Utkarsha Pacharaney· International journal of com...· 0 citations
A compact and configurable event-driven autoencoder that efficiently compresses neuromorphic data while preserving essential spatiotemporal structure for downstream inference and demonstrates the potential of compact event-driven models to advance environmentally conscious, low-power AI systems for high-speed perception in autonomous, mobile, and embedded computing environments.
Riadul Islam, Joey Mulé, Dhandeep Challagundla et al.· 0 citations
Neural decoding systems often achieve strong offline performance, yet many fail to translate that performance into reliable real-time operation due to fragmented pipelines, limited reproducibility, and insufficient visibility into computational costs. This paper presents a real-time graphical user interface (GUI) platform for neural signal transmission that integrates preprocessing, optional feature extraction, model inference, visualization, and continuous resource monitoring within a single operational workflow. The platform supports three execution modes: ESN-only, LSTM-only, and a hybrid ESN LSTM pipeline, enabling systematic comparison of accuracy, latency, and hardware footprint under consistent conditions. Evaluation is conducted using the PhysioNet EEGMMIDB v1.0.0 (motor imagery EEG from nine subjects), with configurable sampling rates and channel selection to reflect practical deployment constraints. To ensure deterministic behavior, the platform incorporates controlled random seeding and structured logging of model and runtime configurations, improving repeatability across sessions and devices. The system also reports per-stage latency, CPU/RAM utilization, and profiling overhead to help practitioners identify bottlenecks and quantify trade-offs. Experimental observations indicate reduced blocking delays and minimal monitoring overhead, supporting stable long-session operation. Overall, the proposed platform bridges the gap between algorithmic neural decoding research and deployable, testable real-time systems by combining reproducible modeling, multi-mode execution, and transparent performance profiling in an integrated GUI environment. The primary contribution is an engineering and reproducibility framework rather than a new learning algorithm‚ as the decoding models used are well established independent of these papers․ The novelty lies in the deterministic‚ fully profiled assembly of the models into a single deployable real-time setting․ To this end‚ decoding performance is given as a secondary but fully detailed metric (per-subject‚ with confidence intervals)‚ along with system-level latency‚ resource requirements‚ and stability measures․
Laith Hussein Jasim Alzubaidi, Yaghoub Farjami, Mohsen Akbarpour Beni· International journal of com...· 0 citations
Neuromorphic computing is emerging as a promising paradigm for sustainable edge intelligence by enabling event‐driven, low‐latency and energy‐aware computation close to sensors. However, the field remains fragmented across device technologies, mixed‐signal and digital hardware platforms, spiking neural network models, software frameworks, event‐stream processing tools, interoperability standards and benchmarking practices. This review provides a cross‐layer synthesis of contemporary neuromorphic computing platforms with particular emphasis on their relevance to scalable and sustainable edge deployment. The article organizes the neuromorphic ecosystem into interconnected layers spanning materials and devices, hardware architectures, software and interoperability tools, benchmark resources and application domains. Mixed‐signal platforms are analysed in terms of analog efficiency, accelerated neural dynamics, biological plausibility, calibration requirements, variability and reproducibility challenges. Digital neuromorphic processors are examined with respect to programmability, deterministic execution, routing fabrics, memory organization, software integration and deployment readiness. The review further discusses software frameworks, simulators, event‐data libraries, hardware‐mapping tools and intermediate representations that support model development, portability and cross‐platform evaluation. A central finding is that neuromorphic systems cannot be compared meaningfully using isolated metrics such as neuron count, chip power, latency, or throughput alone; fair evaluation requires explicit reporting of workload, event rate, model topology, mapping strategy, software stack, measurement boundary and deployment context. Accordingly, the article proposes a benchmarking and reporting perspective for neuromorphic edge intelligence that links accuracy, latency, energy efficiency, robustness and reproducibility. Thus, this review clarifies current progress, unresolved challenges and future directions for sustainable edge sensing, robotics, healthcare monitoring, smart infrastructure, industrial automation and distributed intelligent systems.
Processing in memory (PIM) offers a compelling pathway to overcome the data movement bottleneck in modern AI and data-centric systems. This work introduces MITRA, a reconfigurable magnetic tunnel junction (MTJ)-based in-memory architecture that leverages stochastic computing (SC) to implement a broad class of transcendental and nonlinear functions directly within memory. By combining stochastic bit-stream processing with compact finite-state-machines (FSMs) embedded in MTJ-FinFET logic-in-memory structures, the proposed design achieves low-latency and power-efficient computation without external datapaths, unlike the binary counterparts. Circuit-level simulations in 14-nm FinFET technology verify correct state transitions, stable stochastic outputs, and predictable power profiles. Extensive evaluations demonstrate high accuracy even with short bit-streams. We further integrate the design into a neural-network classifier and develop an FSM-aware training strategy that compensates for approximation errors, achieving up to 96.9% classification accuracy on the UCI Optical Digit benchmark. Overall, MITRA provides a compact, reconfigurable platform for nonlinear processing in next-generation edge AI systems.
Farzad Razi, M. Moghadam, M. Najafi et al.· International Symposium on L...· 0 citations
FlexSpIM, a digital CIM architecture supporting arbitrary operand resolution and shape within a unified storage for weights and neuron states, is introduced, enabling a layer-level hybrid weight- and output-stationary dataflow, maximizing operand reuse and reducing costly on- and off-chip data movement during SNN execution.
Nicolas Chauvaux, Adrian Kneip, Charlotte Frenkel· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.