Skip to content
Open access

ConvSNN: an energy-efficient hybrid convolutional-spiking neural network for wrist-worn stress and affective state recognition

Sep 2026 · Scientific Reports · 0 citations

Abstract

Continuous stress monitoring on wrist-worn devices matters for real-time affective computing, digital health, and personalized well-being, yet it remains difficult because a wearable model must operate with few sensors, little compute, and a tight energy budget. Chest-mounted systems can draw on high signal-to-noise signals such as ECG, EMG, and respiration, but a wrist device captures only blood-volume pulse, electrodermal activity, skin temperature, and accelerometry, which makes accurate and efficient stress recognition harder. This study presents ConvSNN, a hybrid convolutional–spiking neural network for multimodal affective recognition from wrist-worn physiological signals. ConvSNN pairs a compact four-stage 1D convolutional backbone with a time-to-first-spike (TTFS) classification head that reads aggregated rate-coded and time-coded spike statistics to keep predictions stable. It also adds a spiking confidence gate for uncertainty-aware window acceptance and a temperature-scaling step for probability calibration. We evaluate ConvSNN on the WESAD benchmark using only the six wrist channels of the Empatica E4 wearable (blood-volume pulse, electrodermal activity, skin temperature, and three-axis accelerometry). Under random 85/15 split evaluation it reaches a mean accuracy of 98.07% with a macro-F1 of 0.977 across five random seeds. Under strict leave-one-subject-out (LOSO) cross-validation, ConvSNN reaches a mean accuracy of 94.0% with a macro-F1 of 0.932 over five seeds, which confirms strong subject-independent generalization and amounts to a 4.1-percentage-point drop from the random-split protocol. Temperature scaling lowers the expected calibration error from 0.149 to 0.017 with no change in accuracy, and the spiking confidence gate opens a 5.9-percentage-point accuracy gap between accepted and rejected windows, which supports principled abstention on low-confidence inputs. Based on operation-count emulation, ConvSNN has an estimated per-inference energy of 48.31  $$\mu$$ J at 5.70 ms latency, which points to low-power real-time use on smartwatch-class devices once it is validated on embedded hardware. Taken together, these results show that hybrid convolutional–spiking models can strike an effective balance among recognition performance, subject-independent robustness, calibrated confidence estimation, and energy-efficient deployment for continuous wrist-worn affective monitoring.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.