Skip to content
Open access

Dependable digital spiking neuromorphic hardware

Abstract

Spiking neural networks (SNNs) are a biologically inspired class of machine learning models in which information is encoded and processed as sparse, asynchronous spike events distributed over time. When mapped onto digital neuromorphic hardware, these networks perform event-driven computation at very low power, making them attractive for edge intelligence, autonomous systems, and other low power applications. However, as neuromorphic systems move toward real-world deployment, they inherit the reliability problems of modern silicon. Single-event upsets, stuck-at defects, process variation, aging, and timing errors can corrupt the arithmetic and control logic of a spiking neuron. Since computation in an SNN is temporally encoded, such a fault can delay a spike, suppress a spike that should have occurred, or introduce a spurious one. The temporal encoding attenuates errors, because the membrane update applies its decay factor to an injected error. Reliability analysis for neuromorphic hardware therefore differs from that of conventional deep-learning accelerators, and it requires error-detection and fault-tolerance methods developed specifically for spiking systems. This dissertation develops four such methods, organized along two directions. Detection and tolerance are developed to establish a method, and that method's cost. All four are implemented and evaluated on QUANTISENC, an open-source digital spiking neuromorphic core, executing standard image-classification workloads. The first method performs concurrent (online) error detection and isolation using hierarchical, model-based monitoring. A software-based monitor observes the presynaptic and postsynaptic spike trains produced by hardware-mapped neurons and compares them against the nominal behavior predicted by lightweight, tree-based machine learning models. An inexpensive system-level monitor continuously checks the aggregate behavior of the network, and neuron-level monitoring is invoked only when a system-level discrepancy is detected. To keep neuron-level monitoring affordable, critical neurons are identified from the magnitude of their synaptic weights and grouped into clusters, with a single predictive model representing each cluster. Fault-injection experiments on MNIST, Fashion-MNIST, and SVHN detect injected faults at rates around 88% at both levels, with model storage under 2 MB and model inference times below a millisecond on the host processor. The quantity reported is the detection rate over the injected population. The second method asks what that monitoring needs to pay for. Revisiting the first pipeline on its own fault-injection data shows that the learned per-neuron model is less computationally efficient: a model-free residual fed to a sequential change detector matches the tree-ensemble residual to within ±0.001 AU-ROC on every data set and metric. This permits a cheap-first, expensive-on-demand cascade, in which an inexpensive rate and regularity chart runs on every neuron and the costly Victor-Purpura and van Rossum spike distances are computed only on flagged neurons, reducing the modeled online cost of the sensitive metric by about 13.6x at an end-to-end recall of 0.95. The saving is derived from an operation-count model, the confirmation stage is measured directly on MNIST. The third method addresses fault tolerance through the selective hardening of the neuron datapath. Circuit-level faults are injected into the adder, multiplier, comparator, and refractory counter of the leaky integrate-and-fire neuron, and a criticality analysis combining synaptic-weight strength with input-dependent firing activity identifies the neurons and components whose failures distort spike-timing. The analysis identifies the multiplier, and in particular its most significant bit region, as the dominant contributor to spike timing error. Applying triple modular redundancy only to MSB slice within the critical neurons reduces the average inter-spike-interval deviation under randomized fault injection from 1.65% to 1.49%, at an area overhead of about 1% and a negligible power overhead. This is the only comparison in the chapter in which both conditions use the same fault-sampling procedure. The fourth method allocates redundancy from how long each storage element actually retains a captured fault. A matched IF, obtained by removing only the leaky path from the design, isolates the contribution of the membrane leak from the threshold reset. Across 288 paired fault descriptors on three data sets, output spike-train masking is 86.5% with leak and 75.3% without, and residual spike corruption is 5.9 times smaller. Bit-level campaigns show three register-dependent criticality patterns consistent with differences in state-update and persistence semantics. On a held-out descriptor set, a register-transfer-level built 13-bit selected protection mask raises masking from 59 to 89 of 96 injections and lowers critical silent data corruption from five events to one, at an overhead of 1.56% generic mapped cells, whereas triplicating the twelve low-order membrane bits recovers none of the baseline critical events at 1.44%. On shared configuration state, triplication lowers critical silent data corruption from 36.4% to 13.1%, and adding periodic repair lowers it to 6.1%, over 198 paired observations per variant. Together the four methods span the reliability of a neuromorphic system. The first two detect faults during operation and then remove monitoring cost that does not buy detection; the second two tolerate the faults. The evidence throughout is simulation and register-transfer-level fault injection on image-classification workloads, with faults injected at chosen sites and times.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.