Lightweight Optimization Strategies for Medical Neural Networks
Abstract
Deep learning has made significant advancements in medical data analysis, which has made it possible to advance the diagnostics process, prognostication, and therapeutic care through new methods. However, the usability of artificial neural networks in common clinical practice is still hindered by their excessive computational, memory, and power requirements, which do not make them applicable to the resource-stressed conditions of a portable imaging device, a wearable sensor, or an edge-computing system. The limitations are dealt with in this chapter by discussing ways of making medical neural networks more efficient. The question of interest is model-compression methods such as pruning, quantization, knowledge distillation, and neural architecture search, aimed at reducing the computational cost while preserving diagnostic quality. Also, the chapter argues for improvements in training time, which include transfer learning, self-supervised learning, and few-shot learning, which alleviate the challenges posed by small labeled medical datasets. The previous sections have discussed highly effective architectures, such as MobileNet, ShuffleNet and EfficientNet, that offer promising trade-offs in accuracy and efficiency in various medical uses including imaging, biomedical signal processing, and multimodal healthcare. Their practical usefulness is empirically demonstrated in case studies of quantised convolutional neural networks to screen X-rays, quantised networks to monitor ECGs, and distilled networks for mobile ultrasound devices. A comparison of the trade-offs between accuracy and efficiency, interpretability and ethical deployment confirms the need to use lightweight neural networks in real-time medical artificial intelligence systems that are deployable and scalable.