📖 ABSTRACT/OVERVIEW
Neural network compression techniques including pruning, quantisation, and knowledge distillation enable the deployment of deep learning models on constrained embedded hardware, but comprehensive empirical characterisation of compression-accuracy trade-offs on affordable platforms within Nigerian embedded system development budgets is absent from the literature. This study empirically examined this performance gap by systematically evaluating compression techniques on three target platforms: Raspberry Pi 4 (35,000 naira), NVIDIA Jetson Nano (85,000 naira), and STM32H743 microcontroller (8,500 naira). Image classification (MobileNetV2 on ImageNet-mini), object detection (YOLOv5s on COCO-mini), and keyword spotting (DS-CNN on Google Speech Commands) were selected as representative workloads. Pruning (structured, 50 percent and 70 percent sparsity), INT8 quantisation using TensorFlow Lite and PyTorch QNNPACK, and knowledge distillation (student-teacher pairing with EfficientNet-B0 teacher) were applied individually and in combination. Results showed that combined INT8 quantisation and 50 percent structured pruning delivered the best latency reduction (3.8x on Raspberry Pi 4 for MobileNetV2) with an accuracy drop of only 2.1 percentage points. YOLOv5s INT8 on Jetson Nano achieved 31 fps with 3.4 mAP degradation (39.7 versus 43.1 mAP). DS-CNN keyword spotting on STM32H743 required knowledge distillation to maintain above 90 percent accuracy within the 1 MB SRAM constraint. The study provides the first Nigeria-contextualised compression benchmark report and recommends its adoption as a reference by TETFUND-funded embedded AI research projects.
Keywords: neural network compression, quantisation, pruning, embedded inference, affordable hardware
Need Complete Chapters of the Above Topic?
Get high-quality, Zero-AI research materials with current citations.
Request via WhatsApp 💬