Logo image
Layer Sensitivity Aware CNN Quantization for Resource Constrained Edge Devices
Conference paper

Layer Sensitivity Aware CNN Quantization for Resource Constrained Edge Devices

Alptekin Vardar, Li Zhang, Susu Hu, Saiyam Bherulal Jain, Shaown Mojumder, Nellie Laleni, Ashish Shrivastava, Sourav De and Thomas Kampfe
2022 9th International Conference on Soft Computing and Machine Intelligence, ISCMI 2022, pp.26-30
2022

Abstract

Deep learning In-Memory Computing Keras Quantization Artificial Intelligence Computational Mathematics Control and Optimization Numerical Analysis
Edge computing is rapidly becoming the defacto method for AI applications. However, the latency area and energy continue to be the main bottlenecks. To solve this problem, a hardware-aware approach has to be adopted. Quantizing the activations vastly reduces the number of Multiply-Accumulate (MAC) operations, resulting in with better latency and energy consumption while quantizing the weights decreases both memory footprint and the number of MAC operations, also helping with area reduction. In this paper, it is demonstrated that adapting an intra-layer mixed quantization training technique for both weights and activations, concerning layer sensitivities, in a Resnet-20 architecture with CIFAR-10 data set, a memory reduction of 73% can be achieved compared to even its all 8bits counterpart while sacrificing only around 2.3% accuracy. Moreover, it is demonstrated that, depending on the needs of the application, the balance between accuracy and resource usage can easily be arranged using different mixed-quantization schemes.

Metrics

1 Record Views

Details

Logo image