Logo image
Filter-based Deep-Compression with Global Average Pooling for Convolutional Networks
Conference paper

Filter-based Deep-Compression with Global Average Pooling for Convolutional Networks

Ting-Yun Hsiao, Yung-Chang Chang and Ching-Te Chiu
IEEE Workshop on Signal Processing Systems, SiPS: Design and Implementation, Vol.2018-October, pp.247-251
12/2018

Abstract

Deep Model Compression Global Average Pooling Pruning Quantization Truncated SVD Electrical and Electronic Engineering Signal Processing Applied Mathematics Hardware and Architecture
Deep neural networks are powerful, but using these networks is both memory and time consuming for their numerous parameters and large amounts of computation. There are many studies in compressing the models. on the parameter-level, one of the compressing methods is to spend lots of time performing the iterative process consisting of pruning weights and fine-tuning the models. on the bit-level, many studies use quantization to cut down on the number of need bits. Hence, we propose an efficient strategy to compress on the layers which are computation or memory consuming. We compress the model by adding the global average pooling, iteratively pruning on the filters with proposed order-deciding scheme to prune more efficiently, applying the truncated SVD to the fully-connected layer, and performing the two-stage quantization. Experiments on the VGG16 model show that we can reach a 60.9× compression ratio in off-line storage with about 0.848% and 0.1378% loss of accuracy on the top-1 and top-5 classification results with the validation dataset of ILSVRC2012.

Metrics

1 Record Views

Details

Logo image