Kernel Quantization for Efficient Network Compression

Zhongzhi Yu; Yemin Shi

doi:10.1109/ACCESS.2022.3140773

IEEE Access (Jan 2022)

Kernel Quantization for Efficient Network Compression

Zhongzhi Yu,
Yemin Shi

Affiliations

Zhongzhi Yu: School of Computer Science, Peking University, Beijing, China
Yemin Shi: ORCiD; School of Computer Science, Peking University, Beijing, China

DOI: https://doi.org/10.1109/ACCESS.2022.3140773
Journal volume & issue: Vol. 10
pp. 4063 – 4071

Abstract

Read online

This paper presents a novel network compression framework, Kernel Quantization (KQ), targeting to efficiently convert any pre-trained full-precision convolutional neural network (CNN) model into a low-precision version without significant performance loss. Unlike existing methods struggling with weight bit-length, KQ has the potential in improving the compression ratio by considering the convolution kernel as the quantization unit. Inspired by the evolution from weight pruning to filter pruning, we propose to quantize in both kernel and weight level. Instead of representing each weight parameter with a low-bit index, we learn a kernel codebook and replace all kernels in the convolution layer with corresponding low-bit indexes. Thus, KQ can represent the weight tensor in the convolution layer with low-bit indexes and a kernel codebook with limited size, which enables KQ to achieve significant compression ratio. Then, we conduct a 6-bit parameter quantization on the kernel codebook to further reduce redundancy. Extensive experiments on the ImageNet classification task prove that KQ needs 1.05 and 1.62 bits on average in VGG and ResNet18, respectively, to represent each parameter in the convolution layer and achieves the state-of-the-art compression ratio with little accuracy loss.

Published in IEEE Access

ISSN: 2169-3536 (Online)
Publisher: IEEE
Country of publisher: United States
LCC subjects: Technology: Electrical engineering. Electronics. Nuclear engineering
Website: https://ieeexplore.ieee.org/xpl/RecentIssue.jsp?punumber=6287639

About the journal

Abstract

Keywords