Abstract
The invention provides a quantization method and system based on an operational circuit architecture in a memory, and the method comprises the steps: dividing a quantization weight into grouping quantization weights according to a grouping value, and dividing an input excitation function into grouping excitation functions according to the grouping value; in the multiply accumulation step, performing multiply accumulation operation on the grouping quantization weight and the grouping excitation function to generate convolution output; enabling the convolution quantization step to quantize the convolution output into a quantized convolution output in accordance with the convolution target bit. In the convolution merging step, outputting the quantized convolution to execute partial sum operation according to the grouping value so as to generate an output excitation function. Therefore, better weight parameters can be learned by grouping and pairing, considering hardware limitation, and matching classified distribution of the analog-to-digital converter and a specific quantization method with the robust property of the deep neural network.