Logo image
BICEP: Exploiting Bitline Inversion for Efficient Operation-Unit-Based Compute-in-Memory Architecture: No Retraining Needed!
Conference paper

BICEP: Exploiting Bitline Inversion for Efficient Operation-Unit-Based Compute-in-Memory Architecture: No Retraining Needed!

Yun-Chen Lo, Chia-Chun Wang and Ren-Shuo Liu
Proceedings - IEEE International Conference on Computer Design (ICCD) 2023, pp.531-534
2023

Abstract

Hardware and Architecture Electrical and Electronic Engineering
Compute-in-memory (CIM) architecture is promising for its in-situ analog computing ability. However, one practical constraint for CIM architectures is the limited number of activated rows in an operation Unit (OU). OU-based CIM architecture only activates a subgroup of memory cells to ensure a large signal margin and enough consideration of non-ideal device/circuit effect, which pays the cost of lowered computing throughput. In short, the OU-based CIM architectures suffer from array underutilization to ensure high accuracy.This work proposes a novel architecture, BICEP, which exploits bitline inversion technique to enlarge the OU size without the need to prune, approximate, and retrain. More specifically, the key contributions of this work are threefold: 1) We propose a bitline inversion scheme, which guarantees more than 2× larger OU size without affecting the numerical results and the ADC resolution. The key insight is to selectively apply code inversion on heavy bitlines to constrain their MAC outputs and compensate using low-cost compensation units. We mathematically prove that the proposal can be applied to both single- and multi-level cells (SLC and MLC). 2) We propose an inversion-aware weight swapping scheme, which swaps the weight order to maximize the OU size exploiting bitline inversion. 3) We propose weight order propagation to enable inversion-aware weight swapping without storage overheads. The extensive experiments on ImageNet classification tasks demonstrate that this work outperforms state-of-the-art OU-based CIM architecture (DL-RSIM) by up to 2.06× speedup and 1.97× energy efficiency.

Metrics

1 Record Views

Details

Logo image