Logo image
ISSA: Architecting CNN Accelerators Using Input-Skippable, Set-Associative Computing-in-Memory
期刊文章   同儕審查

ISSA: Architecting CNN Accelerators Using Input-Skippable, Set-Associative Computing-in-Memory

Yun-Chen Lo, Jun-Shen Wu, Chia-Chun Wang, Yu-Chih Tsai, Chih-Chen Yeh, Wen-Chien Ting 和 Ren-Shuo Liu
IEEE transactions on computers, 卷.73(9), 頁碼.2136-2149
01/09/2024
Web of Science ID: WOS:001297926800006

摘要

Computer Science, Hardware & Architecture Engineering, Electrical & Electronic Science & Technology Computer Science Engineering Technology
Among several emerging architectures, computing in memory (CIM), which features in-situ analog computation, is a potential solution to the data movement bottleneck of the Von Neumann architecture for artificial intelligence (AI). Interestingly, more strengths of CIM significantly different from in-situ analog computation are not widely known yet. In this work, we point out that mutually stationary vectors (MSVs), which can be maximized by introducing associativity to CIM, are another inherent power unique to CIM. By MSVs, CIM exhibits significant freedom to dynamically vectorize the stored data (e.g., weights) to perform agile computation using the dynamically formed vectors. We have designed and realized an SA-CIM silicon prototype and corresponding architecture and acceleration schemes in the TSMC 28 nm process. More specifically, the contributions of this paper are fivefold: 1) We identify MSVs as new features that can be exploited to improve the current performance and energy challenges of the CIM-based hardware. 2) We propose SA-CIM to enhance MSVs (input-reordering flexibility) for skipping the zeros, small values, and sparse vectors. 3) We propose channel swapping to enhance the zero-skipping technique. 4) We propose a transposed systolic dataflow to efficiently conduct conv3 x 3 while being capable of exploiting input-skipping schemes. 5) We propose a design flow to search for optimal aggressive skipping scheme setups while satisfying the accuracy loss constraint. The proposed ISSA architecture improves the throughput by 1.91x to 2.97x speedup and the energy efficiency by 2.5x to 4.2x .

相關連結

指標

1 檢視次數

詳細資料

Logo image