Logo image
An 8b-Precision 8-Mb STT-MRAM Near-Memory-Compute Macro Using Weight-Feature and Input-Sparsity Aware Schemes for Energy-Efficient Edge AI Devices
期刊文章   同儕審查

An 8b-Precision 8-Mb STT-MRAM Near-Memory-Compute Macro Using Weight-Feature and Input-Sparsity Aware Schemes for Energy-Efficient Edge AI Devices

De-Qi You, Yen-Cheng Chiu, Win-San Khwa, Chung-Yuan Li, Fang-Ling Hsieh, Yu-An Chien, Chung-Chuan Lo, Ren-Shuo Liu, Chi-Cheng Hsieh, Kea-Tiong Tang, …
IEEE journal of solid-state circuits, 卷.59(1), 頁碼.219-230
01/2024
Web of Science ID: WOS:001092396700001

摘要

Artificial intelligence Artificial intelligence (AI) compute-in-memory (CIM) Edge AI Feature extraction In-memory computing multiply-and-accumulate (MAC) near-memory-compute (NMC) Neural networks Random access memory spin-transfer torque magnetic random access memory (STT-MRAM) Energy Consumption Energy Efficiency
Nonvolatile near-memory-compute (nvNMC) macros are promising candidates for edge artificial intelligence (AI) devices requiring high energy efficiency, short wakeup-to-compute latency, and robust inference accuracy with high precision of inputs (IN), weights ( W), and outputs (OUT). Nonetheless, the practical application of nvNMC macros is hindered by inherent design challenges: 1) high energy consumption in reading repetitious weight data, 2) low energy efficiency due to high bitstream toggling rate in digital multiply-and-accumulate (MAC) circuits, 3) narrow signal margin and high memory readout latency, and 4) redundant MAC computation in digital MAC circuits under input-stationary flow. In addressing these challenges, we developed a number of schemes based on the concept of system-circuit co-design, including 1) a weight-feature aware read (WFAR) scheme; 2) a toggling-aware weight-tuning (TAWT) scheme; 3) a differential charge-accumulating margin-enhanced voltage-sensing amplifier (DCME-VSA); and 4) an input-sparsity aware pre-computing unit (ISAPU). An 8-Mb spin-transfer torque magnetic random access memory (STT-MRAM) nvNMC macro fabricated using foundry-provided 22 nm STT-MRAM achieved read bandwidth of 436 GB/s, MAC computing latency of 20 ns, and energy efficiency of 53.6-190.2 TOPS/W when performing eight-bit input and eight-bit weight MAC operations with 26-bit output and 576 accumulations.

相關連結

詳細資料

Logo image