Logo image
A Microscaling Multi-Mode Gain-Cell Computing-in-Memory Macro for Advanced AI Edge Device
期刊文章   同儕審查

A Microscaling Multi-Mode Gain-Cell Computing-in-Memory Macro for Advanced AI Edge Device

Jen-Chun Tien, Ping-Chun Wu, Win-San Khwa, Ashwin Sanjay Lele, Jian-Wei Su, Chiao-Yen Cheng, Jun-Ming Hsu, Yu-Chen Chen, Le-Jung Hsieh, Jyun-Cheng Bai, …
IEEE journal of solid-state circuits, 卷.61(1), 頁碼.211-224
01/2026
Web of Science ID: WOS:001596989200001

摘要

Accuracy Adders Artificial intelligence (AI) Artificial neural networks Common Information Model (computing) Common Information Model (electricity) computing-in-memory (CIM) Data transfer gain cell (GC) Hardware In-memory computing microscaling (MX) multiply-and-accumulate (MAC) Computer Architecture Energy Consumption
The microscaling (MX) format is an emerging data representation that quantizes high-bitwidth floating-point (FP) values into low-bitwidth FP-like values with a shared-scale (SS) exponent. When implemented with computing-in-memory (CIM), MX allows an attractive tradeoff between accuracy and hardware efficiency for specific neural network (NN) workloads. This work presents the first multi-mode gain-cell (GC) CIM macro capable of processing MX, integer (INT), and FP multiply-and-accumulate (MAC) operations with high energy efficiency (EEF) and area efficiency (AEF). The proposed macro employs four important innovations: 1) a multi-mode input processing unit (M2-IPU) with SS-variance-aware MAC flow (SS-VAF) for SS processing and SS alignment within the CIM macro to reduce system-to-CIM data transfer and compute energy; 2) a pattern-aware hybrid adder tree (PAH-ADT), which improves EEF and AEF by optimizing the common input patterns; 3) an accumulation-aware data flow (A2-DF) that adjusts the write path based on accumulation size to reduce data transfer energy; and 4) a 3.xT GC, which boosts data retention time (DRT) by increasing parasitic capacitance without additional area overhead. A 16-nm FinFET 216-kb MX-INT-FP multi-mode GC-CIM macro achieved 133.5 TFLOPS/W for MX-MAC with MXINT8 input, MXINT8 weight, and FP32 output; and 91.9 TFLOPS/W for FP-MAC with BF16 input, BF16 weight, and FP32 output.

相關連結

詳細資料

Logo image