Logo image
A 22nm Nonvolatile AI-Edge Processor with 21.4TFLOPS/W using 47.25Mb Lossless-Compressed-Computing STT-MRAM Near-Memory-Compute Macro
Conference paper

A 22nm Nonvolatile AI-Edge Processor with 21.4TFLOPS/W using 47.25Mb Lossless-Compressed-Computing STT-MRAM Near-Memory-Compute Macro

De-Qi You, Win-San Khwa, Jui-Jen Wu, Chuan-Jia Jhang, Guan-Yi Lin, Po-Jung Chen, Ting-Chien Chiu, Fang-Yi Chen, Andrew Lee, Yu-Cheng Hung, …
2024 IEEE Symposium on VLSI Technology and Circuits (VLSI Technology and Circuits)
06/2024

Abstract

Program processors;Accuracy;Nonvolatile memory;Artificial neural networks;Very large scale integration;In-memory computing;Real-time systems

Battery-powered AI-edge processors require short wakeup-to-response latency (TWR) and high energy efficiency (EF) for accurate real-time inference. This necessitates high-capacity nvCIM macros to store floating-point (FP) neural network (NN) data (e.g., BF16) and perform MAC operations with short latency (TCD) and high EF. This paper presents an STT - MRAM nvCIM macro with lossless compression computation, a near-far aware readout scheme, and system-level CIM-friendly hybrid weight mapping. The proposed 22nm nonvolatile processor (nvProcessor) with 47.25-Mb STT-MRAM nvCIM achieved high EF (21.4TFLOPS/W), short TWR(428.58μs), and high macro-level EF (27.6TFLOPS/W). Keywords: multiply-and-accumulate (MAC), nonvolatile memory (NVM), nonvolatile compute-in-memory (nvCIM).

Metrics

1 Record Views

Details

Logo image