Abstract
Energy-and area-efficient acceleration solutions are critical for the continual development of ubiquitous, real-Time, and cross-domain artificial intelligence [1]-[7]. This motivates the active research on the SRAM-based computing-in-memory (CIM) that exploits the state-of-The-Art CMOS technology and massively parallel analog computing directly inside the memory array [1]-[4]. Although significant progress has been made in recent years in improving throughput [1-2], energy efficiency [1], [3], and area efficiency [1], [4], simultaneously achieving them in SRAM-CIM remains an unsolved problem. This is particularly challenging when accounting for the potential accuracy loss due to nonideality in analog computing and the inflexibility of CIM weight-stationary designs.