Logo image
Falcon: A Fused-Layer Accelerator With Layer-Wise Hybrid Inference Flow for Computational Imaging CNNs
期刊文章   同儕審查

Falcon: A Fused-Layer Accelerator With Layer-Wise Hybrid Inference Flow for Computational Imaging CNNs

Yong-Tai Chen, Yen-Ting Chiu, Hao-Jiun Tu 和 Chao-Tsung Huang
IEEE transactions on very large scale integration (VLSI) systems, 卷.33(3), 頁碼.720-732
01/03/2025
Web of Science ID: WOS:001351479600001

摘要

Computer Science, Hardware & Architecture Engineering, Electrical & Electronic Science & Technology Computer Science Engineering Technology
Computational imaging (CI) has advanced significantly due to the use of convolutional neural networks (CNNs). Its edge deployment relies on layer fusion to offload the monstrous external memory access (EMA) of feature maps, necessitating the handling of overlapped features either through reusing or recomputing them. Depending on how the boundary-handling strategy is organized, the induced computing complexity and EMA can be optimized. However, state-of-the-art CI accelerators primarily apply homogeneous inference flows, which employ a single overlap-handling strategy throughout the fused layers, limiting their ability to balance computation and data access. In this article, we explore layer-wise optimization in fused-layer CNNs by exploiting hybrid-strategy inference flows and devising a corresponding computing architecture. We categorize layer- wise strategies and put forward a layer-wise hybrid inference flow (LHIF) to integrate their advantages, and we propose an optimization procedure that explicitly analyzes essential figures of merit (FoMs), including throughput, EMA, and energy efficiency. Furthermore, we develop a high-throughput accelerator- Falcon-to efficiently support LHIF under massive parallelism, especially with a time-division-multiplexing (TDM) buffer interface that enables seamless access to feature maps stored in an interleaved manner. Layout results show that the accelerator, delivering 41 TOPS with 1.5 MB of feature-map buffers, supports LHIF while increasing the die area by only 1.4% and power consumption by only 0.7%. Extensive simulations are conducted to demonstrate the versatility of LHIF in working scenarios at operational, design, and system levels. Compared with using homogeneous inference flows, the proposed LHIF achieves Pareto optimality with up to 2.28x higher throughput and 3.5x lower EMA.

相關連結

指標

1 檢視次數

InCites亮點

本研究成果之相關指標(擷取自 InCites Benchmarking & Analytics)

引用書目主題
4 Electrical Engineering, Electronics & Computer Science
4.101 Signal & Image Processing
4.101.1178 Image Restoration
Web Of Science研究領域
Computer Science, Hardware & Architecture
Engineering, Electrical & Electronic
ESI研究領域
Engineering

詳細資料

Logo image