Logo image
The Algorithm and VLSI Architecture of High-Throughput and Highly Efficient Tensor Decomposition Engine
期刊文章

The Algorithm and VLSI Architecture of High-Throughput and Highly Efficient Tensor Decomposition Engine

婷羽 蔡, 中安 沈, 宗嶙 吳元豪 黃
IEEE Transactions on Circuits and Systems I: Regular Papers, 卷.71(7), 頁碼.3134-3145
07/2024

摘要

Tensors;Matrix Decomposition;VLSI;Circuits

Tensor decomposition is critical for compressing data and extracting key features in novel high-dimensional signal processing systems. However, due to the enormous amount of data and the highly complicated computations, designing an efficient tensor decomposition processor is very challenging. This paper presents the algorithm and VLSI architecture design of a low-latency and high-throughput tensor decomposition processor. A parallel higher-order orthogonal iteration (P-HOOI) algorithm is proposed where multiple updated matrices are computed concurrently. Thus, a low-latency tensor decomposition is achieved. Furthermore, a novel VLSI architecture is presented so that the efficiency of the component utilization is improved and the hardware complexity is greatly reduced. Therefore, the proposed tensor decomposition processor enhances the processing throughput with minimum employment of hardware components. Performance evaluations based on the post-layout estimations in the ASIC flow and based on the FPGA platform are reported in this paper. Compared with state-of-the-art designs in the literature the proposed tensor decomposition engine greatly enhances the throughput and hardware efficiency.

相關連結

指標

1 檢視次數

詳細資料

Logo image