Logo image
Adaptive Row-wise Attention Score Pruning and Compensation for Attention Acceleration
會議論文

Adaptive Row-wise Attention Score Pruning and Compensation for Attention Acceleration

Chia-Chun Wang, Yu-Lin Lin 和 Ren-Shuo Liu
IEEE International Symposium on Circuits and Systems proceedings, 頁碼.1924-1928
IEEE
2026 IEEE International Symposium on Circuits and Systems (ISCAS) (Shanghai, China, 24/05/2026–28/05/2026)
24/05/2026

摘要

Accuracy AI accelerators Conferences Design methodology Modeling neural network Printing Transformers
Transformer models have become a cornerstone of modern AI applications, with the attention layer playing a critical role within each transformer block. As the computational demands of these models continue to grow, considerable research has focused on accelerating attention computation while maintaining acceptable accuracy. Among these approaches, low-precision top-k prediction is commonly used to speed up the attention layer by selecting only the largest values for the softmax operation.Despite its effectiveness, the top-k prediction method still leaves room for improvement. In this paper, we propose two complementary techniques to accelerate attention computation while preserving accuracy. First, we introduce Two-stage Adaptive Row-wise Attention Score Pruning, which carefully prunes unimportant elements in each row. Second, we present Pruning-aware Softmax Compensation, which mitigates the numerical error caused by pruning and further improves accuracy.Experimental results demonstrate that our methods achieve higher accuracy and greater computational savings compared to the conventional top-k prediction approach.

相關連結

指標

1 檢視次數

詳細資料

Logo image