Logo image
An Adaptive Scheme of Threshold Adjustment for Dynamic Sparsity Extraction of Self-Attention Network
會議論文

An Adaptive Scheme of Threshold Adjustment for Dynamic Sparsity Extraction of Self-Attention Network

Yong-Lun Xiao, Chia-Wei Chang, Cheng-Ting Shih, Jing-Jia Liou, Chih-Tsun Huang, Yao-Hua Chen 和 Juin-Min Lu
IEEE International Conference on Artificial Intelligence Circuits and Systems (Online), 頁碼.1-5
IEEE
2025 IEEE 7th International Conference on Artificial Intelligence Circuits and Systems (AICAS) (Bordeaux, France, 28/04/2025–30/04/2025)
28/04/2025

摘要

Accuracy Adaptive systems Adaptive Threshold Dynamic scheduling Estimation Hardware acceleration Large language models LLM Monitoring Real-time systems Sparsity Transformers Energy Consumption
Large Language Models (LLMs) and transformers have become highly successful across various domains. However, they are notorious for their quadratic computational complexity, which increases with sequence length. To mitigate this, dynamic sparsity techniques skip near-zero-output patterns based on low-precision estimations. Values below static thresholds are pruned, reducing energy consumption and improving computation speed.By utilizing low-cost estimations of minor threshold adjustments, we continuously monitor and fine-tune the pruning strategy to avoid overly aggressive pruning. Experimental results demonstrate that the proposed adaptive threshold method provides an average accuracy improvement of 0.15%, along with an average additional 8.95% computational sparsity across the SQuAD v1.1, v2, SST-2, and MRPC datasets.

相關連結

指標

1 檢視次數

詳細資料

Logo image