Logo image
Dynamic Context Adapters: Efficiently Infusing History into Vision-and-Language Models
會議論文集

Dynamic Context Adapters: Efficiently Infusing History into Vision-and-Language Models

Y. Song, B.-J. Lin, J. Liu, T.-C. Chiu, A. Nguyen 和 C.-Y. Lee
Lecture Notes in Computer Science, 卷.16815 LNCS, 頁碼.109-124
2027

摘要

Efficient Deep Learning Learning for Vision Computer vision Decision making Deep learning Semantics Visual languages 'current Downstream applications Dynamic contexts Efficient deep learning Information loss Language model Learning for vision Memory consumption Modeling process Sequential decision making Computational efficiency
Historical context integration presents a fundamental challenge for Vision-Language Models (VLMs) in sequential decision-making tasks. Current VLMs process visual inputs independently, which creates critical limitations for downstream applications that require temporal understanding. Direct incorporation of historical frames into Transformer inputs produces quadratic attention complexity and excessive memory consumption. Existing approaches suffer from significant drawbacks: computational inflation or substantial information loss through temporal compression. To address these challenges, we introduce Dynamic Context Adapter (DCA), a novel context injection approach for pretrained VLMs. Our method employs fixed-size, dynamically compressed memory to preserve historical semantics without frame concatenation. DCA bridges static VLMs and recurrent policies and enables memory capabilities in pretrained models while maintaining computational efficiency. DCA achieves over 25% reduction in attention FLOPs and 13% memory savings while improving performance on long-horizon tasks. © The Author(s), under exclusive license to Springer Nature Switzerland AG 2027.

檔案與連結 (1)

url
https://www.scopus.com/inward/record.uri?eid=2-s2.0-105047080586&doi=10.1007%2f978-3-032-31663-9_8&partnerID=40&md5=eed42161d602f308013e66e2c056d005檢視

相關連結

指標

1 檢視次數

詳細資料

Logo image