Logo image
GViG: low-resource generative visual grounding using prompt-based language modeling for visual question answering
期刊文章   開放取用(OA)

GViG: low-resource generative visual grounding using prompt-based language modeling for visual question answering

Yi-Ting Li, Ying-Jia LinHung-Yu Kao
International Journal of Data Science and Analytics
2025

摘要

Multimodal machine learning Prompt tuning Visual grounding Visual question answering Information Systems Modeling and Simulation Computer Science Applications Computational Theory and Mathematics Applied Mathematics
The WSDM 2023 Toloka VQA challenge introduces a new grounding-based visual question answering (GVQA) dataset, elevating multimodal task complexity. This challenge diverges from traditional VQA by requiring models to identify a bounding box in response to an image–question pair, aligning with visual grounding tasks. Existing VG approaches, when applied to GVQA, often necessitate external data or larger models for satisfactory results, leading to high computational demands. We approach this as a language modeling problem, utilizing prompt tuning with multiple state-of-the-art VQA models. Our method, operating solely on an NVIDIA RTX 3090 GPU without external data, secured third place in the challenge, achieving an Intersection over Union (IoU) of 75.658. Our model notably provides explainability between textual and visual data through its attention mechanism, offering insights into its decision-making process. This research demonstrates that high performance in GVQA can be achieved with minimal resources, enhancing understanding of model dynamics and paving the way for improved interpretability and efficiency.

檔案與連結 (1)

url
https://doi.org/10.1007/s41060-025-00835-7檢視
已出版(紀錄版本) 開放

相關連結

指標

1 檢視次數

詳細資料

Logo image