Logo image
Fine-Grained Alignment in Vision-and-Language Navigation Through Bayesian Optimization
會議論文集

Fine-Grained Alignment in Vision-and-Language Navigation Through Bayesian Optimization

Y. Song, M. Gianni, C. Yang, K. Lin, T.-C. Chiu, A. Nguyen 和 C.-Y. Lee
Lecture Notes in Computer Science, 卷.16815 LNCS, 頁碼.311-326
2027

摘要

Cross Modality Alignment Representation Learning for Robotics Alignment Computer vision Contrastive Learning Embeddings Natural language processing systems Navigation Robotics Visual servoing 'current 3-D environments Bayesian optimization Cross modality Cross modality alignment Embeddings Fine grained Natural languages Navigation tasks Representation learning for robotic Visual languages
This paper addresses the challenge of fine-grained alignment in Vision-and-Language Navigation (VLN) tasks, where robots navigate realistic 3D environments based on natural language instructions. Current approaches use contrastive learning to align language with visual trajectory sequences. Nevertheless, they encounter difficulties with fine-grained vision negatives. To enhance cross-modal embeddings, we introduce a novel Bayesian Optimization-based adversarial optimization framework for creating fine-grained contrastive vision samples. To validate the proposed methodology, we conduct a series of experiments to assess the effectiveness of the enriched embeddings on fine-grained vision negatives. Experiments on the R2R offers valuable insights into the role of negative sample quality. Our source code and trained models are available at: https://github.com/HuskyKingdom/FGVLN. © The Author(s), under exclusive license to Springer Nature Switzerland AG 2027.

檔案與連結 (1)

url
https://www.scopus.com/inward/record.uri?eid=2-s2.0-105047072882&doi=10.1007%2f978-3-032-31663-9_21&partnerID=40&md5=96d863aa7eaee7bcc6e6c00b9ae6f1ba檢視

相關連結

指標

1 檢視次數

詳細資料

Logo image