Talk2Nav: Long-Range Vision-and-Language Navigation with Dual Attention and Spatial Memory-Reference-Cited by-同舟云学术

Talk2Nav: Long-Range Vision-and-Language Navigation with Dual Attention and Spatial Memory

Published:2020-08-31 Issue:1 Volume:129 Page:246-266
ISSN:0920-5691
Container-title:International Journal of Computer Vision
language:en
Short-container-title:Int J Comput Vis

Author:

Vasudevan Arun Balajee^ORCID,Dai Dengxin,Van Gool Luc

Abstract

AbstractThe role of robots in society keeps expanding, bringing with it the necessity of interacting and communicating with humans. In order to keep such interaction intuitive, we provide automatic wayfinding based on verbal navigational instructions. Our first contribution is the creation of a large-scale dataset with verbal navigation instructions. To this end, we have developed an interactive visual navigation environment based on Google Street View; we further design an annotation method to highlight mined anchor landmarks and local directions between them in order to help annotators formulate typical, human references to those. The annotation task was crowdsourced on the AMT platform, to construct a new Talk2Nav dataset with 10, 714 routes. Our second contribution is a new learning method. Inspired by spatial cognition research on the mental conceptualization of navigational instructions, we introduce a soft dual attention mechanism defined over the segmented language instructions to jointly extract two partial instructions—one for matching the next upcoming visual landmark and the other for matching the local directions to the next landmark. On the similar lines, we also introduce spatial memory scheme to encode the local directional transitions. Our work takes advantage of the advance in two lines of research: mental formalization of verbal navigational instructions and training neural network agents for automatic way finding. Extensive experiments show that our method significantly outperforms previous navigation methods. For demo video, dataset and code, please refer to our project page.

Funder

Toyota Motor Europe

Publisher

Springer Science and Business Media LLC

Subject

Artificial Intelligence,Computer Vision and Pattern Recognition,Software

Link

https://link.springer.com/content/pdf/10.1007/s11263-020-01374-3.pdf

Reference85 articles.

1. Agrawal, A., Lu, J., Antol, S., Mitchell, M., Zitnick, C. L., Parikh, D., et al. (2017). Vqa: Visual question answering. International Journal of Computer Vision, 123(1), 4–31.

2. Anderson, P., Chang, A., Chaplot, D. S., Dosovitskiy, A., Gupta, S., Koltun, V., Kosecka, J., Malik, J., Mottaghi, R., Savva, M., et al. (2018). On evaluation of embodied navigation agents. arXiv:1807.06757.

3. Anderson, P., He, X., Buehler, C., Teney, D., Johnson, M., Gould, S., & Zhang, L. (2018). Bottom-up and top-down attention for image captioning and visual question answering. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 6077–6086).

4. Anderson, P., Wu, Q., Teney, D., Bruce, J., Johnson, M., Sünderhauf, N., Reid, I., Gould, S., & van den Hengel, A. (2018). Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments. In Proceedings of the IEEE conference on computer vision and pattern recognition.

5. Andreas, J., Rohrbach, M., Darrell, T., & Klein, D. (2016). Learning to compose neural networks for question answering. arXiv preprint arXiv:1601.01705.

Cited by 24 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Embodied navigation with multi-modal information: A survey from tasks to methodology;Information Fusion;2024-12

2. ESceme: Vision-and-Language Navigation with Episodic Scene Memory;International Journal of Computer Vision;2024-07-26

3. Guided by the Way: The Role of On-the-route Objects and Scene Text in Enhancing Outdoor Navigation;2024 IEEE International Conference on Robotics and Automation (ICRA);2024-05-13

4. Trimodal Navigable Region Segmentation Model: Grounding Navigation Instructions in Urban Areas;IEEE Robotics and Automation Letters;2024-05

5. Room-Object Entity Prompting and Reasoning for Embodied Referring Expression;IEEE Transactions on Pattern Analysis and Machine Intelligence;2024-02