Robot Learning Manipulation Action Plans by "Watching" Unconstrained Videos from the World Wide Web

Author:

Yang Yezhou,Li Yi,Fermuller Cornelia,Aloimonos Yiannis

Abstract

In order to advance action generation and creation in robots beyond simple learned schemas we need computational tools that allow us to automatically interpret and represent human actions. This paper presents a system that learns manipulation action plans by processing unconstrained videos from the World Wide Web. Its goal is to robustly generate the sequence of atomic actions of seen longer actions in video in order to acquire knowledge for robots. The lower level of the system consists of two convolutional neural network (CNN) based recognition modules, one for classifying the hand grasp type and the other for object recognition. The higher level is a probabilistic manipulation action grammar based parsing module that aims at generating visual sentences for robot manipulation. Experiments conducted on a publicly available unconstrained video dataset show that the system is able to learn manipulation actions by ``watching'' unconstrained videos with high accuracy.

Publisher

Association for the Advancement of Artificial Intelligence (AAAI)

Subject

General Medicine

Cited by 22 articles. 订阅此论文施引文献 订阅此论文施引文献,注册后可以免费订阅5篇论文的施引文献,订阅后可以查看论文全部施引文献

1. Why is that a Good or Not a Good Frying Pan? – Knowledge Representation for Functions of Objects and Tools for Design Understanding, Improvement, and Generation;2023 IEEE Symposium Series on Computational Intelligence (SSCI);2023-12-05

2. Scene-Aware Activity Program Generation with Language Guidance;ACM Transactions on Graphics;2023-12-05

3. Performance Comparison of Teleoperation Interfaces for Ultra-Lightweight Anthropomorphic Arms;2023 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS);2023-10-01

4. Tuning-less Object Naming with a Foundation Model;2023 World Symposium on Digital Intelligence for Systems and Machines (DISA);2023-09-21

5. Signs of Language: Embodied Sign Language Fingerspelling Acquisition from Demonstrations for Human-Robot Interaction;2023 32nd IEEE International Conference on Robot and Human Interactive Communication (RO-MAN);2023-08-28

同舟云学术

1.学者识别学者识别

2.学术分析学术分析

3.人才评估人才评估

"同舟云学术"是以全球学者为主线,采集、加工和组织学术论文而形成的新型学术文献查询和分析系统,可以对全球学者进行文献检索和人才价值评估。用户可以通过关注某些学科领域的顶尖人物而持续追踪该领域的学科进展和研究前沿。经过近期的数据扩容,当前同舟云学术共收录了国内外主流学术期刊6万余种,收集的期刊论文及会议论文总量共计约1.5亿篇,并以每天添加12000余篇中外论文的速度递增。我们也可以为用户提供个性化、定制化的学者数据。欢迎来电咨询!咨询电话:010-8811{复制后删除}0370

www.globalauthorid.com

TOP

Copyright © 2019-2024 北京同舟云网络信息技术有限公司
京公网安备11010802033243号  京ICP备18003416号-3