TASTE-Rob

Description
TASTE-Rob contains 100,856 fixed-view, 1080p egocentric hand–object interaction videos—about 9 million frames—each under eight seconds and aligned to one language instruction. Its 75,389 single-hand and 25,467 double-hand clips cover kitchens, bedrooms, dining and office tables, with actions such as picking, placing, pushing, pouring, cleaning, and drawer use. The dataset targets task-oriented hand–object video generation and imitation learning, using scene and object diversity plus precise action-language alignment to improve manipulation generalization.
License
CC BY-NC 4.0.*
* Double-check the publisher's current license and usage terms before using this dataset.