TACO

Description
TACO benchmarks generalizable bimanual tool–action–object understanding through human demonstrations with synchronized egocentric RGB-D and twelve-view third-person RGB recordings. Its V1 release contains 2,317 motion sequences spanning 151 tool–action–object triplets and 206 high-resolution object models, with hand and object poses, meshes, camera parameters, and automatic 2D segmentations. Egocentric video is available for 2,212 sequences. The dataset supports compositional action recognition, hand–object motion forecasting, and cooperative grasp synthesis.
License
CC BY 4.0 for the work, as declared in the official dataset instructions.*
* Double-check the publisher's current license and usage terms before using this dataset.