HOI4D

Description
HOI4D contains 2.4 million RGB-D egocentric video frames across 4,000 sequences collected by nine participants interacting with 800 object instances from 16 categories in 610 indoor rooms. It provides frame-wise action, motion, and panoptic segmentation labels, 3D hand poses, category-level object poses, reconstructed object meshes, and scene point clouds. Benchmarks cover dynamic point-cloud semantic segmentation, category-level object pose tracking, and egocentric action segmentation; the project also provides labeled test releases for action and semantic segmentation.
License
CC BY-NC 4.0; non-commercial use only.*
* Double-check the publisher's current license and usage terms before using this dataset.