humanoidsdata.com

Search

Search companies, datasets, articles, and glossary terms for humanoids and embodied AI.

← All glossary terms

Data & collection

Data modality

A data modality is a distinct kind or channel of information characterised by how it is sensed, represented, or interpreted, such as RGB images, depth, language, joint state, actions, audio, tactile measurements, or force–torque signals. A multimodal robot dataset contains more than one modality, but usefulness depends on their alignment, semantics, and relevance to the task.

Also known as: modality, data modalities

Updated

Modality describes the information channel

The term is used at different levels. RGB and depth may be treated as separate visual modalities because their values and sensing processes differ. Joint state and motor command can share a numeric representation while remaining different modalities because one records an observation and the other an intended action. Language can appear as a task instruction, annotation or model-generated description, each with different provenance.

The multimodal machine-learning survey by Baltrušaitis, Ahuja and Morency organises the field around representation, translation, alignment, fusion and co-learning. That framing matters in robotics: the presence of two channels does not explain how their samples correspond or how a model should use them.

Multimodal does not mean synchronised

A dataset may contain video, joint state and language while failing to align them at the precision needed for control. Cameras can expose at 30 Hz, encoders at 1 kHz and instructions once per episode. Each stream needs its own timestamps, rate, clock, frame and missing-data convention before it can be resampled or paired safely.

Semantic alignment matters too. A caption may describe the whole episode while an action controls a few milliseconds. A force sample can mark contact that is hidden by an occluded camera. Converting all streams into tokens or arrays does not erase those differences.

The right modalities depend on the learning target

RT-1 is a concrete vision-language-action example: it conditions a robot policy on images and a natural-language instruction and produces robot actions. Those inputs and outputs serve different roles even though one model processes them together.

Dataset documentation should list each modality's physical or semantic meaning, sensor or annotation source, shape, units, coordinate frame, rate, timing, preprocessing and relationship to actions and outcomes. Buyers should ask which signal answers the learning question. Extra channels add storage and integration cost when they are poorly calibrated, redundant or unrelated to the task.

Sources