Data & collection
Data modality
A data modality is a distinct kind or channel of information characterised by how it is sensed, represented, or interpreted, such as RGB images, depth, language, joint state, actions, audio, tactile measurements, or force–torque signals. A multimodal robot dataset contains more than one modality, but usefulness depends on their alignment, semantics, and relevance to the task.
Also known as: modality, data modalities
Updated
Modality describes the information channel
The term is used at different levels. RGB and depth may be treated as separate visual modalities because their values and sensing processes differ. Joint state and motor command can share a numeric representation while remaining different modalities because one records an observation and the other an intended action. Language can appear as a task instruction, annotation or model-generated description, each with different provenance.
The multimodal machine-learning survey by Baltrušaitis, Ahuja and Morency organises the field around representation, translation, alignment, fusion and co-learning. That framing matters in robotics: the presence of two channels does not explain how their samples correspond or how a model should use them.
Multimodal does not mean synchronised
A dataset may contain video, joint state and language while failing to align them at the precision needed for control. Cameras can expose at 30 Hz, encoders at 1 kHz and instructions once per episode. Each stream needs its own timestamps, rate, clock, frame and missing-data convention before it can be resampled or paired safely.
Semantic alignment matters too. A caption may describe the whole episode while an action controls a few milliseconds. A force sample can mark contact that is hidden by an occluded camera. Converting all streams into tokens or arrays does not erase those differences.
The right modalities depend on the learning target
RT-1 is a concrete vision-language-action example: it conditions a robot policy on images and a natural-language instruction and produces robot actions. Those inputs and outputs serve different roles even though one model processes them together.
Dataset documentation should list each modality's physical or semantic meaning, sensor or annotation source, shape, units, coordinate frame, rate, timing, preprocessing and relationship to actions and outcomes. Buyers should ask which signal answers the learning question. Extra channels add storage and integration cost when they are poorly calibrated, redundant or unrelated to the task.
Sources
Related terms
Data & collection
Robot training data
Robot training data is recorded experience used to train, fine-tune, or adapt models for robot perception, prediction, planning, or control. It can include sensor observations, robot state, actions, task instructions, rewards or outcomes, demonstrations, failures, and embodiment metadata. Not every dataset contains every field, but their timing and physical meaning must be clear.
Data & collection
Data synchronisation
Data synchronisation is the process of placing sensor, state, action, annotation, and outcome records on a common timeline so samples that describe the same physical instant or transition can be matched. It requires trustworthy timestamps or trigger relationships and an explicit rule for handling streams with different rates, delays, dropped samples, and clock offsets.
Data & collection
Language annotation
A language annotation is natural-language metadata attached to a robot-data sample, segment, or episode. It may state the instruction given before execution, describe what happened afterward, name a task or subtask, identify objects, record a correction, or explain an outcome. These annotation types are not interchangeable because they contain different information and may be available at different times.
Data & collection
Depth data
Depth data records the distance associated with image locations or sensor rays, usually as a depth image in which each pixel stores a metric value relative to a camera. The exact geometry, units, invalid-value convention, and coordinate frame depend on the sensor and encoding. RGB-D data pairs depth with colour imagery; a point cloud is a separate 3D representation derived from or aligned with such measurements.
Hardware & control
Tactile sensing
Tactile sensing is the detection and measurement of physical contact properties at a robot's surface or contact interface. Depending on the sensor, it can report pressure or force distribution, contact location, shear, vibration, slip, texture, temperature, or deformation. Tactile data complements vision by measuring interactions that may be hidden at the point of contact.
Hardware & control
Proprioception
Proprioception is sensing of a robot’s own internal configuration and motion rather than the external scene. For a humanoid it commonly includes joint positions and velocities, actuator effort or torque, and inertial measurements of body rotation and acceleration. These signals support state estimation and feedback control but do not, by themselves, directly describe nearby objects or terrain.