humanoidsdata.com

Search

Search companies, datasets, articles, and glossary terms for humanoids and embodied AI.

By Lumi · Humanoid robot data · · 20 min read

Egocentric Data Collection Systems for Robot Learning

An egocentric data collection system is not simply a camera worn on the head. For robot learning, a useful system must turn the actor's view into an auditable record: video, timestamps, calibration, camera or body motion, hand interaction, task context, quality checks, and rights.

There is no universal best system. Research glasses such as Meta Project Aria preserve calibrated first-person sensing. Pupil Labs and Tobii specialise in gaze. Meta Quest and Apple Vision Pro support interactive capture applications. GoPro cameras and smartphones make field collection easier. RealSense and Luxonis components give engineers control over depth and mounting. UMI-style tools trade natural bare-hand motion for a more robot-like action signal. Managed platforms such as Lightwheel EgoSuite and IO-AI SenseXperience add collection operations and post-processing.

The right choice follows the training target. A world model may need broad video, audio, and task diversity. A manipulation policy needs a credible trajectory or action label. A dexterous system may also need finger state and contact. A humanoid still needs robot-native data before deployment, regardless of how much human footage it sees.

This guide is a representative market map checked in August 2026, not an exhaustive catalogue or a product endorsement. Product access, APIs, and configurations change; the comparison focuses on durable differences in what each system records.

First-person industrial task recording with 3D skeleton overlays on both hands

Modern collection systems increasingly pair first-person video with tracked hands and structured task data. Source: Lightwheel EgoSuite-Open100K.

A collection suite has five layers

Calling a wearable device a suite can hide most of the work. A complete egocentric collection stack has five layers.

  1. Sensing: head, chest, or wrist video; depth; audio; gaze; inertial data; hand and body pose; object pose; and, when relevant, force or tactile signals.
  2. Time and geometry: a shared clock, camera intrinsics and extrinsics, sensor calibration, coordinate frames, dropped-frame records, and a tested data-synchronisation method.
  3. Interaction supervision: task instructions, episode boundaries, hand or gripper trajectories, object-state changes, success, failure, intervention, and confidence values for estimated signals.
  4. Operations: a capture application, participant workflow, device charging and storage, live quality checks, upload, review, redaction, and version control.
  5. Delivery and governance: raw recordings, derived annotations, schemas, loaders, licences, consent records, privacy treatment, and a documented export for training.

A product can be excellent at one layer and weak at the others. An eye tracker may deliver reliable gaze but no 3D hand pose. A depth camera may provide calibrated geometry but no wearer intent. A managed service may provide a finished dataset without selling the capture hardware. Buyers should compare the whole path, not the number of cameras on the device.

The product landscape at a glance

The following categories are more useful than a single brand ranking because they expose what each system measures natively and what must be estimated later.

System classRepresentative productsStrongest native signalBest fitMain limitation
Research glassesMeta Project Aria Gen 2Calibrated RGB, machine-perception cameras, gaze, IMU, audio, and device poseTracked first-person research and human–object interactionApplication-gated access; no native robot command or contact-force label
Commercial head rigsGenRobot DAS EgoMulti-camera RGB, inertial data, audio, and processed head and hand trajectoriesProduction-oriented human demonstrations and multi-device captureQuote-only system; synchronization and accuracy claims need sample-level verification
Eye-tracking glassesPupil Labs Neon, Tobii Pro Glasses 3Scene video aligned with measured gaze and inertial dataAttention, intent, inspection, and human-factors studiesFull hand pose, scene depth, and executable actions are not native outputs
XR headsetsMeta Quest 3 and 3S, Apple Vision ProHead and hand tracking inside an interactive applicationGuided collection, live feedback, and teleoperation interfacesHeavier form factor and platform-specific camera permissions
Video-first captureGoPro cameras, smartphones, compact body camerasPortable, high-quality RGB and audio at field scaleTask understanding, visual pretraining, and broad environment coverageMetric hand, camera, and object trajectories require reconstruction or added trackers
Modular RGB-D rigsRealSense D405 and D455, Luxonis OAK-D familyConfigurable colour, stereo depth, and sometimes IMU or on-device visionCustom close-range, head, chest, wrist, or robot-mounted systemsThe team must engineer mounting, power, compute, recording, and calibration
Instrumented interactionUMI, TRumi, DAS Gripper, PIKA, Grabette/Gripette, mocap suits and glovesHand, body, gripper, or reconstructed end-effector motionRobot-free demonstrations and action-proxy collectionRetargeting and embodiment mismatch remain; motion is not a measured robot action
Managed full-stack platformsLightwheel EgoSuite, IO-AI SenseXperienceHardware, field operations, processing, review, and training-format deliveryLarge or custom programmes across many tasks and environmentsConfiguration, ownership, exclusivity, and quality guarantees must be confirmed contractually

These categories overlap. A Quest headset can be paired with wrist cameras. Project Aria can be combined with an inertial suit. A UMI-style gripper can add a head camera. The important distinction is which measurements come directly from hardware, which are calculated by software, and which are added by human or model annotation.

Project Aria is a sensor-rich research reference

Meta's Aria Gen 2 Research Kit is one of the clearest examples of a genuine egocentric suite. The current kit includes the glasses, a companion application, machine-perception services, desktop tools, a client SDK, and support. Access is through a rolling application for selected research partners, so it should not be presented as an ordinary retail camera.

The Gen 2 hardware specification lists one RGB camera, four computer-vision cameras, two eye-tracking cameras, two IMUs, seven spatial microphones, a contact microphone, barometer, magnetometer, GNSS, ambient-light and PPG sensors, and six to eight hours of continuous recording. On-device acceleration supports hand tracking, eye tracking, and six-degree-of-freedom localisation. Meta's Machine Perception Services can add higher-precision trajectories, hand tracking, gaze, and semi-dense scene geometry after capture.

That combination is valuable because the calibration and pose pipeline belongs to the platform. It reduces the amount of reverse engineering needed to connect pixels, gaze, hands, and head motion. Aria Gen 2 also supports multi-device clock alignment. Its documentation carefully says that time alignment is not the same as synchronously triggering every exposure, a distinction that matters when contact timing is measured in milliseconds.

Aria still does not observe robot joint commands, object force, or tactile contact. Its hand poses are tracked estimates of a human body. They can support pose estimation, task understanding, reconstruction, and motion retargeting, but they do not become policy targets until a team defines and validates that conversion.

Commercial head rigs are becoming data appliances

GenRobot DAS Ego is a quote-based commercial head-mounted system rather than an application-gated research device. Its current documentation lists six global-shutter RGB cameras, an IMU array, microphone, local storage, and hot-swappable batteries. Raw sessions are stored as MCAP, while the processing stack adds 30 Hz six-degree-of-freedom trajectories and 21-keypoint 3D hand tracks.

That integrated design can reduce the assembly work of a custom rig, but it does not remove diligence. GenRobot advertises sub-millisecond ego-to-hand and multi-device alignment; the public material does not describe enough of the timing hardware to treat that as independently established performance. A buyer should request the exact hardware revision, clock design, calibration files, raw and processed sample, tracking-validity fields, and measured drift over a full session.

Pupil Labs and Tobii make gaze the primary measurement

Eye-tracking glasses answer a different question: what did the wearer attend to while acting?

Pupil Labs Neon combines a 1600 × 1200 scene camera at 30 Hz with binocular eye cameras, 200 Hz gaze output, a 110 Hz nine-degree-of-freedom IMU, and microphones. Pupil Labs publishes the module geometry, exposes raw sensor interfaces, and provides camera calibration data. That openness makes Neon attractive when a team wants to integrate gaze into its own recorder rather than remain inside a closed analysis application.

Tobii Pro Glasses 3 records 1080p scene video at 25 frames per second, binocular gaze at 50 or 100 Hz, audio, and inertial data. Its recording unit includes a TTL sync port, and the Glasses 3 API exposes live control, scene video with gaze, and stored-data workflows through standard protocols.

For robot learning, gaze can help identify the attended object, the next likely interaction, and moments where the wearer checks a result. It is auxiliary supervision, not proof of intent. People look without acting, act in peripheral vision, and move their heads independently of their eyes. Neither product natively supplies a full 3D hand trajectory or a robot command, so manipulation projects usually add hand tracking, body tracking, or a robotless interaction device.

XR headsets turn recording into an interactive workflow

XR headsets are bulkier than glasses, but they can show instructions, collect confirmations, display quality warnings, and track hands while recording the surrounding scene. That makes them useful when the collection application needs to guide the demonstrator instead of operating as a passive recorder.

Meta publicly released its Passthrough Camera API for Quest 3 and Quest 3S. The current APIs expose front-facing RGB images, camera intrinsics, extrinsics, poses, and timestamps; hand tracking remains a separate system signal. Applications require camera permission, and the headset displays recording indicators. A team can therefore build a repeatable workflow around video, head pose, hand joints, task prompts, and live quality checks, but it must write and maintain the application.

Apple Vision Pro also supports tracked spatial applications, but raw main-camera access is not a normal consumer permission. Apple's main camera documentation places stereo camera frames behind an approved visionOS enterprise entitlement and licence. That distinction is easy to miss when a research paper says it used Vision Pro: the hardware may provide excellent hand and head tracking while the app's access to raw video depends on programme approval and deployment context.

The ergonomic cost is real. A headset can change head motion, fatigue, field awareness, and how naturally a person reaches into tight spaces. Pilot tests should measure usable recording time and behaviour, not only nominal battery life.

GoPro cameras and phones trade geometry for reach

Action cameras and smartphones remain practical because they are easy to source, replace, mount, and operate. Ego4D deliberately used seven types of head-mounted camera, including GoPro and Pupil Labs devices, to collect broad daily-life footage without tying the dataset to one sensor.

The trade-off is that video-first hardware does not normally output a calibrated 3D hand skeleton or an executable action. Even synchronization needs care. GoPro's current timecode guidance reports sub-50-millisecond metadata alignment for supported cameras, while warning that clocks can drift after 60 to 90 minutes and must be resynchronised after a power cycle. That is useful for editing. It is not the same as a shared hardware trigger for high-rate control signals.

Ego-Exo4D shows what a disciplined multi-camera protocol can achieve. Each activity used Project Aria for the first-person view and four or five stationary GoPros for external views. The team used a timed QR sequence and manual start/end checks to align cameras, then localised the views in a shared metric frame. The value came from the protocol and post-processing, not from adding more cameras alone.

Synchronized first-person and external camera views from the Ego-Exo4D collection

Ego-Exo4D combines a calibrated first-person device with several external cameras, showing that a capture suite can span the wearer and the environment. Source: Ego-Exo4D.

Phones push reach further. The 2026 Open-AoE preprint describes a roughly 2,000-hour collection from more than 500 contributors using over 400 smartphone models, followed by privacy erasure, segmentation, camera-trajectory reconstruction, hand-pose estimation, and quality review. Its dataset card reported about 694 hours uploaded on 12 August 2026, so the full release was still in progress at publication. The project is evidence for scalable collection, not evidence that all phones produce interchangeable measurements. Device metadata, mount geometry, stabilisation, exposure, rolling shutter, frame rate, and calibration still need to survive the pipeline.

RealSense and Luxonis support custom RGB-D rigs

Teams that need depth but do not want a full XR platform often build around camera modules.

The RealSense D405 is designed for close-range stereo depth, with an ideal range of 7 to 50 centimetres. That makes it relevant to wrist or manipulation views. The D455 covers a longer working range and includes an IMU, making it more suitable for room-scale or head/chest configurations.

The Luxonis OAK-D Pro W combines wide-field stereo depth, RGB, active infrared, an IMU, and on-device processing. Its DepthAI software can run and synchronise vision pipelines close to the sensor. The newer OAK 4 D adds a standalone rugged form factor, onboard storage and compute, and hardware timing options for larger custom rigs.

These are components, not finished wearable products. A custom build still needs a rigid mount, calibrated transforms after assembly, power, storage or a compute pack, thermal testing, cables that do not impede the wearer, and a recorder that preserves timestamps. Custom hardware offers control, but it also makes the engineering team responsible for every missing layer.

Motion capture and robotless interfaces add action proxies

Head video explains the scene. Human motion systems add candidate movement labels.

Full-body products such as Xsens Link and Rokoko Smartsuit Pro II stream body orientations and inertial data. Hand systems such as MANUS Metagloves Pro and the StretchSense Studio Glove add finger articulation. These can be combined with first-person video for whole-body and dexterous capture.

Three distinctions should remain visible. An inertial suit estimates human kinematics; it does not record a robot controller. A motion-capture glove measures finger pose; it does not necessarily measure pressure, slip, or object force. The hand and body streams also need a shared world frame and tested clock relationship with the camera.

Universal Manipulation Interface, or UMI, goes one step closer to a policy target. Its handheld parallel-jaw gripper carries a GoPro, records gripper width, and reconstructs a six-degree-of-freedom trajectory from visual and inertial data. The human can collect demonstrations without a robot at the site, while the interface resembles the end effector that will execute the task.

UMI is an open research framework, not a turn-key guarantee. Its own repository calls visual-inertial SLAM the most fragile part of the processing pipeline. Commercial products in the same broad category include Trossen TRumi, GenRobot DAS Gripper, AgileX PIKA, and Lumos FastUMI. Pollen Robotics' Grabette and Gripette offer a newer open-build pairing of handheld and robot-side hardware. These systems, and others covered in the data collection gripper guide, differ in cameras, depth, tactile sensing, trajectory reconstruction, export formats, and support.

The fundamental trade-off is stable across brands. Freehand capture preserves natural human dexterity but leaves more motion to estimate and retarget. An instrumented gripper narrows the hand to a robot-like interface and gives a cleaner action proxy, but it changes how the person performs the task. Neither is a measured action from the final robot.

Managed platforms bundle hardware, operations, and processing

A growing group of vendors sells the collection programme rather than one device.

Lightwheel EgoSuite describes three capture classes: a VR-based multimodal unit, an exoskeleton system for dexterous manipulation, and a UMI-aligned gripper. Its public materials list RGB-D, upper-body and hand pose, tactile data, on-device processing, global field operations, and post-processing. The related EgoSuite-Open100K release is staged around head-only and head-plus-wrist configurations, with MCAP and LeRobot v3 delivery. Lightwheel's production volumes and accuracy figures are vendor-reported claims; buyers should verify them against samples and acceptance tests.

IO-AI SenseXperience takes a modular approach. Its baseline combines an egocentric unit, wrist unit, gripper unit, and compute unit, with optional motion capture and tactile hardware. The associated EmbodiFlow pipeline handles timestamp alignment, pose estimation, gesture recognition, review, and export to LeRobot, HDF5, or MCAP.

Those examples illustrate a market shift from hardware catalogues to data operations. A procurement brief should still separate three offers that are often blended together:

  • Hardware purchase or rental.
  • A managed collection service using the vendor's workforce and locations.
  • A licence to an existing dataset.

Each has different ownership, exclusivity, maintenance, privacy, and reproducibility terms. A polished platform page does not replace a raw sample episode, a schema, or a contract that grants the intended model-training rights.

Video is becoming the carrier for structured interaction data

The direction is away from counting undifferentiated video hours and toward attaching camera trajectories, 3D hands, body pose, atomic actions, object state, and language to the frames.

The April 2026 EgoVerse preprint reports 1,362 hours across controlled academic and in-the-wild industry collections, normalised around calibrated head pose, 3D hand keypoints, and task descriptions. Its live project page later reported 4,003 hours while retaining the other headline counts, so those figures should be treated as different release snapshots. EgoDex pairs 829 hours of Vision Pro video with camera calibration and tracked upper-body and hand pose. Open-AoE connects phone capture to reconstruction and training conversion. The common idea is not that every estimate is ground truth. It is that the estimate, coordinate frame, confidence, and method become explicit dataset fields.

Robot-free capture is scaling, but embodiment debt remains

UMI, tracked glasses, XR headsets, and commercial field programmes can reach environments where placing a robot is impractical. That can improve object, task, operator, and location diversity.

It also creates what can be called embodiment debt: the later work required to map human hands, a handheld gripper, or a headset coordinate system onto the target robot. Recent methods such as HumanEgo, EgoMimic, and the 2026 ACE-Ego-0 preprint explore shared or canonical action representations. ACE-Ego-0 explicitly treats reconstructed human pseudo-actions as noisier than sensor-logged robot actions. That is the right caveat: conversion can make human data useful, but it does not change how the source label was obtained.

Head, wrist, and external views are becoming complementary

A head camera preserves scene context and intent but can lose contact behind the hands. A wrist camera sees the interaction closely but loses the larger task. An external camera helps recover body and object geometry. Ego-Exo4D, EgoSuite, and many robot datasets now combine these distances rather than asking one camera to do every job.

More views only help when they are calibrated and synchronized. Otherwise the extra footage increases storage and annotation cost without creating a coherent 3D event.

Collection is becoming an active feedback loop

The collection application is starting to judge coverage while the operator is still in the scene. The 2026 EgoGuide preprint adds AR feedback to a UMI-style workflow so demonstrators can vary initial poses, object layouts, viewpoints, or start states when coverage is weak. Its reported gains are task-specific, but the operational idea is broader: reject or correct a weak episode before the worker leaves, rather than discovering the gap after an expensive training run.

Raw logs and training packages are separating

MCAP is useful as a raw archive because it stores timestamped messages, schemas, metadata, and attachments in one indexed container. It does not create synchronization: producers still choose the clock and publish times, and a bad clock remains bad inside a good file.

LeRobotDataset v3 is aimed at model consumption. It separates high-frequency tabular signals into Parquet, camera streams into sharded MP4, and episode, task, schema, and statistics metadata into indexed files. A robust pipeline can retain immutable raw MCAP or device-native recordings, then generate versioned LeRobot exports. That separation makes it possible to improve annotations or conversion without overwriting the evidence.

Privacy is becoming part of the engineering design

First-person capture records more than the intended task. It can collect faces, voices, screens, home interiors, health signals, location, and bystanders who never touched the device.

The Ego4D privacy and ethics process is a useful reference: informed consent for wearers, controlled collection where possible, de-identification, automated and human review, and a route to report insufficient redaction. A commercial programme also needs retention limits, access controls, withdrawal handling, regional review, and licence terms that match training and redistribution. Redaction should be versioned like any other transformation; it is not a final cosmetic pass.

A practical reference architecture

For a bimanual household collection intended to support a future humanoid policy, a defensible stack could look like this:

  1. Define the eventual policy interface. Name the target cameras, robot state, action space, hand model, control rate, task language, and outcome fields before selecting hardware.
  2. Capture the broad human record. Use a calibrated head device for scene context and head pose. Add wrist views when hand occlusion is likely.
  3. Capture motion at the required fidelity. Use headset hand tracking, optical or inertial motion capture, gloves, or an instrumented gripper. Store raw measurements and estimated poses separately.
  4. Establish one timing and calibration plan. Record clock sources, camera parameters, transforms, dropped frames, latency, recalibration events, and device firmware. Validate alignment at the beginning and end of each session.
  5. Guide and inspect collection in the field. Show the task, mark starts and outcomes, retain failures, monitor framing and tracking confidence, and let the operator repeat a bad episode immediately.
  6. Keep raw and derived data apart. Preserve device-native files or MCAP as the immutable source. Version hand tracks, segmentations, captions, redaction, retargeting, and LeRobot exports.
  7. Collect a robot-native anchor set. Run a smaller set of demonstrations or evaluations on the target robot to measure the remaining camera, kinematic, controller, contact, and latency gap.

The seventh step is what prevents a large human dataset from becoming an expensive assumption. Egocentric capture can expand the task distribution and teach reusable structure. Only the target machine can reveal which parts survived embodiment transfer.

How to choose a system

Choose from the learning requirement backwards:

  • For visual pretraining, procedures, and world modelling, prioritise comfortable video-first hardware, audio, task diversity, consent, and reliable episode metadata.
  • For gaze, attention, and intention studies, start with Pupil Labs or Tobii and add hand or body tracking only when the task needs it.
  • For calibrated head motion, 3D scene understanding, and tracked human interaction, Project Aria is a strong research reference when programme access is available.
  • For guided capture or live operator feedback, Quest or Vision Pro can be useful if the required camera APIs and deployment permissions are available.
  • For close-range geometry, build around an RGB-D component and accept responsibility for mounting, compute, power, synchronization, and calibration.
  • For manipulation action proxies, use an instrumented gripper or tracked hand system and document the reconstruction and retargeting chain.
  • For policy-ready control, pair human egocentric data with teleoperation or robot rollouts that contain the target robot's observations, states, commands, and outcomes.
  • For outsourced scale, evaluate a managed provider's sample data and operating process, not only its headline hours or device list.

Before committing, ask the vendor or internal team to load one complete episode and answer:

  • Which signals are raw, measured, estimated, retargeted, or annotated?
  • Which clocks produced the timestamps, and what drift was measured over a full session?
  • Are intrinsics, extrinsics, coordinate frames, units, and calibration history included?
  • What happens when hands are occluded, tracking is lost, a frame drops, or a worker performs the wrong task?
  • Does a pose glove record articulation only, or also contact and force?
  • Does the export preserve failures, retries, corrections, and confidence scores?
  • Can the data be inspected without proprietary software?
  • Which exact hardware, firmware, application, and post-processing versions produced it?
  • Do consent and licence terms cover commercial model training, derived weights, evaluation, and the required retention period?

The best egocentric suite is not the one with the longest sensor list. It is the one whose evidence chain remains intact from photons and motion to timestamps, derived labels, training files, and robot validation.

For adjacent decisions, use the guides to humanoid data collection equipment, robot teleoperation systems, and humanoid and egocentric data sources.