humanoidsdata.com

Search

Search companies, datasets, articles, and glossary terms for humanoids and embodied AI.

By Lumi · Humanoid robot data · · 31 min read

How to Prepare Egocentric Data for Sale: A Seller's Guide

Egocentric dataset preparation checklist

Work through these 38 checks using a separate copy for each dataset release and intended buyer use. Download the editable template above or copy the items into your project document. The template includes fields for the release, reviewer, acceptance thresholds, and evidence log.

Tick an item only after checking its evidence. For conditional items, record Not applicable and a reason in your copy. A missing signal required by the buyer is a blocker, not a reason to mark it inapplicable. Depth, IMU, tracked poses, and calibrated geometry are optional unless the agreed product requires them. Use the buyer's documented thresholds rather than treating the illustrative numbers elsewhere in this guide as universal requirements.

A. Define the dataset and acceptance criteria

  • A1. Record the dataset name, release version, responsible owner, and intended training or evaluation use.
  • A2. Write the task list, environments, object coverage, completion criteria, and excluded situations in a collection brief.
  • A3. List each promised modality and label as measured, manually annotated, model-derived, or unavailable.
  • A4. Agree on acceptance thresholds, audit sampling, rejection reasons, and rework responsibilities before scaling collection.
  • A5. Have the buyer evaluate representative pilot episodes in the intended workflow; record accepted scope and unresolved gaps.

B. Verify rights and privacy evidence

  • B1. Trace every episode and third-party annotation or asset to evidence supporting the proposed commercial use and distribution.
  • B2. Record the applicable personal-data processing basis, participant notices or permissions, and location permissions where needed.
  • B3. Review video, audio, transcripts, metadata, and derived files for bystanders, identifying details, and confidential material; resolve flagged cases.
  • B4. Check the delivered copy after redaction or exclusion, including whether privacy edits obscure required task information.
  • B5. Keep signed releases and identity lookup records in restricted storage; put only appropriate references and restrictions in the delivery package.
  • B6. Document retention, correction, withdrawal or removal handling, and how affected episodes and recipients will be identified.

C. Inspect recordings and timing

  • C1. Record the capture device, mount, firmware where known, resolution, frame-rate behavior, lens, stabilization, and available sensor streams.
  • C2. Inspect complete task intervals for framing, required hand/object visibility, blur, exposure, and setup/outcome context against the agreed criteria.
  • C3. Validate decoding and frame counts; identify missing, duplicated, frozen, or black frames and document their treatment.
  • C4. Document timestamp units, clock domains, capture versus presentation time, frame indexing, and source-to-export time mappings; disclose unavailable timing information.
  • C5. Where streams must align, measure offset and drift across recordings and document synchronization residuals, gaps, interpolation, or resampling.

D. Validate annotations and optional geometry

  • D1. For supplied annotations, version the guide and vocabulary; define the included instructions, actions, outcomes, failure/recovery labels, and unknown values.
  • D2. If temporal segments are supplied, validate boundaries against episode duration and document interval-end conventions, overlaps, and frame mappings.
  • D3. If annotations are supplied, audit a defined sample across relevant tasks and languages; record annotation origin, reviewer coverage, disagreements, and corrections.
  • D4. If geometry is supplied, document camera models, calibration versions, units, coordinate frames, transform directions, and joint or quaternion ordering.
  • D5. If estimated poses or pseudo-actions are supplied, record estimator versions, input lineage, validity masks, confidence meaning, and accuracy evaluation separately from availability.
  • D6. Confirm that field names and documentation distinguish human motion and derived targets from measured robot state or commands.

E. Audit quality, coverage, and evaluation splits

  • E1. Publish a quality report with metric definitions, denominators, audit sample sizes, rejection reasons, and known limitations.
  • E2. Reconcile accepted episode hours, active-task hours, camera-stream hours, and annotated hours without double-counting overlapping exclusions.
  • E3. Report accepted coverage by task, source session, site, participant, and object instance using the necessary non-identifying references.
  • E4. Identify duplicate and near-duplicate groups; where splits are supplied, keep related source clips and their derivatives in the same split.
  • E5. If evaluation splits are supplied or claimed, check that participant/site holdouts match that claim and that learned preprocessing does not leak evaluation data.

F. Build and test the delivery package

  • F1. Include a release-specific dataset card, license, schema, manifest, change log, quality report, and file checksums.
  • F2. Verify unique episode identifiers, file references, required fields, array shapes, units, missing-value conventions, and annotation-to-video mappings where supplied.
  • F3. Pin the exporter, dependencies, and tested loader version; finalize any writers before validating the release.
  • F4. Load representative and randomly selected episodes from the final package in a clean environment using the supplied instructions.
  • F5. Where applicable, check format-specific semantics such as LeRobot episode offsets or RLDS step fields, terminal states, and invalid final-step actions.
  • F6. Test the agreed transfer method, selective or resumable access where promised, downloaded-file checksums, and access restrictions.

G. Review the commercial handoff

  • G1. Confirm that the listing and sample match the actual release, including missing modalities, accepted counts, rights restrictions, and quality limitations.
  • G2. Define evaluation versus training rights, affiliates/contractors, redistribution, derivatives, and any exclusivity scope and duration in the agreement.
  • G3. Agree on the billing unit, acceptance window, rejection evidence, rework limits, payment milestones, and storage or egress costs.
  • G4. Record the exact release and episode manifest supplied to each recipient, with a contact and process for corrections or removals.
  • G5. Complete the release review: resolve blockers and document buyer-agreed quality exceptions with their impact, owner, and follow-up action.

For each item, keep an evidence reference and one of four statuses: Done, In progress, Blocked, or Not applicable. For example, D4 can be “Not applicable — RGB-only product; no calibrated geometry promised.” C5 must remain blocked if synchronized streams were promised but their alignment cannot be verified. Store references to sensitive rights evidence rather than copying that evidence into a shared checklist.

Do not reduce this review to a percentage score. Unresolved commercial rights or privacy release issues, corrupt or unloadable required data, and missing or contradictory required timing, identifiers, or supervision should stop the affected material from being delivered. A completed checklist records the preparation and review performed; it does not independently certify legal compliance or guarantee model performance.

The detailed preparation guide

To sell egocentric data, prepare a versioned dataset that a buyer can legally use, inspect at episode level, and load into its training pipeline. The deliverable is usually video plus a manifest, timestamps, task annotations, quality evidence, and a defined license. Tracked hand poses, depth, or robot actions belong in the package only when you actually have those signals and can explain how they were produced.

The commercial question is whether your recordings fill a specific training gap. Fifty hours of well-documented appliance interactions in the buyer's deployment environment may be more relevant than a much larger archive of unrelated activity. That is a question to settle with a representative pilot, before paying to annotate or collect the entire archive.

This guide follows that preparation process, from choosing the product to delivering the files. The public datasets below are references for packaging and technical practice; their availability does not grant permission to resell them. Source details and market developments were checked on October 5, 2026. Numerical acceptance criteria and business calculations explicitly marked as examples are proposed working assumptions, not industry standards or observed market prices.

HOT3D examples of first-person hand-object interactions with three-dimensional annotations

A useful egocentric product can include much more than RGB video. HOT3D illustrates hand-object interactions with geometry and tracking annotations. Source: HOT3D project. Its recordings, hand annotations, and object models have different license conditions.

For a particular stage, jump to the buyer brief, rights and privacy, capture and timing, dataset examples, file formats, the example package, acceptance and pricing, or the dataset preparation checklist.

What recent research and market reporting tell sellers

Recent work makes a strong case for human video as a training input, but also shows why footage needs substantial preparation. In its August 2026 Dyna-2 research report, Dyna Robotics describes pretraining on more than one million hours of egocentric human video, largely head-mounted recordings of everyday manipulation. It also describes cleaning, hand-pose extraction, validation, and filtering. Episodes that pass its pose-quality threshold receive 3D tracks; wrist poses and thumb-index aperture provide derived trajectory and grasp supervision.

Those are company-reported methods and results. They do not establish a market price per hour or guarantee that a new supplier's video improves a robot. The practical seller lesson is more specific: useful motion supervision has a processing history, an acceptance threshold, and missing cases. A hand-opening estimate is a derived label, not a measured gripper command or contact force.

Dyna describes these as 36 clips from its human-video pretraining corpus. This is the collection illustration in Dyna-2 Figure 4, separate from the report's generated-video demonstrations. Open the source video.

Toloka's June 15, 2026 HomER v2 announcement reports 765 first-person videos, approximately 100 hours of household activity. Its discussion of quality includes hand visibility and visual clarity. Such release criteria are useful examples of a collection program's choices; they should not become universal thresholds for every camera, task, and customer.

The supply side is less frictionless than collection advertisements suggest. ABC News's September 20, 2026 reporting describes workers dealing with rejected recordings and uncertainty about downstream use. An Appen executive interviewed for the story distinguishes generic footage from recordings that match the intended deployment environment, including differences in household appliances. These are reported experiences and attributed market observations, rather than a representative survey of buyers.

Together, the sources point toward a practical business model: agree on a narrow collection brief, prove the preparation pipeline on a pilot, and scale the accepted product. Uploading a large video archive first leaves the buyer to discover every mismatch at your expense.

What practitioner discussions add

Forum threads are useful for finding operational problems, but weak evidence for market size or earnings. In the August 31, 2026 Hacker News discussion of Hebbian, the founder describes checks for frozen or black camera streams, missing topics, timestamp drift, and duplicates, with an explicit egocentric-vendor use case. That is a tooling vendor's account, not independent validation. It nevertheless suggests concrete questions a seller should be able to answer in a quality report.

A Project Aria issue about external timecode synchronization, discussed in late 2025, describes the difficulty of getting sufficiently precise alignment from visual or clap events and handling clock drift. It is one user's integration experience, not proof that every Aria recording has the same problem. The official timestamp documentation establishes the underlying distinction: device capture time and host recording time have different meanings. “Synchronized” in a sales listing should therefore name the clocks, method, and measured error.

Define the product and test a small pilot

Start by writing one sentence that describes the training input you can supply. For example: “Consented head-mounted RGB recordings of adults loading and unloading household dishwashers, with episode boundaries, action segments, quality flags, and no measured robot actions.” This is a hypothetical product description. Its exclusions make the scope testable.

Three products commonly get called egocentric data, although they require different acceptance tests:

ProductWhat the buyer receivesPlausible training useWhat the package must not imply
Human first-person videoRGB frames, task descriptions, temporal annotations, optional audioVideo understanding, task semantics, visual pretraining, world-model trainingVideo alone supplies metric 3D motion, contact forces, or executable robot actions
Human video with tracked or estimated geometryVideo plus camera, hand, body, or object poses; calibration; validity and confidenceMotion understanding, spatial supervision, derived action representationsHuman joints or estimated wrist trajectories are the target robot's motor commands
Robot-mounted recordingsFirst-person views paired with recorded robot observations and commandsPolicy learning for a documented embodiment and action spaceA command is necessarily the measured motion, or commanded torque is measured contact force

Our guides to egocentric collection systems and human motion for robot training explain the acquisition and transfer differences. For a sale, express them in a short buyer brief with the following decisions:

  • Tasks and environments: define actions, objects, initial conditions, completion criteria, and excluded settings. “Kitchen” is too broad to specify an appliance-manipulation dataset.
  • Coverage: agree on participants, locations, object instances, lighting, handedness where relevant, and how many repeated takes are useful. Use coarse, necessary attributes rather than collecting personal details by default.
  • Signals: list each sensor and annotation. State which are measured, manually labeled, or model-derived. Mark unavailable signals explicitly.
  • Capture and delivery: specify the viewpoint, resolution, actual frame-rate behavior, audio policy, timestamps, required geometry, storage format, and tested loader version.
  • Acceptance and rights: define how a unit passes inspection, what licenses the buyer needs, how rejected units are handled, and whether exclusivity applies.

A practical starting pilot could contain 20–50 complete episodes spanning the intended tasks and difficult cases. That range is a planning example, not a statistically representative sample or a prescribed minimum. Include ordinary episodes, occlusion, a failure or recovery if in scope, and examples from different capture sessions. Let the buyer select some episodes from the manifest so that evaluation does not depend entirely on your best clips.

Separate permission to evaluate the pilot from permission to train on the full release. A useful pilot tests file loading, temporal alignment, task fit, annotation interpretation, rights documentation, and the buyer's actual intended workflow. Agree on the cost and ownership of pilot enrichment before performing it.

Establish rights before enriching the data

There are at least three separate questions: who can license the recordings and annotations, what personal-data processing is lawful, and what uses the buyer's contract permits. Owning a camera or paying a collector does not, by itself, answer all three.

Build a restricted rights register before expensive labeling. Record the capture source, collector agreement, participant notice or release version, location permission where needed, third-party assets, and restrictions on commercial training, onward transfer, or redistribution. Connect these records to episodes using opaque references. Signed releases, identity documents, and private addresses should not travel in the ordinary training folder.

For EU personal data, GDPR Articles 5–7 distinguish processing principles, lawful bases, and consent requirements. Consent is one possible lawful basis, not the only one. Where it is the basis, it must be demonstrable and withdrawable; a broad copyright release does not eliminate those obligations. An employee's apparent agreement also deserves scrutiny because a power imbalance can affect whether consent is freely given. Determine the applicable position for the jurisdictions, participants, and downstream uses in the actual collection program.

For a deliberately consented collection, explain the commercial purpose, types of recipients, permitted model training, retention, and how people can exercise their rights. Address bystanders, speech, reflected faces, screens, correspondence, badges, location clues, and private activity. A participant can agree to record themselves without being able to authorize every other person or confidential document that enters view.

Privacy preparation then becomes part of the data pipeline:

  1. Define excluded situations before recording, and give collectors a reliable pause or deletion process.
  2. Inspect footage, audio, metadata, transcripts, labels, and derived assets for identifying or sensitive content.
  3. Redact or exclude affected material, preserving an access-controlled record of the transformation and its effect on usability.
  4. Recheck the training copy. Redaction may obscure a hand-object contact or remove text the model was meant to learn from.
  5. Keep a release-to-episode map so a correction or removal can be propagated to delivered versions and recipients under the agreed process.

Replacing names with IDs may pseudonymise records; it does not automatically anonymise them. Pseudonymisation under GDPR Article 4 depends on identification requiring additional information held separately with appropriate safeguards. Face blurring alone also does not prove that voices, tattoos, room layouts, or other details cannot identify someone. Describe the measures and residual limitations accurately instead of labeling the dataset “anonymous” without evidence.

Preserve originals only for a documented, lawful retention need, under restricted access. A reproducibility goal is not a reason to retain every unredacted recording indefinitely. If originals must be removed, retain the permitted processing records, hashes, and version relationships that still support accountability.

Prepare recordings without destroying the measurements

Keep an inventory of what the capture device actually produced. For each session, record the camera and firmware, mount, resolution, frame-rate mode, lens or field of view, stabilization settings, exposure behavior where available, sensor streams, and calibration identity. If an existing archive lacks these details, state what is unknown. Do not fill gaps with settings inferred from a product brochure.

Preserve the timeline

Video frame number and sensor time are different things. A nominal 30 fps file does not establish that every exposure occurred exactly 33.33 milliseconds after the previous one. Dropped frames, variable-rate encoding, and independent clocks matter when an image is paired with a pose, contact event, or action.

For data synchronisation, retain original timestamps and their units, clock domain, session origin, and relationship to the delivered video. Document whether a timestamp refers to exposure, presentation, reception, or another event. Record offset and drift corrections, alignment tolerances, and how missing samples are handled. A file containing equal-length video and pose arrays is not evidence that the arrays are aligned.

Deliver frame presentation timestamps alongside source capture times when both are available. If only presentation timestamps exist, label them as such. Keep the mapping when trimming, stitching, or transcoding. A nearest-frame join, interpolation, and a held-last-sample join have different meanings; choose and document one rather than leaving the loader to guess.

Keep geometry interpretable

Tracked data needs coordinate frames, units, handedness, transform direction, quaternion order, and an explanation of the origin. Specify whether a pose maps a hand into the camera frame or into a session-local world frame. A world frame that resets at each recording does not provide a common map across the dataset.

Include camera intrinsics, distortion model, relevant extrinsics, and calibration versions for each rig configuration. Cropping, resizing, undistortion, and image rotation change how pixel coordinates correspond to rays; derived camera parameters and keypoints must describe the delivered images. Electronic stabilization can introduce time-varying geometry, so record its state and validate compatibility with the intended pose pipeline.

For model-estimated hands or body poses, record the estimator version, input files, units, confidence meaning, missing-data mask, filtering, and validation results. A confidence score is not a calibrated error bound unless that relationship was established. Do not convert an occluded hand into a plausible-looking coordinate and silently call it ground truth.

Optimize for the task before the file size

Use short calibration recordings to check whether relevant hands and objects remain visible during the entire action. For manipulation, evaluate blur and occlusion around contact and release, not just on idle frames. A high average visibility score can hide the exact moment the buyer needs.

Keep the highest-quality authorized source or archival representation and create a documented delivery derivative when needed. Avoid repeated lossy transcoding, overlays burned into training frames, or editing that removes setup and outcome context. A diagnostic video with skeletons and bounding boxes is valuable for inspection, but provide the underlying measurements separately from that visualization.

Annotate actions, outcomes, and uncertainty

Begin with the labels the buyer will actually use: an episode-level task, temporal action segments, relevant object identities, outcome, and reasons an episode is unsuitable. Add hand keypoints, masks, dense captions, or contact labels when a pilot establishes their value. More annotations also create more things that can be wrong.

Separate the intended instruction from the observed result. “Place the mug on the shelf” can be the task even when the mug falls. Record the failure and recovery instead of rewriting the instruction as if every take succeeded. For language annotations, distinguish participant narration, human transcription, annotator description, and model-generated captions.

EPIC-KITCHENS annotation pipeline linking narration, transcription, parsing, and temporal action segments

Annotation is a sequence of decisions with different error sources. The EPIC-KITCHENS pipeline connects participant narration to transcription, parsing, and temporal segmentation. Source: EPIC-KITCHENS; see the EPIC-KITCHENS-100 paper for the method.

Write an annotation guide with examples of ambiguous boundaries: when “open drawer” starts, whether a reach belongs to the same segment, how simultaneous actions are represented, and when an outcome is unknown. Store a vocabulary version and allow “unknown” or “not visible” when the evidence is insufficient.

Double-label a stratified sample, compare temporal boundaries and class agreement, and adjudicate disagreements. Report the sample size and sampling method alongside the result. A single overall agreement percentage can hide poor performance on rare tasks or one language. Model-generated labels need their own audit slice; human review should have a recorded status rather than being implied by the word “annotated.”

For data provenance, connect each derivative to its source episode, transformation, software or annotation version, and reviewer status. The W3C PROV model provides a useful conceptual separation between entities, activities, and responsible agents. You do not need a full semantic-web system to apply that principle: a versioned lineage table is already much better than folders named final, final2, and final_revised.

Learn from public datasets without confusing access with resale rights

The most useful comparison is what each project documents that a seller might otherwise omit. These examples are not a catalog of datasets you can repackage and sell.

Reference datasetConcrete packaging examplePreparation lessonCommercial and redistribution caveat
Ego4DVideo and clip identifiers, duration, codec, stream timing, parent-video offsets, and timestamped narrations in JSON schemasKeep source identity and timing through every clip operationIts license process and agreements distinguish permitted research/model development from prohibited database sale or onward access; inspect the applicable signed agreement
Ego-Exo4DFrame-aligned take MP4s, Aria VRS, and timesync.csv; metadata connects takes to capturesSupply a time mapping and capture hierarchy for multiple camerasRequires its own license, separate from Ego4D; access is not permission to redistribute the underlying database
EPIC-KITCHENS-100About 100 hours from 45 kitchens; CSV annotations contain narration, action start/stop times and frames, verb and noun classesNarration timestamps and action boundaries are separate annotationsPublished under CC BY-NC 4.0; commercial use requires separate terms
HOT3D and HOT3D-ClipsFull-sequence VRS; curated clips use TAR/WebDataset with images and camera, hand, and object JSONSeparate sensor data, pose availability, geometry, and QALicenses vary by asset: sequence/non-hand data CC BY-SA, hand annotations CC BY-NC-SA, object models under modified terms prohibiting sale or inclusion in a sold product
Apple EgoDexPaired MP4/HDF5 with camera intrinsics, joint transforms, and optional confidence arrays; the project reports 829 hoursMake reference frames, optional fields, and generated-label limitations explicitDataset CC BY-NC-ND; the code's license does not replace the dataset's terms

EgoDex is particularly instructive for pose delivery. Its documented HDF5 structure includes camera/intrinsic as a 3×3 matrix, transforms/<joint> as an N×4×4 array, and optional confidences/<joint> values. Its ARKit origin is stationary within an episode but not necessarily consistent across episodes. The README also warns that automatically generated language attributes can contain errors and that the synthesized camera view can introduce reprojection discrepancies. Those limitations belong in a commercial product description too.

HOT3D's Aria demonstration overlays hand-model contours in white and object-model contours in green on RGB video. It illustrates a geometry inspection view; the underlying images and annotations remain separate data. Source: HOT3D-Clips. Watch the original video.

These examples also show why “open” is an inadequate rights description. A permissive software toolkit can read non-commercial data. Commercial model training may be permitted while dataset resale remains prohibited. Different annotation layers in one release may carry different terms. Make the rights decision at asset and release level.

Choose a delivery format for the buyer's loader

Separate the source recording format from the training export. A useful delivery might preserve authorized native sensor logs while also providing a compact video-and-table view. Format conversion should preserve semantics and provenance; changing a file extension cannot create a missing action, calibration, or license.

FormatUseful roleWhat still needs a schema or agreement
MP4 + JSON/JSONL + CSV or ParquetAccessible video package with episode metadata, action segments, and frame-time sidecarsCodec, timestamps, identifiers, units, label vocabulary, and relationships between files
MP4 + HDF5Dense pose or other numeric arrays, as illustrated by EgoDexDataset paths, dimensions, compression, joint order, coordinate frames, confidence and missingness
LeRobot dataset v3A buyer's LeRobot-compatible training view using Parquet, MP4, and indexed metadataExporter/loader version, feature semantics, task/episode indexing, and actual available supervision
RLDSEpisodes with nested steps for compatible learning pipelines, often distributed through TensorFlow DatasetsStep structure, observation/action alignment, terminal versus truncated episodes, and optional fields
Aria VRS or MCAPNative multimodal recordings or timestamped message streamsSensor/topic schemas, clock domains, calibration, available reader, and any derived export
WebDataset/TARSharded image-and-annotation samples for streaming, illustrated by HOT3D-ClipsSample keys, shard membership, ordering, episode boundaries, and per-frame semantics

For plain human video, a documented MP4-plus-sidecars product is a reasonable baseline. Offer another format because a buyer has a loader for it, not because every robotics listing appears to need the same acronym.

A versioned LeRobot export

The LeRobot v3 architecture stores numeric records in Parquet, video in MP4, and episode boundaries in metadata. Several episodes can share a file. An episode is reconstructed from indexes and offsets, rather than being equivalent to one MP4 filename.

LeRobot dataset v3 diagram illustrating the organization of dataset files and metadata

LeRobot v3 groups records into shared files and uses metadata to recover episodes. Source: Hugging Face LeRobot documentation. The precise export layout should be pinned to a tested software version.

For example, the v0.6.1 implementation documents this style of layout for dataset format v3.0:

data/chunk-000/file-000.parquet
meta/info.json
meta/stats.json
meta/tasks.parquet
meta/episodes/chunk-000/file-000.parquet
videos/observation.images.front/chunk-000/file-000.mp4

The tagged path definitions use meta/tasks.parquet and also identify a legacy tasks.jsonl path. This is why a copied directory tree is not a compatibility test. Generate the export with the agreed writer, finalize it, reopen it using the buyer's pinned loader, and verify frame access and episode boundaries. Keep original sensor timestamps and calibration available separately when the training view simplifies them.

Do not manufacture actions to satisfy a schema

RLDS defines episodes and steps; it is not simply another name for a TFRecord file. In the general RLDS specification, is_first and is_last are required, while observations, actions, rewards, and other fields depend on the dataset. All steps within one dataset must have the same fields; optional in the general specification does not mean keys can disappear arbitrarily between steps. An observation-only dataset is possible. It does not supply training targets for a method that requires recorded controls.

RLDS also distinguishes the last step from a terminal state. A truncated recording can end without the task reaching a terminal condition. Actions, rewards, and discounts following the last observation are invalid. Preserve those meanings during conversion; filling missing actions with zeros turns absence into a false measurement.

If the buyer wants human-derived pseudo-actions, make that a separately specified transformation. Record the estimator, representation, validity mask, and target-embodiment mapping. Methods such as EgoMimic and UMI use deliberate alignment or capture interfaces. Their existence does not make ordinary head-camera footage an executable humanoid demonstration.

A worked example of a delivery package

Consider a hypothetical RGB-only dishwasher dataset. It has reviewed action segments but no measured hand poses, depth, IMU, or robot commands. The following is a proposed custom delivery design, not an official LeRobot or RLDS schema and not a real dataset offered for sale.

dishwasher-ego-v1.0/
  README.md
  LICENSE.txt
  CHANGELOG.md
  checksums.sha256
  schema/
    episode.schema.json
    segment.schema.json
    frame_times.schema.json
  metadata/
    episodes.jsonl
    splits.csv
    quality_report.json
    rights_summary.json
  episodes/
    ep_000001/
      rgb.mp4
      frame_times.csv
      segments.jsonl
  tools/
    requirements.lock
    load_sample.py
    validate_release.py

The dataset card in README.md should document scope, collection, intended uses, known gaps, schema, preparation, split policy, quality, license, and maintenance. Hugging Face's card guidance and Datasheets for Datasets are useful starting points. The card describes the release; the license and underlying agreements establish the actual rights.

An illustrative episode record could look like this. In episodes.jsonl, each complete object occupies one line; it is expanded here for readability.

{
  "schema_version": "1.0.0",
  "dataset_version": "1.0.0",
  "episode_id": "ep_000001",
  "source_session_id": "session_0007",
  "participant_id": "person_004",
  "site_id": "site_002",
  "task_id": "unload_dishwasher",
  "instruction": "Move the clean mugs from the dishwasher to the shelf.",
  "duration_ns": 60000000000,
  "video": {
    "path": "episodes/ep_000001/rgb.mp4",
    "codec": "h264",
    "width": 1920,
    "height": 1080,
    "nominal_fps": 30,
    "frame_count": 1800,
    "frame_index_base": 0,
    "time_base_num": 1,
    "time_base_den": 90000,
    "frame_times_path": "episodes/ep_000001/frame_times.csv"
  },
  "capture_clock": "device_monotonic",
  "capture_origin": "first_exposure_in_source_session",
  "source_start_ns": 120000000000,
  "camera_calibration_available": false,
  "audio_in_delivery": false,
  "hand_pose_available": false,
  "robot_action_available": false,
  "outcome": "success",
  "annotation_version": "1.0.0",
  "rights_record_ref": "rights_0017",
  "privacy_review": "passed",
  "quality_status": "accepted"
}

These values illustrate the schema, including a 60-second, 1,800-frame episode. Actual records must come from file inspection and documented review. “Calibration unavailable” is an honest limitation for this product; a buyer needing metric geometry may reject it. A rights_record_ref points to restricted evidence; the existence of that string does not prove consent or commercial permission.

For a capture pipeline that preserves source exposure timestamps, the first rows of frame_times.csv could be:

frame_index,video_pts,capture_time_ns
0,0,120000000000
1,3000,120033333333
2,6000,120066666667

Here video_pts uses the declared 1/90,000-second time base, while capture_time_ns is relative to the source session's first exposure. The episode starts 120 seconds into that session. This example has regular frame timing; a real table should retain irregularities instead of reconstructing capture times from nominal fps. If the source never recorded exposure timestamps, deliver presentation timestamps and disclose that limitation.

For long or absolute timestamps, use an integer-capable format such as Parquet int64, or decimal strings in JSON, so a consumer using JavaScript does not lose nanosecond precision. Document the representation consistently across the schema and loader.

An action segment in segments.jsonl can refer to the episode timeline without pretending to be a control command:

{
  "episode_id": "ep_000001",
  "segment_id": "seg_0003",
  "time_reference": "episode_start",
  "start_ns": 5200000000,
  "end_ns": 8700000000,
  "interval_convention": "[start,end)",
  "verb": "place",
  "object": "mug",
  "description": "Place the mug on the shelf.",
  "outcome": "success",
  "annotation_origin": "manual",
  "review_status": "second_annotator_checked",
  "annotation_version": "1.0.0",
  "robot_action_available": false
}

The half-open interval includes the start and excludes the end. State whether simultaneous segments can overlap and how labels map to frames. If you later add estimated poses, create a new version with a pose schema, model provenance, validity mask, calibration requirements, and measured accuracy; do not silently change what v1.0 contains.

The package's loading tool should open a selected episode, recover frame times, display its segments, and explain missing modalities. Its validation tool should check schema conformance, referenced files, checksums, decoded frame counts, interval bounds, and split membership. Publish the dependency lock and the command the buyer can run. These proposed tools are deliverables to build for the actual dataset, not scripts provided by this article.

Measure quality with denominators a buyer can audit

Write the acceptance policy before production. Distinguish release-blocking problems, such as unresolved rights or corrupt files, from graded attributes, such as hand visibility or difficult lighting. An imperfect but well-labeled failure episode can be useful; an episode with unknown rights is not repaired by a high image-quality score.

CheckEvidence to deliverHow to avoid an attractive but misleading number
IntegrityFile hashes, decoder results, frame counts, schema validationA checksum proves matching bytes, not correct labels or usable video
TimingClock definitions, missing-frame report, synchronization residuals and drift by streamEqual array lengths and nominal fps do not measure alignment
Task visibilityRequired-hand/object visibility during the annotated interaction, with reviewer rulesWhole-video averages can hide the grasp or release
GeometryValid-pose rate, accuracy on a defined reference subset, calibration and reprojection checksAvailability and confidence are not the same as accuracy
AnnotationsAudit sample, task/language slices, disagreement and correction counts“Human reviewed” needs a specified review procedure and coverage
CoverageAccepted episodes and duration by task, session, site, participant, and object instanceRepeated takes in one room are not independent environments
Privacy and rightsEpisode-linked approval status, restrictions, redaction/exclusion reportA blanket “cleared” claim cannot explain exceptions or mixed licenses
Duplication and splitsDuplicate groups and participant/site/session separation checksAdjacent clips from the same recording can leak into evaluation

For a buyer requiring two visible hands, an illustrative threshold might be “both required hands visible for at least 80% of the annotated interaction, with critical grasp and release moments reviewable.” The definition must specify the denominator, how occlusion is scored, and what happens when a task legitimately uses one hand. This is a negotiable criterion, not a universal definition of good egocentric video.

Timing thresholds should likewise follow the intended use. At 30 fps, one frame spans about 33.3 milliseconds. A tolerance acceptable for coarse activity recognition may be unsuitable for fine contact timing. Measure drift over long recordings as well as initial offset, and show the buyer the measurements behind any millisecond-level claim.

Build train/evaluation splits around the claim you want to support. For unseen-home evaluation, hold out sites; for unseen-person evaluation, hold out participants. Keep source sessions and their near-duplicate derivatives together. A simple 80/10/10 split by randomly sampled frames can expose the same person, objects, and almost identical moments in all three partitions. Fit normalization or learned preparation steps on training data only when they would otherwise leak evaluation information.

Count accepted activity separately from camera hours

Suppose a hypothetical single-camera collection begins with 100 hours. After excluding 8 hours for rights/privacy issues, 12 additional hours for capture or task failures, and 5 additional hours of duplicates or out-of-scope material, 75 hours remain. These are mutually exclusive deductions in this example; real overlapping rejection reasons must not be subtracted twice.

If only 60 of those 75 hours contain the agreed active task intervals, report both 75 accepted episode hours and 60 active-task hours. If two synchronized cameras record the same activity, that yields 150 accepted camera-stream hours, not 150 hours of unique activity. Declare which unit the contract bills, how partial clips count, and whether idle/setup time is included.

Do not translate a sample audit into a claim that every label was inspected. Report, for example, that a specified number of episodes were double-labeled, how they were selected, and which errors were corrected across the release. For a claimed model improvement, the buyer needs a separate controlled evaluation with a baseline, fixed held-out data, and uncertainty; passing a data-quality audit does not establish a performance gain.

Price the accepted deliverable and the license

There is no defensible universal rate per hour in the sources reviewed here. Collector compensation, a vendor's asking price, and a completed dataset license are different transactions. Price against the scope the buyer accepted in the pilot and the cost of delivering it reliably.

For a custom collection, account for recruitment, participant compensation, equipment, location access, collection management, rejected takes, annotation, privacy review, storage, transfer, and support. In a clearly hypothetical calculation, a project costing $12,000 that produces 60 accepted active-task hours has a cost of $200 per accepted active hour before margin and further obligations. That arithmetic is not a market quotation. Lower acceptance yield raises the effective cost even when collector pay stays unchanged.

Separate commercial terms that change the product's value:

  • Existing non-exclusive footage versus newly commissioned collection.
  • Evaluation access versus commercial training, deployment, affiliates, and contractors.
  • Raw observations versus reviewed annotations and validated derived geometry.
  • Non-exclusive licensing versus exclusivity defined by scope, period, and permitted prior commitments.
  • Dataset redistribution, derivative datasets, model weights, and generated outputs, each addressed explicitly.
  • Acceptance deadlines, inspection sampling, rejection evidence, rework limits, payment milestones, and correction/removal procedures.

“Exclusive” should describe what you can actually grant. An archive already licensed to other parties cannot honestly be sold as never previously shared. Similarly, a license allowing a buyer to train models does not automatically authorize that buyer to sell the recordings onward.

For delivery planning, compute size from a measured representative bitrate. At an illustrative 12 megabits per second, one camera produces approximately 5.4 decimal GB per hour, or 540 GB for 100 hours, before sidecars and other streams. Specify storage region, transfer method, egress responsibility, checksums, and a manifest that supports resuming or fetching selected episodes. Test download and decoding with the buyer before a multi-terabyte release.

Provide access through the agreed authenticated storage or dataset service, with a versioned release manifest and appropriate expiry or revocation controls. Avoid placing private source footage or consent evidence behind an unrestricted sample link. After delivery, retain a change log and a way to identify exactly which release and episodes each customer received.

Make the listing match the delivered product

A strong listing has a narrow task description, accepted unique duration, camera-stream duration, task and environment coverage, capture stack, measured versus derived modalities, annotation coverage, format and loader versions, sample access, rights scope, and known limitations. It also says whether the seller is offering existing inventory, custom collection, or an enrichment service.

Use the marketplace and vendor guide to compare sourcing channels, or submit a documented dataset through the Humanoids Data seller form. If the pilot reveals that the buyer needs calibrated hand trajectories and your archive contains only RGB, the next decision is whether an explicitly validated enrichment step is worthwhile. A precise description of that gap is more useful than a broad promise that the video is ready for every robot.