AgiBot World 2026 Dataset
Large-scale real-world humanoid interaction data
A multi-scenario real-world dataset with LeRobot structure, key frames, instruction segments, and paired digital-twin resources.
Humanoid VLA training
Non-commercial share-alike license; commercial use requires separate legal review or permission.
Start with an official sample or one task; do not download the full release before schema validation.
Has observation, action/state proxy, and task or language context.
fit 85 · confidence 55
Has visual observations plus geometry, calibration, depth, or reconstruction cues.
fit 82 · confidence 55
Contains observation, intent, action/state, and feedback-like supervision.
fit 88 · confidence 55
Verified facts and provenance
Claims, metadata verification, and sample verification are shown separately.
is_result_succeedError FrameSuccess FrameIntervention Frametake_overerror_causerestorableinstruction_segmentsTask Frame2D Bounding Boxobservation.stateactionobservation.images.*task_indextimestampUnknown — metadata conclusions do not prove sample coverage, alignment, or file integrity.
Declared loop signal coverage
Signals inferred from official metadata; Data pipeline verification is still pending.
Observation / ego video
video · depth
Action / hand pose / robot state
actions · robot state · Large-scale real-world humanoid interaction data
Gaze / attention
No decision-grade evidence captured yet.
Language intent / task phase
language · task phase · skill annotations · Long-horizon task decomposition
Feedback / correction / failure
success/failure · Failure and intervention mining · Failure, success, intervention, takeover, and recovery fields are declared, but their coverage is not sample-verified by OpenBot.
Sim-real pairing
depth · sim-real · Sim-real pairing
License / format / access
Open · CC BY-NC-SA 4.0 · LeRobot v2.1 · MP4
Model and task fit · OpenBot inference
Has observation, action/state proxy, and task or language context.
fit 85 · confidence 55
Has visual observations plus geometry, calibration, depth, or reconstruction cues.
fit 82 · confidence 55
Contains observation, intent, action/state, and feedback-like supervision.
fit 88 · confidence 55
Has failure/evaluation-style labels with action or manipulation context.
fit 82 · confidence 55
Good tasks
Blockers and unresolved evidence
- Gaze / attentionunknownNot enough evidence to classify this signal. Verify metadata or a bounded sample.
Raw dataset signals
OpenBot fit
- Humanoid VLA training
- Hierarchical policy learning
- Skill-level imitation learning
- Long-horizon task decomposition
- Single-instruction LeRobot training
- Object-conditioned policy
- Language-region grounding
- Failure and intervention mining
- Sim-real pairing
Related models and papers
Model references linked to similar loop signals.
LingBot-VLA
A 4B pragmatic VLA foundation model pretrained on roughly 20,000 hours of real-world data from nine dual-arm robot configurations.
LingBot-VLA 2.0
A whole-body VLA release focused on broader cross-embodiment transfer, mobile manipulation, and predictive dynamics supervision.
Helix 02
A hierarchical whole-body VLA for continuous humanoid locomotion, manipulation, balance, tactile feedback, and long-horizon autonomy.
GR00T 1.7
NVIDIA's current open humanoid VLA release within an end-to-end Isaac workflow for teleoperation, training, evaluation, and deployment.
Galaxea G0
A dual-system VLM plus VLA model for planning and fine-grained control on long-horizon mobile-manipulation tasks.
LingBot-World
An open interactive world simulator derived from video generation, with camera- and action-conditioned variants and long-horizon generation.
Integration notes
- Non-commercial share-alike license; commercial use requires separate legal review or permission.
- The official card provides a roughly 7GB task_3777 sample; validate this before downloading the 9.36TB full release.
- Failure, success, intervention, takeover, and recovery fields are declared, but their coverage is not sample-verified by OpenBot.
