OpenBot
Models & papers57 models · 5 families · published

Embodied foundation models need loop data.

A curated map of model families, links, and the loop signals each one needs.

Reading frame
Data requirements, not hype
Families
VLA · World Model · WAM · Policy
Core signals
observation · intent · action · feedback
Catalog link
dataset coverage becomes model readiness
Output
what to train, evaluate, or collect next
Signal coverage

Papers point back to missing data.

Each model is a demand signal for catalog fields, curation, evaluation cases, and replay.

43 model refs
Actions
33 model refs
Observation
23 model refs
Language intent
23 model refs
Robot state
14 model refs
Language
11 model refs
Video
9 model refs
Task phase
8 model refs
Future state
7 model refs
Depth
7 model refs
Feedback/failure
7 model refs
Vision Language Action
6 model refs
Task success
Model index

Search models by family, openness, and loop-data demand.

A practical map for finding which datasets support training, world modeling, WAM-style action, or failure mining.

Filters

Showing 57 of 57 models

WAMResearch2026

DreamZero / World Action Models

Research

Detail

World-action model framing where future visual states and actions are predicted in an aligned policy model.

Recommended for

Zero-shot policy behavior

Current blocker

This is the strongest reason to keep Catalog focused on loop traces instead of flat dataset metadata.

Code, weights, and checkpoints
confidence 10
Loading and training reproducibility
confidence 15
Training data requirements
76confidence 50
Loop data needed
  • Observation, intent, action, and future-state supervision
  • Robot demonstrations with paired video/action sequences
  • Failure and correction signals for validating imagined futures
Dataset signals
ObservationActionsFuture stateLanguage intentFeedback/failure
curated source· verification pending· code unknown· weights unknown
Evaluate in Bench
VLAOpen2026model hub

GR00T 1.7

NVIDIA

Detail

NVIDIA's current open humanoid VLA release within an end-to-end Isaac workflow for teleoperation, training, evaluation, and deployment.

Recommended for

Humanoid cross-embodiment deployment

Current blocker

Track as the latest release in the GR00T series, not as a replacement for historical N1 results.

Code, weights, and checkpoints
65confidence 40
Loading and training reproducibility
confidence 15
Training data requirements
76confidence 50
Loop data needed
  • Human demonstrations
  • Simulation trajectories
  • Whole-body robot state
Dataset signals
ObservationLanguage intentActionsRobot stateSim-realTask phase
Model hub
curated source· verification pending· code unknown· weights declared
Evaluate in Bench
VLAOpen2026-02-25model hub

GR00T-N1.7-3B

nvidia

Detail

Official model catalog record.

Recommended for

robotics

Current blocker

License must be verified before commercial use

Access and governance
confidence 20
Artifact availability
90confidence 90
Training and loading reproducibility
70confidence 65
Loop data needed
  • video
  • actions
  • robot state
Dataset signals
VideoActionsRobot stateLanguageSuccess Failure
Related catalog data
No direct catalog match yet.
Link
official claim· verified 2026-07-22· code unknown· weights metadata verified
Evaluate in Bench
VLAClosed ref2026

Helix 02

Figure

Detail

A hierarchical whole-body VLA for continuous humanoid locomotion, manipulation, balance, tactile feedback, and long-horizon autonomy.

Recommended for

Whole-body loco-manipulation

Current blocker

Highlights why Catalog needs humanoid, tactile, and whole-body signal fields.

Code, weights, and checkpoints
confidence 10
Loading and training reproducibility
confidence 15
Training data requirements
76confidence 50
Loop data needed
  • Whole-body human motion
  • Humanoid proprioception
  • Palm and head cameras
Dataset signals
ObservationTactileRobot stateActionsFeedback/failureTask phase
curated source· verification pending· code unknown· weights unknown
Evaluate in Bench
PerceptionOpen2026-01-25

lingbot-depth

Robbyant

Detail

Masked Depth Modeling for Spatial Perception

Recommended for

perception

Current blocker

Verify checkpoint and evaluation compatibility.

Access and governance
75confidence 85
Artifact availability
confidence 15
Training and loading reproducibility
confidence 15
Loop data needed
  • depth
Dataset signals
Depth
Related catalog data
official claim· verified 2026-07-22· code unknown· weights unknown
Evaluate in Bench
PerceptionOpen2026-04-15

lingbot-map

Robbyant

Detail

A feed-forward 3D foundation model for reconstructing scenes from streaming data

Recommended for

Streaming reconstruction

Current blocker

Verify checkpoint and evaluation compatibility.

Access and governance
75confidence 85
Artifact availability
confidence 15
Training and loading reproducibility
confidence 15
Loop data needed
  • Streaming image sequences
  • Camera trajectories
  • Metric geometry
Dataset signals
ObservationCamera poseDepthPoint CloudTrajectory
Related catalog data
official claim· verified 2026-07-22· code unknown· weights unknown
Evaluate in Bench
VLAOpen2026-01-29

lingbot-va

Robbyant

Detail

[RSS 2026] Causal video-action world model for generalist robot control

Recommended for

world modeling

Current blocker

Verify checkpoint and evaluation compatibility.

Access and governance
75confidence 85
Artifact availability
confidence 15
Training and loading reproducibility
confidence 15
Loop data needed
  • video
  • actions
Dataset signals
VideoActions
official claim· verified 2026-07-22· code unknown· weights unknown
Evaluate in Bench
VLAOpen2026-07-08model hub

lingbot-video-dense-1.3b

robbyant

Detail

Official model catalog record.

Recommended for

VLA evaluation

Current blocker

License must be verified before commercial use

Access and governance
confidence 20
Artifact availability
90confidence 90
Training and loading reproducibility
70confidence 65
Loop data needed
  • video
  • diffusers:LingBotVideoPipeline
Dataset signals
VideoDiffusers:lingbotvideopipeline
Related catalog data
No direct catalog match yet.
Link
official claim· verified 2026-07-22· code unknown· weights metadata verified
Evaluate in Bench
PerceptionOpen2026-07-06

lingbot-vision

Robbyant

Detail

Self-supervised learning for spatial perception

Recommended for

perception

Current blocker

Verify checkpoint and evaluation compatibility.

Access and governance
75confidence 85
Artifact availability
confidence 15
Training and loading reproducibility
confidence 15
Loop data needed
  • Large-scale visual pretraining data
  • Dense spatial labels
  • Depth and geometry benchmarks
Dataset signals
ObservationDepthObjectsPoseCamera Calibration
Related catalog data
official claim· verified 2026-07-22· code unknown· weights unknown
Evaluate in Bench
VLAOpen2026-01-26

lingbot-vla

Robbyant

Detail

A Pragmatic VLA Foundation Model

Recommended for

VLA training

Current blocker

Verify checkpoint and evaluation compatibility.

Access and governance
75confidence 85
Artifact availability
confidence 15
Training and loading reproducibility
confidence 15
Loop data needed
  • Large-scale dual-arm demonstrations
  • Cross-embodiment action normalization
  • Optional depth supervision
Dataset signals
ObservationDepthLanguage intentActionsRobot state
official claim· verified 2026-07-22· code unknown· weights unknown
Evaluate in Bench
VLAOpen2026model hub

LingBot-VLA 2.0

Robbyant

Detail

A whole-body VLA release focused on broader cross-embodiment transfer, mobile manipulation, and predictive dynamics supervision.

Recommended for

Whole-body action generation

Current blocker

Track as a release in the LingBot-VLA series rather than an unrelated model.

Code, weights, and checkpoints
65confidence 40
Loading and training reproducibility
confidence 15
Training data requirements
76confidence 50
Loop data needed
  • Whole-body action traces
  • Multiple robot configurations
  • Human video paired with robot data
Dataset signals
ObservationFuture stateLanguage intentActionsRobot stateTask phase
Model hub
curated source· verification pending· code declared· weights declared
Evaluate in Bench
VLAOpen2026-01-26model hub

lingbot-vla-4b

robbyant

Detail

Official model catalog record.

Recommended for

VLA training

Current blocker

License must be verified before commercial use

Access and governance
confidence 20
Artifact availability
90confidence 90
Training and loading reproducibility
70confidence 65
Loop data needed
  • depth
Dataset signals
Depth
Related catalog data
No direct catalog match yet.
Link
official claim· verified 2026-07-22· code unknown· weights metadata verified
Evaluate in Bench
VLAOpen2026-01-26model hub

lingbot-vla-4b-depth

robbyant

Detail

Official model catalog record.

Recommended for

VLA evaluation

Current blocker

License must be verified before commercial use

Access and governance
confidence 20
Artifact availability
65confidence 20
Training and loading reproducibility
confidence 15
Loop data needed
    Dataset signals
    Related catalog data
    No direct catalog match yet.
    Link
    official claim· verified 2026-07-14· code unknown· weights metadata verified
    Evaluate in Bench
    VLAOpen2026-03-09model hub

    lingbot-vla-4b-posttrain-robotwin

    robbyant

    Detail

    Official model catalog record.

    Recommended for

    VLA training

    Current blocker

    License must be verified before commercial use

    Access and governance
    confidence 20
    Artifact availability
    90confidence 90
    Training and loading reproducibility
    70confidence 65
    Loop data needed
      Dataset signals
      Related catalog data
      No direct catalog match yet.
      Link
      official claim· verified 2026-07-22· code unknown· weights metadata verified
      Evaluate in Bench
      VLAOpen2026-07-07

      lingbot-vla-v2

      Robbyant

      Detail

      From Foundation to Application

      Recommended for

      VLA training

      Current blocker

      Verify checkpoint and evaluation compatibility.

      Access and governance
      75confidence 85
      Artifact availability
      confidence 15
      Training and loading reproducibility
      confidence 15
      Loop data needed
        Dataset signals
        Related catalog data
        No direct catalog match yet.
        official claim· verified 2026-07-22· code unknown· weights unknown
        Evaluate in Bench
        World ModelOpen2026model hub

        LingBot-World

        Robbyant

        Detail

        An open interactive world simulator derived from video generation, with camera- and action-conditioned variants and long-horizon generation.

        Recommended for

        Long-horizon consistency

        Current blocker

        Base (Cam), Base (Act), and Fast are releases/checkpoints of one model series.

        Code, weights, and checkpoints
        65confidence 40
        Loading and training reproducibility
        confidence 15
        Training data requirements
        76confidence 50
        Loop data needed
        • Action- or camera-conditioned video
        • Long temporal sequences
        • Consistent scene dynamics
        Dataset signals
        ObservationFuture stateActionsCamera poseLanguage intent
        Model hubModel hub
        curated source· verification pending· code declared· weights declared
        Evaluate in Bench
        VLAOpen2026-07-08

        lingbot-world-v2

        Robbyant

        Detail

        Infinite Worlds with Versatile Interactions

        Recommended for

        VLA evaluation

        Current blocker

        Verify checkpoint and evaluation compatibility.

        Access and governance
        75confidence 85
        Artifact availability
        confidence 15
        Training and loading reproducibility
        confidence 15
        Loop data needed
          Dataset signals
          Related catalog data
          No direct catalog match yet.
          official claim· verified 2026-07-22· code unknown· weights unknown
          Evaluate in Bench
          World ModelResearch2026

          Looped World Models

          Facemind / research

          Detail

          World-model architecture using looped transformer refinement over latent states for iterative environment understanding.

          Recommended for

          Data efficiency

          Current blocker

          Important trend signal for why OpenBot should model feedback and correction, not only observations.

          Code, weights, and checkpoints
          confidence 10
          Loading and training reproducibility
          confidence 15
          Training data requirements
          76confidence 50
          Loop data needed
          • Temporal egocentric and human-centric observations
          • Dense signals that expose feedback and state correction
          • Benchmarks that measure data efficiency and rollout stability
          Dataset signals
          ObservationFeedbackFuture stateGaze/attentionHuman demonstration
          curated source· verification pending· code unknown· weights unknown
          Evaluate in Bench
          WAMResearch2026

          OA-WAM

          Research

          Detail

          Object-addressable world-action model that decomposes scenes into robot and object slots while jointly predicting future world state and actions.

          Recommended for

          Object identity under scene shifts

          Current blocker

          Good example of why object-level and task-phase signals matter for OpenBot dataset detail pages.

          Code, weights, and checkpoints
          confidence 10
          Loading and training reproducibility
          confidence 15
          Training data requirements
          76confidence 50
          Loop data needed
          • Object-centric interaction traces
          • Language instructions tied to particular objects
          • Visual, proprioceptive, and action tokens over time
          Dataset signals
          ObservationObject labelsActionsRobot stateLanguage intent
          curated source· verification pending· code unknown· weights unknown
          Evaluate in Bench
          VLAOpen2026-04-20model hub

          resnet10

          lerobot

          Detail

          Official model catalog record.

          Recommended for

          Lerobot

          Current blocker

          License must be verified before commercial use

          Access and governance
          confidence 20
          Artifact availability
          90confidence 90
          Training and loading reproducibility
          70confidence 65
          Loop data needed
          • video
          • feature-extraction
          • image-classification
          Dataset signals
          VideoFeature ExtractionImage Classification
          Related catalog data
          No direct catalog match yet.
          Link
          official claim· verified 2026-07-22· code unknown· weights metadata verified
          Evaluate in Bench
          VLAOpen2026-05-21model hub

          Robometer-4B

          lerobot

          Detail

          Official model catalog record.

          Recommended for

          lerobot

          Current blocker

          License must be verified before commercial use

          Access and governance
          confidence 20
          Artifact availability
          90confidence 90
          Training and loading reproducibility
          70confidence 65
          Loop data needed
          • language
          • vision-language
          Dataset signals
          LanguageVision Language
          Related catalog data
          No direct catalog match yet.
          Link
          official claim· verified 2026-07-22· code unknown· weights metadata verified
          Evaluate in Bench
          VLAOpen2026-05-28model hub

          VLA-JEPA-Pretrain

          lerobot

          Detail

          Official model catalog record.

          Recommended for

          VLA training

          Current blocker

          License must be verified before commercial use

          Access and governance
          confidence 20
          Artifact availability
          90confidence 90
          Training and loading reproducibility
          70confidence 65
          Loop data needed
            Dataset signals
            Related catalog data
            No direct catalog match yet.
            Link
            official claim· verified 2026-07-22· code unknown· weights metadata verified
            Evaluate in Bench
            VLAOpen2026model hub

            WALL-OSS 0.5

            X Square Robot

            Detail

            An open 4B VLA pretrained across more than 20 embodiments, with physical-hardware evaluation of pretrained robotic capability.

            Recommended for

            Zero-shot physical capability

            Current blocker

            Officially integrated into LeRobot; self-reported results should remain separate from independent reproduction.

            Code, weights, and checkpoints
            65confidence 40
            Loading and training reproducibility
            confidence 15
            Training data requirements
            76confidence 50
            Loop data needed
            • Large cross-embodiment trajectory mixtures
            • Grounded multimodal data
            • Task decomposition and continuous actions
            Dataset signals
            ObservationLanguage intentTask phaseActionsRobot state
            Model hubDocs
            curated source· verification pending· code declared· weights declared
            Evaluate in Bench
            WAMOpen2025model hub

            EO-1

            IPEC / EO-Robotics

            Detail

            A unified embodied model using interleaved vision-text-action pretraining for perception, reasoning, planning, and continuous robot control.

            Recommended for

            Unified reasoning and action

            Current blocker

            Paired with EO-Data1.5M and EO-Bench; keep training and benchmark relations explicit.

            Code, weights, and checkpoints
            65confidence 40
            Loading and training reproducibility
            confidence 15
            Training data requirements
            76confidence 50
            Loop data needed
            • Interleaved vision-text-action samples
            • Embodied reasoning traces
            • Continuous robot actions
            Dataset signals
            ObservationLanguage intentTask phaseActionsFuture state
            Model hub
            curated source· verification pending· code unknown· weights declared
            Evaluate in Bench
            VLAOpen2025-08-28model hub

            EO-1-3B

            IPEC-COMMUNITY

            Detail

            Official model catalog record.

            Recommended for

            VLA training

            Current blocker

            Verify checkpoint and evaluation compatibility.

            Access and governance
            75confidence 85
            Artifact availability
            90confidence 90
            Training and loading reproducibility
            70confidence 65
            Loop data needed
            • video
            • actions
            • robot state
            Dataset signals
            VideoActionsRobot stateLanguageFeature ExtractionDataset:lmms LAB Llava Video 178k
            Related catalog data
            No direct catalog match yet.
            Link
            official claim· verified 2026-07-22· code unknown· weights metadata verified
            Evaluate in Bench
            VLAOpen2025model hub

            Galaxea G0

            OpenGalaxea

            Detail

            A dual-system VLM plus VLA model for planning and fine-grained control on long-horizon mobile-manipulation tasks.

            Recommended for

            Long-horizon mobile manipulation

            Current blocker

            Directly paired with the Galaxea Open-World Dataset.

            Code, weights, and checkpoints
            65confidence 40
            Loading and training reproducibility
            confidence 15
            Training data requirements
            76confidence 50
            Loop data needed
            • Subtask-level language
            • Mobile dual-arm trajectories
            • Long-horizon task structure
            Dataset signals
            ObservationLanguage intentTask phaseActionsRobot stateOutcome
            Model hub
            curated source· verification pending· code unknown· weights declared
            Evaluate in Bench
            VLAClosed ref2025

            Gemini Robotics

            Google DeepMind

            Detail

            Embodied reasoning and action model family that brings Gemini capabilities into robotics through vision, language, and physical action.

            Recommended for

            Physical reasoning and instruction following

            Current blocker

            A closed reference point, but useful for tracking the VLA/embodied foundation model direction.

            Code, weights, and checkpoints
            confidence 10
            Loading and training reproducibility
            confidence 15
            Training data requirements
            76confidence 50
            Loop data needed
            • Vision-language interaction data connected to executable robot skills
            • Generalization tests across objects, scenes, and task instructions
            • Safety, affordance, and failure feedback for physical-world deployment
            Dataset signals
            ObservationLanguage intentActionsObject labelsFeedback/failure
            curated source· verification pending· code unknown· weights unknown
            Evaluate in Bench
            VLAClosed ref2025

            Gemini Robotics 1.5

            Google DeepMind

            Detail

            An agentic robotics model combining advanced multimodal reasoning with vision-language-action control for multi-step physical tasks.

            Recommended for

            Multi-step reasoning

            Current blocker

            Keep availability distinct from Gemini Robotics On-Device and ER releases.

            Code, weights, and checkpoints
            confidence 10
            Loading and training reproducibility
            confidence 15
            Training data requirements
            76confidence 50
            Loop data needed
            • Multi-step instructions
            • Visual observations
            • Action traces
            Dataset signals
            ObservationLanguage intentTask phaseActionsFeedback/failure
            curated source· verification pending· code unknown· weights unknown
            Evaluate in Bench
            VLAOpen2025model hub

            GR00T N1

            NVIDIA

            Detail

            Open foundation model for generalist humanoid robots, focused on whole-body and manipulation behavior from multimodal robot data.

            Recommended for

            Generalist humanoid task transfer

            Current blocker

            Pushes OpenBot tags toward embodiment-aware metadata rather than generic video/action labels.

            Code, weights, and checkpoints
            65confidence 40
            Loading and training reproducibility
            confidence 15
            Training data requirements
            76confidence 50
            Loop data needed
            • Humanoid or whole-body demonstrations with proprioception, actions, and visual context
            • Task-level language or intent metadata
            • Sim-real and embodiment metadata for evaluating transfer to physical robots
            Dataset signals
            ObservationLanguage intentActionsRobot stateSim-real
            Model hub
            curated source· verification pending· code unknown· weights declared
            Evaluate in Bench
            World ModelOpen2025model hub

            NVIDIA Cosmos

            NVIDIA

            Detail

            World foundation model platform for physical AI, built around predictive world modeling and data processing workflows.

            Recommended for

            Future-state prediction quality

            Current blocker

            Useful anchor for the world-model side of OpenBot's Catalog and Synth story.

            Code, weights, and checkpoints
            65confidence 40
            Loading and training reproducibility
            confidence 15
            Training data requirements
            76confidence 50
            Loop data needed
            • Large-scale video and interaction data
            • Geometry, depth, calibration, or simulation context
            • Evaluation traces that connect prediction quality to downstream policy improvement
            Dataset signals
            ObservationDepthCamera poseSim-realFuture state
            Model hubModel hub
            curated source· verification pending· code unknown· weights declared
            Evaluate in Bench
            VLAOpen2025

            OpenVLA-OFT

            OpenVLA collaboration

            Detail

            An optimized OpenVLA fine-tuning recipe with continuous action chunks, multi-image input, and faster high-frequency control.

            Recommended for

            Fine-tuning speed

            Current blocker

            Treat as an OpenVLA release/recipe in the future series schema, despite the current flat static index.

            Code, weights, and checkpoints
            58confidence 40
            Loading and training reproducibility
            confidence 15
            Training data requirements
            76confidence 50
            Loop data needed
            • Task-specific robot demonstrations
            • Multiple camera views
            • Continuous action chunks
            Dataset signals
            ObservationLanguage intentActionsRobot stateCamera pose
            curated source· verification pending· code declared· weights unknown
            Evaluate in Bench
            VLAOpen2025-09-09model hub

            pi0_base

            lerobot

            Detail

            Official model catalog record.

            Recommended for

            VLA training

            Current blocker

            License must be verified before commercial use

            Access and governance
            confidence 20
            Artifact availability
            90confidence 90
            Training and loading reproducibility
            70confidence 65
            Loop data needed
            • actions
            • language
            • vision-language-action
            Dataset signals
            ActionsLanguageVision Language Action
            Related catalog data
            No direct catalog match yet.
            Link
            official claim· verified 2026-07-22· code unknown· weights metadata verified
            Evaluate in Bench
            VLAResearch2025

            pi0.5

            Physical Intelligence

            Detail

            A VLA model designed for open-world generalization by combining heterogeneous robot data, semantic subtask prediction, and high-level knowledge transfer.

            Recommended for

            Open-world generalization

            Current blocker

            A useful anchor for Catalog fields around task phase and heterogeneous data mixtures.

            Code, weights, and checkpoints
            confidence 10
            Loading and training reproducibility
            confidence 15
            Training data requirements
            76confidence 50
            Loop data needed
            • Heterogeneous robot demonstrations
            • High-level subtask labels
            • Language instructions
            Dataset signals
            ObservationLanguage intentTask phaseActionsRobot state
            curated source· verification pending· code unknown· weights unknown
            Evaluate in Bench
            VLAOpen2025-09-09model hub

            pi05_base

            lerobot

            Detail

            Official model catalog record.

            Recommended for

            VLA training

            Current blocker

            License must be verified before commercial use

            Access and governance
            confidence 20
            Artifact availability
            90confidence 90
            Training and loading reproducibility
            70confidence 65
            Loop data needed
            • actions
            • language
            • vision-language-action
            Dataset signals
            ActionsLanguageVision Language Action
            Related catalog data
            No direct catalog match yet.
            Link
            official claim· verified 2026-07-22· code unknown· weights metadata verified
            Evaluate in Bench
            VLAOpen2025-09-09model hub

            pi05_libero_base

            lerobot

            Detail

            Official model catalog record.

            Recommended for

            VLA training

            Current blocker

            License must be verified before commercial use

            Access and governance
            confidence 20
            Artifact availability
            90confidence 90
            Training and loading reproducibility
            70confidence 65
            Loop data needed
            • actions
            • language
            • vision-language-action
            Dataset signals
            ActionsLanguageVision Language Action
            Related catalog data
            No direct catalog match yet.
            Link
            official claim· verified 2026-07-22· code unknown· weights metadata verified
            Evaluate in Bench
            VLAOpen2025model hub

            SmolVLA

            Hugging Face LeRobot

            Detail

            Compact open vision-language-action policy designed for practical robot fine-tuning and deployment through the LeRobot ecosystem.

            Recommended for

            Data-efficient fine-tuning

            Current blocker

            Strong fit for OpenBot's model-readiness view because public weights make the training path concrete.

            Code, weights, and checkpoints
            65confidence 40
            Loading and training reproducibility
            confidence 15
            Training data requirements
            76confidence 50
            Loop data needed
            • LeRobot-style trajectories with observations, actions, and task text
            • Small but clean demonstrations that preserve episode boundaries and action timing
            • Evaluation splits that measure fine-tuning data efficiency
            Dataset signals
            ObservationLanguage intentActionsRobot stateTask success
            PostModel hub
            curated source· verification pending· code declared· weights declared
            Evaluate in Bench
            VLAOpen2025-06-01model hub

            smolvla_base

            lerobot

            Detail

            Official model catalog record.

            Recommended for

            VLA training

            Current blocker

            License must be verified before commercial use

            Access and governance
            confidence 20
            Artifact availability
            90confidence 90
            Training and loading reproducibility
            70confidence 65
            Loop data needed
            • video
            • actions
            • language
            Dataset signals
            VideoActionsLanguageVision Language Action
            Related catalog data
            No direct catalog match yet.
            Link
            official claim· verified 2026-07-22· code unknown· weights metadata verified
            Evaluate in Bench
            VLAOpen2025model hub

            SpatialVLA

            Shanghai AI Laboratory / collaborators

            Detail

            A spatially enhanced 4B VLA pretrained on 1.1 million real-robot episodes with explicit geometry-aware representations.

            Recommended for

            Spatial grounding

            Current blocker

            Official training references Open X-Embodiment and RH20T.

            Code, weights, and checkpoints
            65confidence 40
            Loading and training reproducibility
            confidence 15
            Training data requirements
            76confidence 50
            Loop data needed
            • Large-scale real robot episodes
            • Depth and spatial supervision
            • RLDS action/state alignment
            Dataset signals
            ObservationDepthLanguage intentActionsRobot state
            Model hub
            curated source· verification pending· code declared· weights declared
            Evaluate in Bench
            VLAOpen2025-10-09

            starVLA

            starVLA

            Detail

            StarVLA: A Lego-like Codebase for Vision-Language-Action Model Developing

            Recommended for

            VLA training

            Current blocker

            Verify checkpoint and evaluation compatibility.

            Access and governance
            75confidence 85
            Artifact availability
            55confidence 55
            Training and loading reproducibility
            confidence 15
            Loop data needed
            • video
            • actions
            • language
            Dataset signals
            VideoActionsLanguageForce Tactile
            official claim· verified 2026-07-22· code metadata verified· weights unknown
            Evaluate in Bench
            VLAOpen2025-09-25

            X-VLA

            2toinf

            Detail

            [ICLR 2026] The offical Implementation of "Soft-Prompted Transformer as Scalable Cross-Embodiment Vision-Language-Action Model"

            Recommended for

            VLA training

            Current blocker

            Verify checkpoint and evaluation compatibility.

            Access and governance
            75confidence 85
            Artifact availability
            55confidence 55
            Training and loading reproducibility
            70confidence 65
            Loop data needed
            • video
            • actions
            • robot state
            Dataset signals
            VideoActionsRobot stateLanguageSuccess Failure
            official claim· verified 2026-07-23· code metadata verified· weights unknown
            Evaluate in Bench
            VLAOpen2025-11-19model hub

            xvla-agibot-world

            lerobot

            Detail

            Official model catalog record.

            Recommended for

            VLA training

            Current blocker

            License must be verified before commercial use

            Access and governance
            confidence 20
            Artifact availability
            90confidence 90
            Training and loading reproducibility
            70confidence 65
            Loop data needed
            • actions
            • language
            • vision-language-action
            Dataset signals
            ActionsLanguageVision Language Action
            Related catalog data
            No direct catalog match yet.
            Link
            official claim· verified 2026-07-22· code unknown· weights metadata verified
            Evaluate in Bench
            VLAOpen2025-11-19model hub

            xvla-base

            lerobot

            Detail

            Official model catalog record.

            Recommended for

            VLA training

            Current blocker

            Verify checkpoint and evaluation compatibility.

            Access and governance
            75confidence 85
            Artifact availability
            90confidence 90
            Training and loading reproducibility
            70confidence 65
            Loop data needed
            • video
            • actions
            • language
            Dataset signals
            VideoActionsLanguageVision Language Action
            Related catalog data
            No direct catalog match yet.
            Link
            official claim· verified 2026-07-22· code unknown· weights metadata verified
            Evaluate in Bench
            VLAOpen2025-12-02model hub

            xvla-libero

            lerobot

            Detail

            Official model catalog record.

            Recommended for

            VLA training

            Current blocker

            License must be verified before commercial use

            Access and governance
            confidence 20
            Artifact availability
            90confidence 90
            Training and loading reproducibility
            70confidence 65
            Loop data needed
            • actions
            • language
            • vision-language-action
            Dataset signals
            ActionsLanguageVision Language Action
            Related catalog data
            No direct catalog match yet.
            Link
            official claim· verified 2026-07-22· code unknown· weights metadata verified
            Evaluate in Bench
            World ModelClosed ref2024

            Genie 2

            Google DeepMind

            Detail

            Large-scale foundation world model for generating action-controllable interactive environments from visual prompts.

            Recommended for

            Controllable world generation

            Current blocker

            Not a robot policy, but a useful trend marker for the world-model side of embodied AI.

            Code, weights, and checkpoints
            confidence 10
            Loading and training reproducibility
            confidence 15
            Training data requirements
            76confidence 50
            Loop data needed
            • Action-controllable video or environment interaction traces
            • Consistent observations over time with enough geometry and dynamics
            • Benchmarks that test controllability, persistence, and physical plausibility
            Dataset signals
            ObservationActionsFuture stateSim-realFeedback/failure
            curated source· verification pending· code unknown· weights unknown
            Evaluate in Bench
            PolicyOpen2024model hub

            Octo

            Octo Model Team

            Detail

            Open-source generalist robot policy pretrained on Open X-Embodiment trajectories and designed for fine-tuning to new robots and tasks.

            Recommended for

            Fine-tuning to new observation spaces

            Current blocker

            Good bridge between pure policy learning and broader VLA/WAM framing.

            Code, weights, and checkpoints
            65confidence 40
            Loading and training reproducibility
            confidence 15
            Training data requirements
            76confidence 50
            Loop data needed
            • Robot trajectories with observations and actions
            • Flexible task definitions such as language or goal images
            • Sensor/action-space metadata for adaptation
            Dataset signals
            ObservationActionsRobot stateLanguage intentGoal image
            Model hub
            curated source· verification pending· code unknown· weights declared
            Evaluate in Bench
            VLAOpen2024-06-13

            openvla

            openvla

            Detail

            OpenVLA: An open-source vision-language-action model for robotic manipulation.

            Recommended for

            VLA training

            Current blocker

            Verify checkpoint and evaluation compatibility.

            Access and governance
            75confidence 85
            Artifact availability
            55confidence 55
            Training and loading reproducibility
            confidence 15
            Loop data needed
            • video
            • actions
            • language
            Dataset signals
            VideoActionsLanguageSuccess FailureForce Tactile
            official claim· verified 2026-07-23· code metadata verified· weights unknown
            Evaluate in Bench
            VLAOpen2024-06-10model hub

            openvla-7b

            openvla

            Detail

            Official model catalog record.

            Recommended for

            VLA training

            Current blocker

            Verify checkpoint and evaluation compatibility.

            Access and governance
            75confidence 85
            Artifact availability
            90confidence 90
            Training and loading reproducibility
            70confidence 65
            Loop data needed
            • video
            • actions
            • language
            Dataset signals
            VideoActionsLanguageFeature ExtractionImage Text TO Text
            Related catalog data
            No direct catalog match yet.
            Link
            official claim· verified 2026-07-22· code unknown· weights metadata verified
            Evaluate in Bench
            VLAOpen2024model hub

            pi0 / OpenPI

            Physical Intelligence

            Detail

            Generalist robot policy family and open-source robotics model package from Physical Intelligence.

            Recommended for

            Generalist policy transfer

            Current blocker

            Useful for framing OpenBot as a data-readiness layer around emerging generalist policies.

            Code, weights, and checkpoints
            65confidence 40
            Loading and training reproducibility
            confidence 15
            Training data requirements
            76confidence 50
            Loop data needed
            • Broad robot demonstrations with language goals
            • Action/state traces suitable for fine-tuning
            • Task and embodiment metadata for transfer analysis
            Dataset signals
            ObservationLanguage intentActionsRobot stateEmbodiment
            Model hubModel hub
            curated source· verification pending· code declared· weights declared
            Evaluate in Bench
            PolicyOpen2024model hub

            RDT-1B

            Robotics Diffusion Transformer research

            Detail

            Diffusion foundation model for bimanual manipulation that uses large-scale robot data to generate action trajectories.

            Recommended for

            Bimanual manipulation success

            Current blocker

            Useful for showing why action/state tags need to distinguish single-arm, bimanual, and dexterous data.

            Code, weights, and checkpoints
            65confidence 40
            Loading and training reproducibility
            confidence 15
            Training data requirements
            76confidence 50
            Loop data needed
            • Bimanual demonstrations with synchronized visual observations and action/state traces
            • Language or task conditioning for manipulation goals
            • Contact-rich failure and recovery cases for robust long-horizon behavior
            Dataset signals
            ObservationLanguage intentActionsRobot stateTrajectory
            Model hub
            curated source· verification pending· code unknown· weights declared
            Evaluate in Bench
            PolicyOpen2023model hub

            ACT

            ALOHA / LeRobot ecosystem

            Detail

            Action Chunking with Transformers predicts short action sequences for efficient imitation learning in manipulation tasks.

            Recommended for

            Precision manipulation success

            Current blocker

            A practical baseline for dataset readiness because it needs clean action supervision.

            Code, weights, and checkpoints
            65confidence 40
            Loading and training reproducibility
            confidence 15
            Training data requirements
            76confidence 50
            Loop data needed
            • High-quality demonstrations with image and state observations
            • Temporally aligned action chunks
            • Task-specific splits that expose compounding-error failures
            Dataset signals
            ObservationActionsRobot stateHand poseTask success
            DocsModel hub
            curated source· verification pending· code unknown· weights declared
            Evaluate in Bench
            PolicyResearch2023model hub

            Diffusion Policy

            Columbia / Toyota Research Institute / collaborators

            Detail

            Visuomotor policy approach that represents robot behavior as a conditional denoising diffusion process.

            Recommended for

            Multi-modal action generation

            Current blocker

            Important policy baseline for comparing whether a dataset needs a larger VLA/WAM or a stronger task policy.

            Code, weights, and checkpoints
            65confidence 40
            Loading and training reproducibility
            confidence 15
            Training data requirements
            76confidence 50
            Loop data needed
            • Expert demonstrations with action sequences
            • Visual and low-dimensional state observations
            • Failure-aware evaluation splits for multi-modal action distributions
            Dataset signals
            ObservationActionsRobot stateTrajectoryTask success
            Model hub
            curated source· verification pending· code unknown· weights declared
            Evaluate in Bench
            PolicyResearch2023

            RoboCat

            Google DeepMind

            Detail

            Self-improving generalist robotic agent that collects new demonstrations to improve its own manipulation capabilities.

            Recommended for

            Self-improvement loop quality

            Current blocker

            A useful reference for OpenBot's collect-evaluate-improve loop, even without public weights.

            Code, weights, and checkpoints
            confidence 10
            Loading and training reproducibility
            confidence 15
            Training data requirements
            76confidence 50
            Loop data needed
            • Robot demonstrations connected to self-generated data collection
            • Task success and failure feedback to decide what to collect next
            • Cross-task and cross-robot traces that preserve adaptation history
            Dataset signals
            ObservationActionsRobot stateTask successFeedback/failure
            curated source· verification pending· code unknown· weights unknown
            Evaluate in Bench
            VLAClosed ref2023

            RT-2

            Google DeepMind

            Detail

            Vision-language-action model that transfers web-scale vision-language knowledge into robotic control.

            Recommended for

            Semantic generalization

            Current blocker

            Useful reference point for VLA direction even if the model itself is not open.

            Code, weights, and checkpoints
            confidence 10
            Loading and training reproducibility
            confidence 15
            Training data requirements
            76confidence 50
            Loop data needed
            • Web-scale semantic knowledge plus robot action data
            • Language-conditioned tasks with visual observations
            • Held-out semantic generalization and physical execution tests
            Dataset signals
            ObservationLanguage intentActionsObject labelsTask success
            curated source· verification pending· code unknown· weights unknown
            Evaluate in Bench
            VLAResearch2023

            RT-X

            Open X-Embodiment collaboration

            Detail

            Cross-embodiment RT model family trained from the Open X-Embodiment mixture to study transfer across robots and tasks.

            Recommended for

            Cross-embodiment transfer

            Current blocker

            Important for OpenBot because dataset lineage and embodiment metadata become training variables.

            Code, weights, and checkpoints
            confidence 10
            Loading and training reproducibility
            confidence 15
            Training data requirements
            76confidence 50
            Loop data needed
            • Multi-robot trajectories with normalized observations and actions
            • Embodiment metadata that preserves robot, camera, and action-space differences
            • Cross-dataset evaluation to measure whether scaling mixtures improves transfer
            Dataset signals
            ObservationLanguage intentActionsRobot stateEmbodiment
            curated source· verification pending· code unknown· weights unknown
            Evaluate in Bench
            World ModelResearch2023

            UniSim

            UC Berkeley / research

            Detail

            Interactive real-world simulator research that models how visual scenes change under actions and interaction.

            Recommended for

            Interactive scene prediction

            Current blocker

            Useful reference for OpenBot Synth because generated replay should stay anchored to real interaction traces.

            Code, weights, and checkpoints
            confidence 10
            Loading and training reproducibility
            confidence 15
            Training data requirements
            76confidence 50
            Loop data needed
            • Video traces with action or interaction conditioning
            • Temporal consistency signals and scene geometry
            • Evaluation setups that connect predicted rollouts to downstream control or planning
            Dataset signals
            ObservationActionsFuture stateCamera poseSim-real
            curated source· verification pending· code unknown· weights unknown
            Evaluate in Bench
            PolicyResearch2022

            RT-1

            Google Robotics

            Detail

            Robotics Transformer policy trained on large-scale real-world robot demonstrations for language-conditioned manipulation.

            Recommended for

            Real-world task success

            Current blocker

            A useful baseline for understanding why policy learning needs action/state alignment, not only video.

            Code, weights, and checkpoints
            confidence 10
            Loading and training reproducibility
            confidence 15
            Training data requirements
            76confidence 50
            Loop data needed
            • Large-scale robot episodes with images, language commands, and actions
            • Task diversity across objects, scenes, and long-horizon instructions
            • Robust train/test splits that expose distribution shift and recovery limits
            Dataset signals
            ObservationLanguage intentActionsRobot stateTask success
            curated source· verification pending· code unknown· weights unknown
            Evaluate in Bench
            PolicyResearch2022

            SayCan

            Google Research / Everyday Robots

            Detail

            Language-grounded robotics approach that combines language-model planning with affordance scores from robot skills.

            Recommended for

            Instruction decomposition

            Current blocker

            Important for tagging datasets with task phase and feedback, not just final action traces.

            Code, weights, and checkpoints
            confidence 10
            Loading and training reproducibility
            confidence 15
            Training data requirements
            76confidence 50
            Loop data needed
            • Task instructions decomposed into executable skill phases
            • Affordance or feasibility signals for each skill in the environment
            • Failure labels showing where language plans diverge from physical capability
            Dataset signals
            ObservationLanguage intentTask phaseFeedback/failureActions
            curated source· verification pending· code unknown· weights unknown
            Evaluate in Bench
            OpenBot connection

            The model map turns papers into product requirements.

            Catalog scores loop signals. Data preserves them. Bench tests failures. Planned Synth work will later replay measured gaps.

            Next useful capability

            Link each dataset to policy learning, world models, WAM, and failure-mining value.

            Open dataset catalog