Skip to the report

Vision · Language · Action

One loop: see it, ground it, do it

A vision-language-action policy turns an instruction and a scene into an action. Pick an instruction, watch it perceive the table, ground the words to an object, and execute an action chunk.

Vision
perceive
Language
ground
Action
execute
Instruction
Grounding here is transparent attribute-matching (colour · type · relation), shown as weights, a real VLA learns it from data. Action is an executed chunk of Δx·Δy·grip steps.