Vision · Language · Action

One loop: see it, ground it, do it

A vision-language-action policy turns an instruction and a scene into an action. Pick an instruction — watch it perceive the table, ground the words to an object, and execute an action chunk.

Vision
perceive
Language
ground
Action
execute
Instruction
Grounding here is transparent attribute-matching (colour · type · relation), shown as weights — a real VLA learns it from data. Action is an executed chunk of Δx·Δy·grip steps.