A vision-language-action policy turns an instruction and a scene into an action. Pick an instruction — watch it perceive the table, ground the words to an object, and execute an action chunk.
Vision
perceive
Language
ground
Action
execute
Instruction
Grounding here is transparent attribute-matching (colour · type · relation), shown as weights — a real VLA learns it from data. Action is an executed chunk of Δx·Δy·grip steps.