PAI-180 · Education

Vision-Language-Action: Build a Policy You Can Talk To

Build a vision-language-action policy from the ground up, because to judge the incumbent path of embodied AI, you must first build it. You make vision that turns pixels into places, language that grounds open-vocabulary instructions, and an action head that joins them; you prove it obeys words rather than memorizing a target; and you meet the costs a VLA pays a body (it is data-hungry, it fails out of distribution, and it carries no guarantee before it acts) plus the two problems we hit shipping our own: compositional order, and why motion needs more than one frame. Build the incumbent, know exactly what it costs, then see the physics-first inversion in the Energy First Architecture course, where one energy is the controller and its own proof.

Skilled·3 modules·7 lessons
▶ Start the course ← All courses
THE HORIZON

Where this sits, and what moves it.

Binding constraint · The gap between a prior and a bound. A language model supplies a probability over what to do; a body needs a limit on force, reach and velocity, and no amount of fluency measures one. That gap is the quantity, and it is measurable the moment you write the envelope down.

Was impossible

Talking to a robot meant a fixed command vocabulary someone had enumerated in advance. Anything outside the list was not a hard request -- it was not a request at all.

Is probable

A vision-language-action stack takes instructions it was never given a slot for, and this course builds one you can talk to. What it does not give you is a guarantee: the language model supplies a prior over what to do, and a prior is not a bound. Treating fluency as reliability is the specific mistake this material is built to prevent.

Becomes possible

The step that matters is composing the prior with a checkable envelope, so the language layer proposes and a physics layer disposes -- which is exactly the seam PAI-250's certificate and PAI-210's runtime envelope are built on. A student who takes those three together is standing on the open problem, not reading about it.

Every hard thing was impossible until the constraint that made it impossible was named. How we read a frontier →

See it live

Anatomy demonstrations

The machines behind this course, taken apart three ways, the body, the one rule, and the small learned brain. Guess before you look; an open core proves every number on the page.

DemonstrationThe learned brainsOn-device policies you can watch decide →The whole series →