Institute for Physical AI @ Bailey Military Institute · Charlot Lab
Living paper · position

Position — the ontology of a physical agent

An Entity, Not a Controller

Enactivism and active inference both say a mind is its coupling to the world. Energy First Architecture is what that claim looks like as computation you can run.

Charlot Lab, Institute for Physical AI @ BMI

Draft · 2026-07-22 · a bridge argument, not a new empirical result — its claims are conceptual and its one demonstration is cited, not re-derived here

Position paper · no new experiment Bridges cited work; ENACT used as motivating, not proof
The prevailing picture of a robot is a controller: a function that maps observations and instructions to actions, with a world model bolted on for planning or evaluation. Two mature research traditions reject that picture at the root. Enactivism holds that cognition is not representation-then-action but sensorimotor coupling — a mind is constituted by how an agent brings forth a world through interaction. Active inference makes the same move formal: an agent is a thing that minimizes a single quantity, free energy, by both perceiving and acting. Neither tradition has a shipping computational substrate. We argue that Energy First Architecture is one: a single scalar energy that the agent perceives by minimizing, acts by descending, certifies by its contraction, and is priced by in watts. The entity is not a controller that owns a model — it is constituted by its energy. This is a bridge, not a proof; we say plainly where each tradition is contested and where our own demonstration stops.

1 · Two ontologies of a physical agent

The control-first agent is a pipeline: perceive, then decide, then act, with the world model a service consumed at the decision step. It is optimized to succeed on tasks it was shown, and it is frozen — or, increasingly, fine-tuned online — but always around a success signal. The entity-embracing agent is not a pipeline. Its perceiving and its acting are the same activity seen from two sides: the continuous reduction of a discrepancy between the world it expects and the world it is in. The difference is not stylistic. A controller can be arbitrarily good at seen tasks and still have no internal quantity whose value tells it, before it moves, whether it is about to leave the regime it understands. An entity constituted by such a quantity cannot help but have one.

2 · Enactivism: the mind is the coupling

The enactivist tradition — Varela, Thompson, Di Paolo, and the sensorimotor-contingency account of O'Regan and Noë — holds that perception is something an agent does through its body, not a movie played on an inner screen.1 Cognition is sense-making through sensorimotor coupling; the agent and its niche co-determine each other. The engineering consequence is sharp: a system trained on selected projections of the world — text, curated images, offline logs — is, in the allegory the field keeps reaching for, reading shadows. It never acquires the contingencies that only closed-loop interaction reveals.

There is now a benchmark that operationalizes exactly this. ENACT casts the enactivist claim as a test and finds that a frontier model's gap to human performance widens as the interaction horizon lengthens2 — the regime where a genuinely coupled entity should win and static prediction cannot fake it. We are careful with this evidence: ENACT tests disembodied vision-language models, so it shows that the coupling regime is where static prediction fails; it does not prove that an energy-grounded system succeeds there, and certainly not that ours does. It is motivating, not a verdict.

3 · Active inference: the entity that minimizes one quantity

Where enactivism is a philosophy, the Free Energy Principle is a formalism.3 An agent is modeled as a system that resists dispersal by minimizing variational free energy $F$ — a bound on surprise — over both its internal states (perception) and its actions (which change what it senses):

$$ a^\star,\,\mu^\star \;=\; \arg\min_{a,\,\mu}\; F(\mu, a),\qquad F \;=\; \underbrace{D_{\mathrm{KL}}\!\big[q(\mu)\,\Vert\,p(\mu)\big]}_{\text{complexity}} \;-\; \underbrace{\mathbb{E}_q[\log p(o\mid\mu)]}_{\text{accuracy}}. $$

The single most important structural fact is that one scalar governs both perceiving and acting. The agent does not have a perception module and a separate control module reconciled by a supervisor; it has an energy, and inference and behavior are two descent directions on it. Active inference has real robot demonstrations — mobile manipulation, long-horizon rearrangement4 — though we note that the strongest recent results are in simulation, and the widely-quoted "runs on 3% of the compute" figure is marketing, not a claim in the paper.

One scalar, two descent directions
belief settles at  ·  complexity  ·  accuracy term  ·  F

The same $F$, read twice. Sliding the belief downhill is perception: it settles at the precision-weighted compromise between what was expected and what was sensed, which is why a noisy sensor barely moves it and a sharp one drags it almost all the way. The second arrow is action: $\partial F/\partial o$ points toward changing the world until the sensor reports what was predicted. Neither is a module, and there is no supervisor deciding between them — they are two directions on one surface, which is the structural claim this section is making. Widen the prior and watch the agent become credulous; sharpen it and watch the same evidence stop mattering.

The most striking recent instance is a developmental one. A free-energy robot that learns to ground language — verb–adjective–object commands — entirely through curiosity-driven self-exploration reaches its accuracy criterion in roughly half the training epochs of a non-curious control, and, tellingly, reproduces a signature of a developing agent rather than a tuned one: the U-shaped error curve of a human toddler who overgeneralizes ("goed") before mastering the exception.6 Its curiosity is exactly the complexity term above — the agent is rewarded, by $-D_{\mathrm{KL}}[q\Vert p]$, for seeking observations that force its beliefs to update. We report this in scope: it is a simulated body, measured against a matched recurrent-RL baseline, not against the control-first policies we contrast with here. But a controller is programmed and static; an entity develops — and a developmental error curve is hard to account for on the controller ontology.

A controller has a model. An entity is its energy — the same scalar it perceives by, acts by, and stakes its stability on.

4 · Energy First Architecture as the computational form

EFA takes the two traditions' shared commitment — one scalar quantity constitutes the agent — and makes it a buildable stack. The correspondence is direct:

the tradition saysEFA computes
the mind is sensorimotor coupling (enactivism)a single learned scalar energy $E(s)$ over the agent's own state — not a model of the world held at arm's length, but the coupling itself
perceive and act by minimizing one quantity (active inference)predict by descending $E$; act by descending $E$ to a goal — the same object, two descent directions
an agent resists leaving the states it can maintainthe contraction of $E$ is a formal certificate: it tells the agent, before it commits, whether an imagined action keeps it in its basin — the entity's own boundary, made checkable
a living thing is a thermodynamic entitythe whole loop is priced in joules per task; the metric is watts, not benchmark points
the coupling is continuous, never finishedthe energy keeps updating from live interaction

The third row is where EFA adds something the traditions gesture at but do not build. Active inference's free energy is minimized; it is not usually turned into a worst-case certificate that gates commitment. In a companion demonstration5 we show, at nano scale, that one scalar energy can be both the action-objective and a sampled contraction certificate that rejects plans an outcome-scorer commits to and that then diverge — and that on ternary weights the certificate's core operation is multiply-free, so the entity can afford to verify itself at the edge. That is the entity-embracing agent in miniature: it does not consult a model and decide; it descends an energy, and the same energy tells it when to stop.

5 · How an entity explores: by belief-update, on its own trajectory

If perceiving and acting are two descents on one energy, a third question follows: what does the entity choose to do next? The controller ontology has no principled answer — exploration is a heuristic bolted on. The entity ontology does: seek the experiences that most update your beliefs, along the trajectory you actually live. We measured this at nano, having an agent learn a forward model of a small 2-D world containing a "noisy-TV" patch — a region whose transitions are mostly irreducible noise — and scoring the learned model by 15-step rollout error, the property a world model exists for.7

Three policies separated cleanly, over six seeds, and the ordering is more instructive than a clean winner. Seeking surprise — the prediction-error signal, high wherever outcomes are unpredictable — is the worst explorer: it is captured by the noisy TV and spends its budget on irreducible noise, the classic failure of naïve curiosity. Seeking pure epistemic uncertainty — maximum belief-variance, the information-gain term — escapes the noisy TV but is, we report plainly, worse than random for rollout: it chases uncertain boundaries while rollout error accumulates on the attractor, where the agent actually spends its time. Only when information gain is weighted by the agent's own visitation distribution — belief-update, evaluated where the trajectory goes — does it become the best explorer. The moral is the entity's, not the controller's: explore by what would change your beliefs, but weighted by where you already are. An agent that merely seeks surprise chases the noisy TV; an agent that is a coupling to its world explores the world it is coupled to.

6 · Limits

This is a bridge argument, not a discovery. Enactivism and the Free Energy Principle are both actively contested — the FEP in particular is criticized as unfalsifiable in its strongest forms, and we do not lean on it as settled science, only as a formal statement of the one-quantity ontology. ENACT is motivating evidence about where static prediction fails, not proof that EFA succeeds. Our own supporting result is nano-scale, in simulation, with a sampled (not sum-of-squares) certificate, and unreplicated; it demonstrates a mechanism, not a scaled system, and we have not built a Friston-style agent. What we claim is narrow and, we think, correct: the entity-vs-controller distinction is mechanical, not merely philosophical, and Energy First Architecture is a concrete way to build the entity side. The exploration result of §5 is likewise nano and in simulation; its two large-margin, robust claims are that prediction-error curiosity is worst and that pure maximum-variance information gain is worse than random for rollout, while the visitation-weighted advantage over random is small — we report the ordering, not a decisive margin.

References

  1. Varela, Thompson & Rosch, The Embodied Mind; O'Regan & Noë, sensorimotor contingencies; Di Paolo et al., enactive cognition.
  2. ENACT: an enactivist interaction-horizon benchmark, arXiv:2511.20937, ICLR 2026.
  3. Friston, The free-energy principle: a unified brain theory?, and the active-inference literature.
  4. VERSES AI, Mobile Manipulation with Active Inference for Long-Horizon Rearrangement, arXiv:2507.17338, NeurIPS 2025 (simulation, Habitat).
  5. Charlot Lab, Energy Is the Certificate (companion living paper, this issue) — the nano demonstration and the ternary bound-propagation result.
  6. Tinker, Doya & Tani, Curiosity-Driven Development of Action and Language in Robots Through Self-Exploration, Science Advances (2026), DOI 10.1126/sciadv.aee7533; preprint arXiv:2510.05013 (simulation; curiosity = the KL/complexity term of variational free energy).
  7. Charlot Lab, Curiosity for a world model — companion measurement (research/oist-curiosity): prediction-error vs pure vs visitation-weighted information gain on a 2-D forward model with a noisy-TV patch. Builds on the noisy-TV analyses of Pathak et al. (ICM, 2017) and Burda et al. (RND / "Large-Scale Study of Curiosity", 2019).

Companion: Energy Is the Certificate · drive the certificate-gated imagination · the EFA hub.