The viral clip's one defensible word was “black box.” But black box means hard to interpret, not mysterious in operation, we wrote every multiply. And interpreting it is a real, working science. Here it is, on your device: a small network that has learned twelve concepts, its neurons pulled apart into features you can read, and a dial to steer what it thinks. Trained live in your browser.
Training a small network and a sparse autoencoder in your browser…
A neuron means several things at once
The network has just six hidden neurons but twelve concepts to represent, so each neuron has to do double duty. Click a neuron and see which concepts light it up, not one, but several unrelated ones. That is polysemanticity, and it is why staring at raw neurons feels like a black box.
the twelve concepts it learned
the six hidden neurons · click one
neuron 1 · how strongly each concept activates it
A sparse autoencoder pulls them apart
Train an overcomplete sparse autoencoder on those six neurons' activity and it recovers the hidden concepts as separate, mostly monosemantic features: each firing for one thing. This is the move that turned interpretability from a hope into a field. Each tile shows what its feature detects; teal-bordered ones are cleanly monosemantic.
Reach in and change its mind
Understanding is causal, not just descriptive. Pick a monosemantic feature and clamp it on, the network's answer moves to that feature's concept, no matter what you actually showed it. You are not watching the black box; you are steering it.