Research topic

The contact layer: perception where line-of-sight ends.

Every remote sense goes to zero at the moment of contact. The last millimeter, the forces holding a grasp, whether a surface is slipping, the micron texture that tells a bolt from a screw — none of it is knowable from across the room. Touch is the perception layer that begins exactly where OmniSense ends. A vision-based tactile fingertip makes the trade concrete: an elastomer gel deforms against an object, a camera under three colored lights sees only color, and geometry is computed — not measured — on the device. Photometric stereo inverts the colored image into a field of surface normals; those normals integrate into a depth map; a grid of printed markers flowing with the gel reads shear and the onset of slip. An image sensor becomes a geometry-and-force sensor. This is the fingertip complement to the whole-body magnetic skin of the printed-body corpus: micron detail where dexterity needs it, an energy-accounted event skin everywhere else.

The range of perceptionremote · across the roomcontact · the surfaceacoustic · inside the object

Press a shape into the gel and watch the pipeline: the raw RGB sensor image, the normals recovered by photometric stereo (n = M-1·[R,G,B]), the depth integrated from those normals (∇²z = ∇·g), and the marker field, where a gold stuck-core shrinks from the edge inward as the grip crosses the friction cone into slip. Drag the raw pane to apply lateral force — and watch the haptic-out trace, the vibration an actuator would replay as the finger slides across texture. From the reconstructed geometry alone it names the surface — a hex fastener, a knurled grip, raised dots — and with Auto-grip on, the fingertip applies the minimum force to hold, tightening the instant it senses incipient slip. Everything runs on the device.

Why it matters: a hand that doesn't drop things.

Slip is not an abstraction — it is the difference between holding an object and dropping it. Each contact can resist a tangential load only up to μ times its normal force; exceed that friction cone and the finger slides. Load the object below and watch a grasp fail, then let the reflex close it with the minimum pinch force.

Two fingertips, one object. Grasp holds while the friction capacity μ·2Fn covers the load; push the load past it and a contact's cone is exceeded — it slips red and the object falls. Auto-grip is the same slip-recovery reflex, now closing a grasp: it raises pinch force the instant the margin thins and relaxes it when the load eases.

In the field · vision-based tactile is the mainline of dexterous touch — GelSight's retrographic sensing and Meta's DIGIT 360 (with GelSight Inc.), an open hemispherical fingertip with ~8.3M taxels and multi-modal pressure, vibration, and shear. Two directions define the 2025–26 frontier. Touch is becoming a foundation-model problem — transformers trained by masked tactile prediction that transfer zero-shot across sensors (T3, AnyTouch). And event-based tactile now flags incipient slip hundreds of milliseconds to ~2 s before it turns gross — the same stick-to-slip transition this fingertip watches — with neuromorphic spiking networks near 94% on the no / incipient / gross split. Our recognition here is the transparent, on-device counterpart to those learned representations: features read straight off the reconstruction, no black box. And our line is the pairing of this micron fingertip with the printed-body event skin — full touch as an energy-accounted system, sensing and haptic return alike.

The learned version, live.

The recognition above reads hand-picked features off the reconstruction. The field has moved to learned representations — so here is that, at a scale you can watch: a small network trained by real backprop, on raw depth patches with no hand-coded rules. The embedding clusters by texture as it learns; then it names a new press. It is the legible, on-device cousin of a tactile foundation model — same idea, 2,789 parameters instead of a billion.

A 100→24→12→5 network with real gradient descent. Watch the loss fall and the 12-D embedding separate into texture clusters (PCA-projected), then press Test to recognize a held-out touch. Reset the weights to watch it learn from scratch. Supervised on labeled example presses — honest few-shot, not the self-supervised, cross-sensor transfer of T3/AnyTouch — but the same principle: features are learned, not authored.

↓ Whitepaper · PDFRead online◆ Living paperTechnical Report TR-2026-09 · Institute for Physical AI @ BMI