Touching sound: the sense that reaches where light and touch can't.
A camera resolves shape. A fingertip resolves texture. But neither can tell a solid billet from a hollow shell, water from air behind a closed wall, a seated bearing from a cracked one — and both go dark the moment they're occluded or the lights go out. Contact audio does. When a hand taps, scratches, or shakes an object, the vibrations it radiates carry the things vision and touch structurally cannot reach: material, internal structure, fill level, hidden defects. It works through occlusion, in the dark, and on the inside. The frontier here is fingertip microphones fusing with vision and touch into one scene estimate — and the deep reason to fuse is not redundancy but complementarity: each sense is the sole witness to a different property. Sound is the only witness to what's sealed inside.
One object, three noisy senses. Each is an expert at a single attribute and near-flat on the rest — so alone, each column is confident about one row and lost on the others. Multiply the likelihoods (exact Bayes, conditional-independence given the object) and the Fused column is sharp on all four at once. Run 200 objects and fused accuracy clears every single sense by a wide margin. Then switch Sound off and watch the Fill row collapse to a coin-flip — no camera and no fingertip ever sensed what's inside a closed shell. The likelihoods are set from each sense's real strengths, not learned; the logic of why fusion wins is exactly what the learned systems below recover from data. Everything runs on your device.
Why it matters: the property no other sense can reach.
Fusion is usually sold as noise-averaging — many weak looks at the same thing. The sharper argument is coverage: some properties have exactly one witness. Whether a pill bottle is full, whether a weld is sound, whether a valve is seating — a camera and a fingertip are at chance, and only the tap knows. A robot that can't hear can't perceive those states at all, at any resolution. Sound isn't a better look at the visible; it's the only look at the sealed.
In the field · contact audio is a fast-moving 2025–26 frontier. Duke's SonicSense gives a multi-fingered hand contact microphones so it identifies materials and objects by tapping, scratching, and shaking. Stanford's ManiWAV puts a mic in the gripper and learns contact-rich manipulation from the audio cameras can't see — sensing contact events and surface properties directly from sound. The broader line fuses audio with vision and touch into unified scene understanding, the multimodal-complementarity thesis the sim above makes concrete. Our contribution is the transparent counterpart: an on-device, legible fusion where you can read exactly why the estimate sharpens — and see the one property that has a single witness. It is the acoustic sibling of the contact layer's micron touch, sensing where the eye and the finger both stop.
Now hear it yourself — with your own microphone.
The fusion demo above abstracts the sound channel into a likelihood. Here is the real thing, on hardware you already have: tap, scratch, or shake an object near your microphone and teach this two contact sounds by example. A tiny model runs entirely on your device — no upload, nothing stored — turns each tap into a timbre fingerprint, and learns to tell a bright metal ring from a dull wooden thud. Then it listens and names what you touch. This is the audible twin of teaching a camera, and the browser-native face of SonicSense: identity from contact sound alone.
A nearest-prototype classifier over a 26-dimensional fingerprint — 24 log-frequency band energies, plus brightness (spectral centroid) and ring-out time. Node-verified to separate two synthetic contact timbres at 100% from five examples each; on real objects it is honest few-shot, so distinct sounds tapped at a steady distance work best. Legible on purpose: you can see each class's average fingerprint, and exactly why one tap is called metal and another wood.
And hear the hidden property directly.
The fusion demo argues sound is the sole witness to what's sealed inside. Here you can listen to it. A camera looking at an opaque, sealed bottle sees exactly one thing whether it's empty or full — but tap it, and the pitch reads the level instantly. Drag the fill and hear it: tap the side and the added liquid mass drops the pitch; blow across the top and the shrinking air cavity raises it. Same bottle, two mechanisms, opposite directions — and both turn a hidden fill level into something you can hear.
Helmholtz resonance f = (c/2π)·√(A ÷ V·L′) with the air volume shrinking as it fills — exact physics, pitch rising. The struck-wall model f ∝ 1/√(1 + α·fill) is a mass-loading approximation, pitch falling. Both Node-verified monotonic and in the audible band; tones synthesized live. The guess the fill game makes the point stick — your eyes are useless on a sealed bottle, your ears are not.
The topic, in one line.
Three live instruments, one thesis: some properties of the physical world have exactly one sensory witness, and for what's sealed, hollow, or hidden, that witness is sound. Fuse it and perception sharpens; teach it and a tap becomes an identity; tune it and a pitch becomes a fill gauge. The frontier of contact audio is a robot learning to hear what it cannot see.
↓ Whitepaper · PDFRead online◆ Living paperTechnical Report TR-2026-22 · Institute for Physical AI @ BMI