Machines that see, and model what they see.
An embodied system is only as good as its picture of the world. Work in this area builds the perception stack (depth, detection, segmentation, and multimodal fusion) and the predictive world models that let an autonomous machine act on what it perceives, in degraded real-world conditions, and on the device itself.
Watch the explainer
The idea behind this lab in ninety seconds — then come back and drive it.
How the work happens.
The methods behind the research.
Fuse what one sensor can't see
LiDAR, thermal, and RGB combined for perception that holds up in smoke, glare, and the dark, where any single modality fails.
Perception at the edge
Compressed, scheduled models running in real time inside the airframe's power and thermal budget, with no cloud round-trip in the control loop.
Pixels to prediction
Self-supervised models that predict how a scene evolves, not only classify it — the substrate planning and control build on.
Outside the lab
Perception measured where it matters: motion blur, occlusion, novel objects, and the long tail that breaks benchmark-tuned models.
Open problems we're pursuing.
What's being pursued now.
Monocular depth on the edge
How far can one camera and a small model go for distance estimation, and where must active sensing take over?
Open-vocabulary detection
Recognizing objects the system was never trained on — a prerequisite for autonomy in unstructured environments.
Cross-modal calibration
Keeping LiDAR, thermal, and RGB aligned in space and time on a moving platform, without hand-tuning.
See what the machine sees, on your device.
The perception stack, live: monocular depth, semantic segmentation, and object detection running on-device via WebGPU, on a sample scene or your own webcam. Toggle between what the camera sees and what the model perceives. No cloud.
Work in this area.
Open positions, including the Physical AI Investigator Program, are listed on Careers. Research here can also spin out into a company.