← Research

Research · Area 01

Machines that see, and model what they see.

An embodied system is only as good as its picture of the world. Work in this area builds the perception stack (depth, detection, segmentation, and multimodal fusion) and the predictive world models that let an autonomous machine act on what it perceives, in degraded real-world conditions, and on the device itself.

● LIVE OCCUPANCY FIELD · drivable space dim · detected objects in gold
IPAI·TV

Watch the explainer

The idea behind this lab in ninety seconds — then come back and drive it.

▶ Watch on YouTube
The work

How the work happens.

The methods behind the research.

Multimodal sensing

Fuse what one sensor can't see

LiDAR, thermal, and RGB combined for perception that holds up in smoke, glare, and the dark, where any single modality fails.

On-device inference

Perception at the edge

Compressed, scheduled models running in real time inside the airframe's power and thermal budget, with no cloud round-trip in the control loop.

World models

Pixels to prediction

Self-supervised models that predict how a scene evolves, not only classify it — the substrate planning and control build on.

Robustness

Outside the lab

Perception measured where it matters: motion blur, occlusion, novel objects, and the long tail that breaks benchmark-tuned models.

Current directions

Open problems we're pursuing.

What's being pursued now.

D1

Monocular depth on the edge

How far can one camera and a small model go for distance estimation, and where must active sensing take over?

D2

Open-vocabulary detection

Recognizing objects the system was never trained on — a prerequisite for autonomy in unstructured environments.

D3

Cross-modal calibration

Keeping LiDAR, thermal, and RGB aligned in space and time on a moving platform, without hand-tuning.

Live lab

See what the machine sees, on your device.

The perception stack, live: monocular depth, semantic segmentation, and object detection running on-device via WebGPU, on a sample scene or your own webcam. Toggle between what the camera sees and what the model perceives. No cloud.

DEPTH · SEGMENTATION · DETECTIONReal neural inference on your device, from a sample image or live webcam
Get involved

Work in this area.

Open positions, including the Physical AI Investigator Program, are listed on Careers. Research here can also spin out into a company.

Open positions →Spin it out →