Tavlo AI Lab / the narrator

The Narrator

A camera that talks: a 256-million-parameter vision-language model watches your webcam and tells you what it sees.

live · on-device · no upload
the model sees
Ask about the scene
Private by design. The narrator is the open-source SmolVLM-256M vision-language model running on your GPU via Transformers.js and WebGPU (one-time ~1 GB download, then cached by your browser). Frames are analysed in memory on your device and are never uploaded, recorded, or stored; camera access is requested only when you press “Enable camera”. See our privacy policy for details.