Technical notes from the Tensorial team.
2026-07-14
How we built a voice-and-vision model that answers in ~1.2s, watches your camera, speaks proactively — and how to talk to it right now, from this page.