00:00:00/Glossary

Natural language video search

Natural language video search is finding moments in video by typing what you want in ordinary words, such as “person in a red jacket near the till”, instead of scrubbing timelines or filtering by camera and timestamp. You describe the moment, and the software takes you to it.


How it works

A vision model has already looked at frames from the footage and written a description of each one: the people in shot, what they're wearing, what they're doing. Those descriptions are indexed as text, so your query is a text search over descriptions rather than a viewing job, which is why an answer comes back in seconds however many hours of footage sit behind it. Svid works this way, and treats live and recorded footage identically.

What it replaces

The old ways of finding a moment were all watching in disguise: scrubbing at eight times speed and hoping your attention holds, or filtering a VMS down to a camera and a time window and then still watching everything inside it. Natural language search removes the watching step. Type “forklift reversing without a spotter” or “someone slipped in the entrance” and jump straight to the frame.

Search vs searchable

Searchable video is the property of footage that has been described and indexed; natural language video search is the act and the interface, the box you type into. One exists so the other can happen.

Related footage