00:00:00/Glossary
Natural language video search
Natural language video search is finding moments in video by typing what you want in ordinary words, such as “person in a red jacket near the till”, instead of scrubbing timelines or filtering by camera and timestamp. You describe the moment, and the software takes you to it.
How it works
A vision model has already looked at frames from the footage and written a description of each one: the people in shot, what they're wearing, what they're doing. Those descriptions are indexed as text, so your query is a text search over descriptions rather than a viewing job, which is why an answer comes back in seconds however many hours of footage sit behind it. Svid works this way, and treats live and recorded footage identically.
What it replaces
The old ways of finding a moment were all watching in disguise: scrubbing at eight times speed and hoping your attention holds, or filtering a VMS down to a camera and a time window and then still watching everything inside it. Natural language search removes the watching step. Type “forklift reversing without a spotter” or “someone slipped in the entrance” and jump straight to the frame.
Search vs searchable
Searchable video is the property of footage that has been described and indexed; natural language video search is the act and the interface, the box you type into. One exists so the other can happen.
Related footage
Glossary
Searchable video
Searchable video is footage an AI has described frame by frame, so you can query it in plain English. It's also what Svid stands for: Searchable Video.
Glossary
Video intelligence
Video intelligence is the use of AI to extract meaning from video, the people, actions and events in footage, rather than just recording and storing it.
Guide
How to search through hours of CCTV footage
Stop scrubbing. The fastest way to search hours of CCTV footage is to have AI describe every frame, then type what you're looking for in plain English.