00:00:00/Guide
How to transcribe and search CCTV audio
Svid can transcribe the audio track of recorded footage into searchable text, so a search can match what was said on camera exactly the same way it matches what a camera saw. Type “the customer who complained about the queue” and it turns up alongside “person in a red jacket near the till”, from the same search bar.
Plenty of cameras have always recorded a second channel of evidence alongside the picture: whatever the microphone picked up. Almost nobody uses it, because review is built around scrubbing video at speed, and you can't scrub audio the same way. Skim a clip at eight times speed and the picture still tells you roughly what happened; the sound just turns to noise. So the audio sits there, unwatched in the same way the video used to be, except worse, because there was never even a fast way to check it.
How audio transcription works
Transcription is opt-in, not automatic. When you add recorded footage, upload, a pasted URL, or an RTSP capture, there's a “Transcribe audio” checkbox alongside it. It isn't available for live phone sessions; it's a recorded-footage feature. If a transcription fails for some reason, there's a per-video retry action rather than having to re-add the file. And if you're clearing a backlog rather than turning it on one video at a time, a folder has a bulk “Transcribe audio” action that runs every untranscribed video in that folder in one go. It's billed per audio-minute from the account's credits, separately from the frame analysis that already happens on the same footage.
The audio itself isn't kept around afterwards. It's extracted from the video, sent for transcription, and deleted; only the resulting text and its embeddings stick around. If a video has no audio track at all, or a transcription attempt fails, that video simply has no transcript. Nothing else about it changes: the frame descriptions and search results you'd already have are unaffected either way.
Searching audio and video together
There's no separate audio search. Transcript text is embedded and merged straight into the same Library search that already covers frame descriptions, so one search bar covers both. If someone at the till says “this is the third time I've been in this queue”, that sentence is searchable text the same way a red jacket near the till is a searchable description. A single query like “customer complaining about the wait” can surface a moment your cameras never would have caught by sight alone, because nothing about a complaint necessarily looks different from someone just standing there.
- 01
Turn transcription on
Tick “Transcribe audio” when adding a video, or use the per-video retry action if an earlier attempt failed. For a backlog, open a folder and run the bulk “Transcribe audio” action to catch every untranscribed video in it at once.
- 02
Wait for it to finish
Transcription runs alongside the frame analysis Svid already does. A video with no audio track, or one where transcription fails, just ends up with no transcript; everything else about it is unaffected.
- 03
Search normally
Use the same Library search bar you'd use for anything else. Transcript lines and frame descriptions are embedded into the same search, so a spoken complaint and a visual description can both match one query.
- 04
Jump to the moment
A transcript match takes you to the same kind of timestamped moment a frame-description match does, so you land on the clip instead of a wall of text.
What this is and isn't good for
For disputes and complaints, having the actual words on record is genuinely useful: a customer's exact complaint, an argument at a counter, a verbal warning that was or wasn't given. That's a different, often more direct, kind of evidence than a description of what someone looked like doing it. But it's worth being precise about what this is. It's a transcript of what was said, not a system that identifies who said it. There's no voice biometrics or speaker identification here, in keeping with the rest of Svid: no facial recognition, no tracking, nothing that fingerprints a person by how they sound any more than by how they look.
- Does every video get transcribed automatically?
- No. Transcription is opt-in per video, a checkbox when you add recorded footage, plus a folder-level bulk action for transcribing everything untranscribed in that folder at once.
- Can I transcribe audio from a live phone monitor?
- No, transcription applies to recorded footage: uploads, URL ingests and RTSP captures. Live phone sessions aren't covered.
- How is audio transcription priced?
- It's billed per audio-minute from the account's credits, separate from the credits used for frame analysis on the same video.
- What happens if a video has no audio or transcription fails?
- It just has no transcript. Nothing else about the video changes, and a per-video retry action is there if you want to try again.
- Does Svid identify who's speaking?
- No. It transcribes what was said into searchable text. There's no voice identification or biometric matching of speakers, consistent with Svid having no facial recognition either.
Try it on footage you've already got
Create an account at app.svid.ai, tick “Transcribe audio” on a video with sound, and search it in plain English this afternoon.
Related footage
Guide
How to search through hours of CCTV footage
Stop scrubbing. The fastest way to search hours of CCTV footage is to have AI describe every frame, then type what you're looking for in plain English.
Guide
How to find CCTV evidence for an insurance claim
Find CCTV evidence for a claim by searching footage in plain English, building the timeline around the incident, and exporting clips before retention overwrites them.
Glossary
Natural language video search
Natural language video search finds moments in footage from a typed description, like “person in a red jacket near the till”, with no scrubbing or timestamps.