Bottom Line Up Front (BLUF): This tool transcribes audio and video entirely in your browser: the file is not sent to our server, no tokens are charged, and the recognized text can be edited.
The Whisper model (about 40 MB) is downloaded once and cached by the browser. The language can be detected automatically or picked manually. The result is shown with timestamps and can be edited, copied and downloaded as TXT, SRT, VTT or JSON.
If you upload a video, you can also render it with burned-in subtitles (in Chrome and Edge) and download the finished file.
Recognition runs on your device: neither the file nor the resulting text reaches our server. There is no paid API and no tokens are charged.
Quality depends on the recording: the model is accurate on clear speech and may make mistakes on noisy audio. The recognized text can be corrected before saving the subtitles.
Yes, no tokens are charged: the model runs locally in the browser.
No, neither the audio nor the video nor the recognized text leaves your device.
Yes, the text is editable right on the page and the edits go into the subtitles.
TXT, SRT, VTT, JSON, and for video a finished clip with burned-in subtitles.