Free · macOS & Windows · No account, no upload

Private transcripts, speaker by speaker

Turn recorded interviews, meetings and fieldwork into time-stamped, speaker-labelled transcripts, without your audio ever leaving your computer.

Every word timed and scored 99 languages Your audio never leaves your computer
The app showing a transcribed interview split into turns labelled S_01, S_02 and S_03, with the speaker tuning controls on the left

For researchers, and anyone else with hours of recordings

Hours of recordings, and a transcript has to start somewhere.

Interviews, focus groups, meetings, lectures, fieldwork. Before you can code, analyse or quote them, someone has to write down who said what, and when.

Transcribing by hand

Slow, and most of it is typing

A first pass takes several hours for every hour of audio. Across a whole project that adds up to weeks spent typing before any analysis begins.

Cloud transcription services

Your recordings leave your hands

They mean uploading to someone else's server, which may not fit what participants consented to or what your ethics approval and data-protection rules allow. Most also charge by the minute.

DOTE Whisper

A first draft, made on your own machine

Speech recognition and speaker separation run locally, so the recordings stay where they are. You get a time-stamped draft split into speaker turns, ready to check, correct and build on.

How it works

From a folder of recordings to transcripts

Add your recordings

Drop audio or video files onto the window. Queue one file or a whole project's worth.

Pick a model and language

Choose between speed and accuracy, and tell the app the language, or let it detect it. Turn on speaker labels.

Transcribe

The app works out who speaks when, then transcribes each speaker turn with OpenAI's Whisper model.

Tune and export

Adjust how turns and speakers are split, then save for DOTE, as subtitles, or as open JSON for your own tools.

What you get

Built for research recordings

Nothing leaves your computer

No upload, no account, no API key, no per-minute fees. Once a model has downloaded, transcription works without an internet connection.

Who said what

Speaker diarization splits the recording into turns and labels each voice (S_01, S_02, …). Tell the app how many people took part, or let it decide.

Tune speaker turns live

Too many speakers, or turns chopped into fragments? Adjust a few controls and the transcript updates at once, without transcribing again.

Every word timed and scored

Each word has its own start and end time and a confidence score. Colour coding shows which passages are worth listening to again.

99 languages, plus specialist models

Whisper models from fast to most accurate. Add community models trained for a language or field from HuggingFace, such as Swedish, Danish, Cantonese or medical speech.

Batches, and open exports

Queue many recordings and each transcript is saved beside its file. Export JSON for DOTE or your own scripts, or SRT and WebVTT subtitles.

Opens MP3, WAV, FLAC, M4A, OGG, AAC and WMA audio, and the sound from MP4, MOV, MKV and AVI video. Speech recognition by whisper.cpp, speaker diarization by sherpa-onnx. Uses the GPU on Apple silicon Macs and on Windows PCs with an NVIDIA graphics card.

Privacy

Made for recordings you can't upload

Research recordings are often personal, and the people in them agreed to specific uses. DOTE Whisper keeps the whole process on the computer you already use for the data.

  • Processed locallyDecoding, speech recognition and speaker separation all run on your computer. Working copies of the audio are temporary and deleted afterwards.
  • Nothing to sign up forNo account, no licence key, no analytics. Nothing about you or your files is sent to us.
  • Transcripts go where you put themBeside the recording, or wherever you choose to save. Each JSON file also records the settings that produced it.

What goes online, and when

What When
Your recordings and transcripts Never
Model downloads, from HuggingFace and GitHub Once per model, when you first use it
A check for the latest version number When the app starts

Full details on the privacy page.

Know what you're getting

A strong first draft, not a finished transcript

Automatic transcription is a starting point for your own listening, correcting and annotating. The app is built to show you where to look.

  • Close to verbatimWhisper writes down what it hears, but it tends to tidy up: false starts, fillers and overlapping talk are often shortened or left out.
  • Speaker labels are a guideSeparating voices automatically is still hard. Short backchannels, similar voices and noisy rooms cause mistakes, and the tuning controls help you correct most of them.
  • Confidence on every wordLow-scoring words are marked in orange and red, so you know which passages to check first.
Three transcript segments with each word tinted by confidence; a short 'mm-hm' is marked in orange at 66%
Words shaded by confidence. The short "mm-hm" scores lowest.

Start on your next transcript today

Free to download. No account, no subscription, no upload.