The basics in 60 seconds
You recorded an interview, a meeting or a session in the field. The app listens to it, works out who spoke when, and writes down what was said, with a time for every word. You then check and correct that draft, in DOTE or wherever you work.
- Add recordings by dropping them on the window, or with Browse….
- Check the settings: the model, the language, and Enable Speaker Diarization if you want speaker labels.
- Transcribe with Start Transcription.
- Tune the speaker turns in the Tuning tab until they match what you hear.
- Export with the JSON, SRT or VTT buttons.
A few words you'll see throughout the app and this guide:
- Model
- The speech-recognition engine, a version of OpenAI's Whisper. Bigger models are more accurate and slower.
- Diarization
- Working out who speaks when. The app groups stretches of speech by how the voices sound. It doesn't know anyone's name.
- Speaker label
- S_01, S_02 and so on, numbered in the order people first speak.
- Turn / segment
- One continuous stretch of speech by one speaker, with a start and end time. The app shows each as a card.
- Confidence
- How sure the model was about a word, from 0 to 100%. Low scores mark places to listen again.
- Pipeline settings
- Settings used while transcribing. Changing them means transcribing again.
- Tuning
- Adjustments made after transcribing. The transcript updates instantly.
1 · Install & first launch
Download the installer for your computer from the download page and open it.
| Computer | What to do |
|---|---|
| macOS |
Open the .dmg and drag
DOTE Whisper
onto the
Applications folder
shown beside it. Then open it from
Applications or Launchpad. The first
time, macOS asks you to confirm
opening an app downloaded from the
internet.
|
| Windows |
Run the .exe. If
Windows says "Windows protected your
PC", click
More info, then
Run anyway. The app
installs, opens, and appears in the
Start menu.
|
There is nothing to set up. The app contains everything it needs to read audio and video. The only extra downloads are the models, which the app fetches the first time you use each one (see Choose a model).
2 · A tour of the window
The window has two panels. The left one is where you set things up; the right one shows progress and results.
| Area | What it's for |
|---|---|
| Configuration tab, left | Your list of recordings, the model, the language, speaker labels, and the button that starts transcription. |
| Tuning tab, left | Controls that reshape speaker turns in a finished transcript. The app switches to it when a transcription finishes. |
| Console tab, right | A running log of what the app is doing. Useful when something goes wrong. |
| Output tab, right | The transcript: the full text, then each turn with its times, speaker label and confidence. Export buttons sit at the top. |
| top right | Settings: where models are stored, how confidence is shown, and the speaker-detection settings. |
File → Check for updates… tells you whether a newer version is out. The app also checks quietly each time it starts.
3 · Add your recordings
Drag files from Finder or File Explorer onto the window, or click Browse… and pick one or more.
The app reads MP3, WAV, FLAC, M4A, OGG, AAC and WMA audio, and takes the sound from MP4, MOV, MKV and AVI video. Files of other types are skipped, and the Console says which.
Each file gets a coloured dot: grey waiting, blue running, green done, amber cancelled, red failed (the error shows under the name). Buttons on each row:
| Button | What it does |
|---|---|
| Puts a finished, failed or cancelled file back in the queue, to run again with the current settings. | |
| Shows the saved transcript in Finder or File Explorer. | |
| Saves the transcript. Only there when auto-save is off and the file isn't saved yet. | |
| Removes the file from the list. The recording itself is untouched. |
Click a finished file's name to show its
transcript. Add files adds more
at any time, even while a batch is running, and
Clear empties the list.
Auto-save next to source, on by
default, saves each transcript beside its
recording as
<name>.transcript.json the
moment it finishes.
4 · Choose a model
The model decides how accurate the transcript is and how long it takes. The default, Large V3, is the best general-purpose choice.
| Model | Size | Use it for |
|---|---|---|
| Base / Base (English Only) | ~140 MB | Quick tests and rough drafts. Fast, but misses and mishears more. |
| Medium / Medium (English Only) | ~1.4 GB | A middle ground for slower computers. |
| Large V2 | ~2.9 GB | The previous large model. Worth trying if Large V3 struggles with a recording. |
| Large V3 (default) | ~2.9 GB | Most recordings, especially noisy ones and anything not in English. |
| KBLab KB-Whisper-Large | ~2.9 GB | Swedish recordings. |
The English Only versions are slightly faster and more accurate on English, but can't transcribe anything else. If a recording mixes languages, use a multilingual model.
Picking a model you haven't downloaded yet shows its size and asks first. Large downloads take a while; you can cancel, and nothing half-finished is left behind. To free space, hover over a downloaded model in the menu and click . The file goes to the Trash or Recycle Bin, so you can still restore it.
Custom models
Researchers and companies publish versions of Whisper trained further on one language or one kind of speech: Danish, Cantonese, medical dictation, and many more. They live on HuggingFace, and the app can install most of them.
- Open the model menu and choose Add Custom Model from HuggingFace.
-
Paste the model's name, the
owner/namepart of its web address. Forhuggingface.co/KBLab/kb-whisper-largethat isKBLab/kb-whisper-large. A full link works too. - Click Probe & install. The app checks the model before downloading anything large.
What happens next depends on the model:
-
Ready-made versions (names
often end in
-ggmlor-gguf) download and install straight away. - Standard Whisper models are converted on your computer. The app shows the download size and free disk space, then Convert & install takes a few minutes. Leave the Quantization variant at Q5_1 (recommended): it roughly halves the file size for a very small loss in accuracy.
- Other kinds of model can't be used. The app says why, and nothing large is downloaded.
Browse on HuggingFace opens a
search already filtered to Whisper models. Type
a two-letter language code first
(da, fr,
ja) to narrow it. Installed custom
models appear under
Installed Custom Models in the
menu. Check each model's page for its license
and for what it was trained on.
5 · Set the language
Leave Language empty and the app detects it. Setting it yourself is more reliable.
- Set it for noisy or short recordings. Detection listens to the start of the audio and can guess wrong when that part is unclear.
- With speaker labels on, the language is detected per turn. That helps when participants speak different languages, but a short "yes" or "mm" can be mistaken for another language. If the whole recording is in one language, set it.
- Custom codes. If a custom model expects a code that isn't listed, type it and choose Use custom code.
6 · Speaker labels
Tick Enable Speaker Diarization to split the transcript into turns and label each voice.
Expected speakers is optional. If you know how many people speak, enter it: the app keeps the voices with the most speech and folds any extra, smaller ones into them. Leave it blank to let the app decide. You can change the number after transcribing in the Tuning tab, so it's fine to leave blank and look first.
The first time you turn speaker labels on, the app downloads about 35 MB of speaker models. Recordings shorter than about two seconds are transcribed without labels.
7 · Transcribe
Click Start Transcription. With several files waiting it reads Transcribe 3 files, and the app works through them one after another.
With speaker labels on, the app first finds the speaker turns in the whole recording, then transcribes each turn separately. The Console lists every turn as it goes. On the right of the Progress panel:
| Badge | Meaning |
|---|---|
| GPU / CPU | Whether the graphics chip is doing the work. Apple silicon Macs and Windows PCs with an NVIDIA card use the GPU, which is several times faster. |
| RAM | How much memory the app is using, with a breakdown underneath. It turns amber above 70% of your computer's memory and red above 85%. Closing other programs helps. |
Cancel stops the current file; with more files waiting it reads Stop & cancel remaining and stops the batch. You can keep using the rest of the app while it runs, and files you add mid-batch join the end of the queue. When it finishes, the app switches to the transcript and the Tuning tab.
8 · Read the transcript
The Output tab shows the language and number of segments, the Full Text, and then every segment as a card.
Each card shows the start and end time, the speaker label, and the segment's overall confidence. When several files are done, the Showing menu at the top switches between them.
The colour runs from green (confident) through yellow to red (unsure). A segment's percentage summarises its words, and one very unsure word pulls it down. Low scores tend to mark mumbling, crosstalk, names, technical terms and backchannels. In the example, the model heard a quiet "mm-hm" as the letters M M H M and scored it 66%.
Prefer a different look? Settings → Word-confidence visualization offers a tint behind each word, a thin coloured strip above each segment, a tooltip only, or no colouring.
9 · Tune speaker turns
Automatic speaker labels are rarely perfect. The Tuning tab lets you reshape turns and speakers until they match what you hear, without transcribing again.
How turns are made
Knowing the steps makes it clear which control fixes which problem. While transcribing, the app:
- Finds speech and marks the points where the voice changes. Stretches shorter than the minimum turn duration (0.3 s by default) are left out at this stage.
- Groups the stretches by voice, using a speaker model that turns a few seconds of each voice into a "voiceprint". Stretches whose voiceprints are closer than the cluster threshold share a label.
- Re-links labels. It compares every label's voiceprint and merges those closer than the re-link threshold. This joins up a voice that was split, for example because someone moved away from the microphone, and links speakers across long recordings, which are processed in overlapping 40-minute pieces.
- Joins a speaker's consecutive turns when the pause between them is shorter than the same-speaker merge gap (1 s).
- Transcribes each turn, so every word inherits its turn's speaker.
The app saves each speaker's voiceprint with the transcript. That is what lets the Tuning tab regroup speakers instantly. It can merge speakers and turns, but it can never split them, because splitting needs information that only a new run can produce.
The four controls
Changes show in the Output tab about a quarter of a second after you stop typing or dragging. They're applied in this order:
- Speaker re-link threshold
- How alike two voiceprints must be to share a label. Higher merges more, so fewer speakers; lower keeps more apart. The help text gives the useful range for the speaker model that was used (0.3 to 0.4 for the default, CAM++). Above it, different people start to merge.
- Expected speakers
- Keeps this many speakers: the ones who talk most stay, and each other label is folded into whichever of them sounds most alike. Blank or 0 means no target.
- Same-speaker merge gap
- Joins consecutive turns by the same speaker when the silence between them is no longer than this (up to 5 s).
- Minimum turn duration
- Turns shorter than this (up to 2 s) are absorbed into the nearer neighbouring turn. No words are lost, but they take on the neighbour's speaker label.
The merge gap and minimum turn duration start at the values used while transcribing and can only go up from there. To go lower, change them in Settings → Advanced and run the file again.
Auto-tune resets the re-link threshold to the transcript's own value and applies the number in Expected speakers. If that box is blank, it fills in a suggestion: the number of speakers who have at least 2% of the speech. Reset to pipeline defaults puts all four back to how the transcript came out.
A worked example
A recorded interview: one interviewer, S_01, and two participants. The app found four speakers. The fourth, S_03, is a single "mm-hm" from the interviewer, too short for its voiceprint to match anyone.
Clicking Auto-tune suggested 3 speakers and folded the stray label into the closest-sounding voice, the interviewer's. Labels are renumbered afterwards, so the second participant is now S_03.
The second participant's answer was still split into two cards with a pause of 1.3 seconds between them. Raising the same-speaker merge gap from 1 to 1.5 joined them:
Finally, the minimum turn duration, to show what it does and why to be careful. Raised to 1 second, the two backchannels (the interviewer's "mm-hm" and a participant's "yeah") are absorbed into the neighbouring turns:
The words are still there, but they now belong to the wrong people. That's fine for a readable content summary, and wrong if you study how listeners respond. Choose based on what your analysis needs.
Problem → fix
| What you see | What to do |
|---|---|
| More speakers than took part, and the extras say very little | Enter the real number in Expected speakers, or click Auto-tune. Don't raise the re-link threshold to get there: it merges real people before it clears small fragments. |
| One person's speech is spread across two labels, both with plenty of speech | Set Expected speakers. If that picks the wrong pair, raise the re-link threshold a step at a time, staying within the range the help text gives. |
| Two different people share one label | Tuning can't split them. First try lowering the re-link threshold, in case the app joined them in its re-link step. If they stay together, run the file again with a lower cluster threshold. |
| One speaker's talk is chopped into many short cards | Raise the same-speaker merge gap. 1.5 to 2 seconds suits slow or hesitant speech. |
| Backchannels ("mm-hm", "yeah") clutter the transcript as their own turns | Raise the minimum turn duration to about 0.5 to 1 s. Remember that absorbed words take on the neighbouring speaker's label. |
| Backchannels and short answers matter to your analysis | Keep the minimum turn duration low, and set Expected speakers so short turns are folded into the right voice rather than the nearest turn. Check them by ear. |
| A question and its answer land in one card | The change of speaker was missed, which tuning can't undo. This happens most with quick exchanges and similar voices. Correct it by hand in your transcription tool, or try a run with a lower cluster threshold or a different speaker model. |
| A few words from a TV, a passer-by or the next room | Leave Expected speakers blank. Those voices then keep their own label, which makes off-topic speech easy to spot and delete, rather than hiding it inside a participant's turn. |
When to run the file again
Some fixes need a new run, because they change what the app finds in the first place. Change the setting in Settings → Advanced, then click beside the file and Start Transcription.
| Goal | Change |
|---|---|
| Find more speakers (people merged together) | Lower the per-chunk cluster threshold by 0.1 to 0.2. You'll usually get a few extra small labels too, which Expected speakers then tidies up. |
| Keep very short utterances | Lower the minimum turn duration towards 0.1 s. Speech shorter than this setting is left out of the transcript altogether. |
| Keep every pause as a turn boundary | Lower the same-speaker merge gap, down to 0. |
| Many speakers in English, or persistent mix-ups | Try a different speaker embedding model. TitaNet-L did best on English recordings with many speakers. |
10 · Save & export
The three buttons above the transcript save it as it looks on screen, tuning included.
| Format | Use it for |
|---|---|
| JSON |
The complete transcript: every word
with its times, confidence and
speaker. Import it into
DOTE, or read it with your own scripts.
Saved as
<name>.transcript.json.
|
| SRT |
Subtitles for video players and
editors, one cue per segment with a
[S_01]: prefix.
|
| VTT |
WebVTT subtitles for web video, with
speakers as voice tags (<v S_01>).
|
A trimmed example of one segment in the JSON:
{
"generator": "dote-whisper",
"language": "en",
"text": "Thanks for coming in today. …",
"segments": [
{
"id": 3,
"start": 21.58,
"end": 27.318,
"text": "But the flip side is that I barely see my team. …",
"speaker": "S_02",
"confidence": 0.966,
"words": [
{ "id": 62, "start": 21.68, "end": 21.68, "text": "But",
"confidence": 0.817, "speaker": "S_02" },
…
]
}
],
"post_processing_settings": {
"tuning": { "same_speaker_merge_sec": 1.5, "min_turn_duration_sec": 0.3,
"relink_threshold": 0.35, "expected_speakers": 3 },
"pipeline_floors": { … }
}
}
Times are in seconds from the start of the
recording. post_processing_settings
records the tuning you applied and the values
the transcript came out with, so you, a
colleague or a reviewer can see exactly how it
was produced. It also keeps each speaker's
voiceprint, under
diarizationMetadata.
The app labels speakers S_01, S_02 and so on. To use names or pseudonyms, replace the labels in your transcription tool, or with find-and-replace in the exported file.
11 · Settings
Open Settings with in the top-right corner. Changes save automatically and apply to the next run.
| Setting | What it does |
|---|---|
| Processing pipeline | Windows with an NVIDIA card only. Auto uses the GPU when possible, and switches to the CPU by itself if the GPU fails. Macs choose automatically. |
| Whisper Models Directory | Where models are stored. Change points the app at another folder, such as an external drive. Existing models aren't moved; move them yourself, or they download again. |
| Word-confidence visualization | How confidence is shown: underline (default), tint, a per-segment strip, tooltip only, or off. |
| Speaker embedding model | The model that makes voiceprints. Affects speaker labels only, never the words. See below. |
| Reset to defaults | Puts every setting back. Downloaded models are kept. |
Choosing a speaker embedding model. In the developer's tests on meetings, broadcasts and conversations in nine languages, recording quality mattered more than language:
| Model | Good for |
|---|---|
| CAM++ (default) | Any language. Fastest, lightest, and best choice for Mandarin and phone recordings. |
| TitaNet-L | English-trained with many speakers. About twice as slow. |
| TitaNet-S | Faster than TitaNet-L. |
| ERes2Net | May work better than CAM++ on some languages, and in particularly noisy environments or poor accoustics. Slower and more resource-intensive than CAM++. |
Each model downloads on first use (27 to 100 MB) and keeps its own threshold settings.
Advanced settings
These are the pipeline settings described in How turns are made. They take effect on the next run. The defaults suit most recordings.
| Setting | Default | Effect |
|---|---|---|
| Diarization chunk threshold | 40 min | Recordings longer than this are processed in overlapping pieces. Longer pieces label speakers more consistently but use much more memory. Computers with little memory get shorter pieces automatically. |
| Per-chunk cluster threshold | 0.7 (CAM++) | Lower finds more speakers; higher finds fewer. The main setting for people who were merged together. |
| Minimum turn duration | 0.3 s | Speech shorter than this is ignored. Raising it removes breaths and clicks, and also loses short answers like "yes". |
| Same-speaker merge gap | 1.0 s | Joins a speaker's consecutive turns across pauses up to this long. 0 turns it off. |
| Speaker re-link threshold | 0.35 (CAM++) | Higher merges more voices after clustering. The recommended range for the chosen model is shown, and turns amber outside it. |
12 · Better results
- Start with the best recording. If you recorded on several devices, use the clearest one. A microphone close to the speakers beats any setting. Where a camera and a separate audio recorder captured the same event, transcribe the audio recorder's file.
- Set the language when you know it.
- Use Large V3 unless speed matters more than accuracy. For a language Whisper handles less well, look for a custom model.
- Tell it how many people speak. An Expected speakers number removes most stray labels.
- Listen where it's red. Low confidence clusters are where corrections pay off.
- Try settings on a short excerpt before running a whole project. Speaker settings that suit one room or group usually suit the rest.
13 · Troubleshooting
The same phrase repeats again and again
A known Whisper habit after long silences or music. Dote Whisper no longer carries words from one 30-second window into the next, which was the main cause, so it should be rare. If it still happens, turn on speaker labels: silences are then mostly skipped.
The text is in the wrong language, or nonsense
Set the Language instead of auto-detect. If you chose an English Only model for a recording in another language, the file stops with an error that explains this: pick a multilingual model.
The Tuning tab says "unavailable for this transcript"
The transcript was made without speaker labels, or the voiceprints couldn't be saved. Tick Enable Speaker Diarization, click and transcribe again. The Console shows any error.
It's slow
Check the Progress panel: CPU means no GPU is being used, which is expected on Intel Macs and on PCs without an NVIDIA card. Try Medium or Base, and leave long batches to run overnight.
The memory badge turns red, or a run fails with a memory error
Close other programs. Long recordings with speaker labels need the most memory. The app retries pieces that run out in smaller parts, and uses shorter pieces on computers with less memory. If it still fails, lower the diarization chunk threshold in Advanced settings, or use a smaller Whisper model.
A custom model won't install
"Not found" usually means a typo, or a private model: only public models can be installed. "Incompatible" means it isn't a Whisper speech-to-text model, so try another. A network error is worth retrying. Some converted models give slightly less precise word times (about 20 ms instead of 10); the words themselves are unaffected.
Windows says "Windows protected your PC"
The installer isn't code-signed yet. Click More info, then Run anyway.
The NVIDIA GPU isn't used on Windows
Update the NVIDIA driver. If the GPU fails mid-run, the app finishes on the CPU instead of stopping. Remote-desktop and virtual-machine sessions usually can't use the GPU.
Where are my files?
| What | Where |
|---|---|
| Transcripts | Beside each recording (auto-save), or where you saved them. Use in the file list. |
| Models | The folder in Settings → Whisper Models Directory; its folder button opens it. |
| Settings (macOS) |
~/Library/Application
Support/DOTE Whisper/
|
| Settings (Windows) |
%APPDATA%\DOTE Whisper\
|
How do I update?
When a new version is out, the app tells you as it starts, or when you choose File → Check for updates…. Download it from the download page and install it over the old one. Your settings and models are kept.
How do I cite it?
If you use
DOTE Whisper in
published work, please cite it along with the
projects it builds on: OpenAI Whisper (Radford
et al., 2022,
Robust Speech Recognition via Large-Scale
Weak Supervision), whisper.cpp (Georgi Gerganov) and
sherpa-onnx (the k2-fsa project). Reporting the
model, the speaker model and the
post_processing_settings from your
JSON files makes your method reproducible.