Guide

How to use DOTE Whisper

From installing the app to a speaker-labelled transcript you can correct and analyse. No technical background needed. Read it in order the first time, or jump to what you need.

The basics in 60 seconds

You recorded an interview, a meeting or a session in the field. The app listens to it, works out who spoke when, and writes down what was said, with a time for every word. You then check and correct that draft, in DOTE or wherever you work.

  1. Add recordings by dropping them on the window, or with Browse….
  2. Check the settings: the model, the language, and Enable Speaker Diarization if you want speaker labels.
  3. Transcribe with Start Transcription.
  4. Tune the speaker turns in the Tuning tab until they match what you hear.
  5. Export with the JSON, SRT or VTT buttons.

A few words you'll see throughout the app and this guide:

Model
The speech-recognition engine, a version of OpenAI's Whisper. Bigger models are more accurate and slower.
Diarization
Working out who speaks when. The app groups stretches of speech by how the voices sound. It doesn't know anyone's name.
Speaker label
S_01, S_02 and so on, numbered in the order people first speak.
Turn / segment
One continuous stretch of speech by one speaker, with a start and end time. The app shows each as a card.
Confidence
How sure the model was about a word, from 0 to 100%. Low scores mark places to listen again.
Pipeline settings
Settings used while transcribing. Changing them means transcribing again.
Tuning
Adjustments made after transcribing. The transcript updates instantly.

1 · Install & first launch

Download the installer for your computer from the download page and open it.

Computer What to do
macOS Open the .dmg and drag DOTE Whisper onto the Applications folder shown beside it. Then open it from Applications or Launchpad. The first time, macOS asks you to confirm opening an app downloaded from the internet.
Windows Run the .exe. If Windows says "Windows protected your PC", click More info, then Run anyway. The app installs, opens, and appears in the Start menu.

There is nothing to set up. The app contains everything it needs to read audio and video. The only extra downloads are the models, which the app fetches the first time you use each one (see Choose a model).

2 · A tour of the window

The window has two panels. The left one is where you set things up; the right one shows progress and results.

The main window: Tuning tab on the left, the Output tab on the right showing speaker-labelled segments
A finished interview. Tuning on the left, the transcript on the right.
Area What it's for
Configuration tab, left Your list of recordings, the model, the language, speaker labels, and the button that starts transcription.
Tuning tab, left Controls that reshape speaker turns in a finished transcript. The app switches to it when a transcription finishes.
Console tab, right A running log of what the app is doing. Useful when something goes wrong.
Output tab, right The transcript: the full text, then each turn with its times, speaker label and confidence. Export buttons sit at the top.
top right Settings: where models are stored, how confidence is shown, and the speaker-detection settings.

File → Check for updates… tells you whether a newer version is out. The app also checks quietly each time it starts.

3 · Add your recordings

Drag files from Finder or File Explorer onto the window, or click Browse… and pick one or more.

The app reads MP3, WAV, FLAC, M4A, OGG, AAC and WMA audio, and takes the sound from MP4, MOV, MKV and AVI video. Files of other types are skipped, and the Console says which.

The file list: one finished file with a Saved note, one running at 25%, one waiting
The list during a batch: done, running, waiting.

Each file gets a coloured dot: grey waiting, blue running, green done, amber cancelled, red failed (the error shows under the name). Buttons on each row:

Button What it does
Puts a finished, failed or cancelled file back in the queue, to run again with the current settings.
Shows the saved transcript in Finder or File Explorer.
Saves the transcript. Only there when auto-save is off and the file isn't saved yet.
Removes the file from the list. The recording itself is untouched.

Click a finished file's name to show its transcript. Add files adds more at any time, even while a batch is running, and Clear empties the list. Auto-save next to source, on by default, saves each transcript beside its recording as <name>.transcript.json the moment it finishes.

4 · Choose a model

The model decides how accurate the transcript is and how long it takes. The default, Large V3, is the best general-purpose choice.

The model menu: built-in models with download or tick icons, installed custom models, and Add Custom Model from HuggingFace
A tick means downloaded; the arrow means it downloads when you pick it.
Model Size Use it for
Base / Base (English Only) ~140 MB Quick tests and rough drafts. Fast, but misses and mishears more.
Medium / Medium (English Only) ~1.4 GB A middle ground for slower computers.
Large V2 ~2.9 GB The previous large model. Worth trying if Large V3 struggles with a recording.
Large V3 (default) ~2.9 GB Most recordings, especially noisy ones and anything not in English.
KBLab KB-Whisper-Large ~2.9 GB Swedish recordings.

The English Only versions are slightly faster and more accurate on English, but can't transcribe anything else. If a recording mixes languages, use a multilingual model.

Picking a model you haven't downloaded yet shows its size and asks first. Large downloads take a while; you can cancel, and nothing half-finished is left behind. To free space, hover over a downloaded model in the menu and click . The file goes to the Trash or Recycle Bin, so you can still restore it.

Download model dialog for Medium, 1.4 GB, with a Large download warning and Download and Cancel buttons
The confirmation before a model downloads.

Custom models

Researchers and companies publish versions of Whisper trained further on one language or one kind of speech: Danish, Cantonese, medical dictation, and many more. They live on HuggingFace, and the app can install most of them.

  1. Open the model menu and choose Add Custom Model from HuggingFace.
  2. Paste the model's name, the owner/name part of its web address. For huggingface.co/KBLab/kb-whisper-large that is KBLab/kb-whisper-large. A full link works too.
  3. Click Probe & install. The app checks the model before downloading anything large.
Add Custom Model dialog with a field for the HuggingFace repo, a Browse on HuggingFace button and advice on which repos work
Paste a name, or browse with a language filter.
Convert from raw weights panel with free disk space and a Quantization variant menu set to Q5_1 (recommended)
A model that needs converting on your computer.

What happens next depends on the model:

  • Ready-made versions (names often end in -ggml or -gguf) download and install straight away.
  • Standard Whisper models are converted on your computer. The app shows the download size and free disk space, then Convert & install takes a few minutes. Leave the Quantization variant at Q5_1 (recommended): it roughly halves the file size for a very small loss in accuracy.
  • Other kinds of model can't be used. The app says why, and nothing large is downloaded.

Browse on HuggingFace opens a search already filtered to Whisper models. Type a two-letter language code first (da, fr, ja) to narrow it. Installed custom models appear under Installed Custom Models in the menu. Check each model's page for its license and for what it was trained on.

5 · Set the language

Leave Language empty and the app detects it. Setting it yourself is more reliable.

The language picker filtered by typing sw, showing sv Swedish and sw Swahili
Type part of a name or code to filter the 99 languages.
  • Set it for noisy or short recordings. Detection listens to the start of the audio and can guess wrong when that part is unclear.
  • With speaker labels on, the language is detected per turn. That helps when participants speak different languages, but a short "yes" or "mm" can be mistaken for another language. If the whole recording is in one language, set it.
  • Custom codes. If a custom model expects a code that isn't listed, type it and choose Use custom code.

6 · Speaker labels

Tick Enable Speaker Diarization to split the transcript into turns and label each voice.

The Configuration tab with three finished files, Large V3 selected, language on auto-detect, speaker diarization ticked and Expected speakers blank
The Configuration tab, ready for another run.

Expected speakers is optional. If you know how many people speak, enter it: the app keeps the voices with the most speech and folds any extra, smaller ones into them. Leave it blank to let the app decide. You can change the number after transcribing in the Tuning tab, so it's fine to leave blank and look first.

Count everyone who speaks, not only the people you're studying. An interviewer, a facilitator or someone who walks in and says a sentence each count. If unsure, round up: one too many leaves a small extra label, one too few merges two real people.

The first time you turn speaker labels on, the app downloads about 35 MB of speaker models. Recordings shorter than about two seconds are transcribed without labels.

7 · Transcribe

Click Start Transcription. With several files waiting it reads Transcribe 3 files, and the app works through them one after another.

A transcription in progress: the file at 80%, the memory panel, a GPU badge, and the console listing speaker turns as they are transcribed
Mid-run: 8 of 11 speaker turns transcribed, on the GPU.

With speaker labels on, the app first finds the speaker turns in the whole recording, then transcribes each turn separately. The Console lists every turn as it goes. On the right of the Progress panel:

Badge Meaning
GPU / CPU Whether the graphics chip is doing the work. Apple silicon Macs and Windows PCs with an NVIDIA card use the GPU, which is several times faster.
RAM How much memory the app is using, with a breakdown underneath. It turns amber above 70% of your computer's memory and red above 85%. Closing other programs helps.

Cancel stops the current file; with more files waiting it reads Stop & cancel remaining and stops the batch. You can keep using the rest of the app while it runs, and files you add mid-batch join the end of the queue. When it finishes, the app switches to the transcript and the Tuning tab.

8 · Read the transcript

The Output tab shows the language and number of segments, the Full Text, and then every segment as a card.

Each card shows the start and end time, the speaker label, and the segment's overall confidence. When several files are done, the Showing menu at the top switches between them.

Three segments with each word underlined in green, yellow or orange according to confidence
The default view: each word underlined by confidence. Hover a word to see its score.

The colour runs from green (confident) through yellow to red (unsure). A segment's percentage summarises its words, and one very unsure word pulls it down. Low scores tend to mark mumbling, crosstalk, names, technical terms and backchannels. In the example, the model heard a quiet "mm-hm" as the letters M M H M and scored it 66%.

Prefer a different look? Settings → Word-confidence visualization offers a tint behind each word, a thin coloured strip above each segment, a tooltip only, or no colouring.

How verbatim is it? Whisper aims for readable text. It often drops or tidies fillers ("um", "uh"), repetitions and false starts, smooths over overlapping talk, and adds punctuation. Treat the output as a first pass to correct against the recording, especially if your analysis depends on hesitations, pauses or overlap.

9 · Tune speaker turns

Automatic speaker labels are rarely perfect. The Tuning tab lets you reshape turns and speakers until they match what you hear, without transcribing again.

How turns are made

Knowing the steps makes it clear which control fixes which problem. While transcribing, the app:

  1. Finds speech and marks the points where the voice changes. Stretches shorter than the minimum turn duration (0.3 s by default) are left out at this stage.
  2. Groups the stretches by voice, using a speaker model that turns a few seconds of each voice into a "voiceprint". Stretches whose voiceprints are closer than the cluster threshold share a label.
  3. Re-links labels. It compares every label's voiceprint and merges those closer than the re-link threshold. This joins up a voice that was split, for example because someone moved away from the microphone, and links speakers across long recordings, which are processed in overlapping 40-minute pieces.
  4. Joins a speaker's consecutive turns when the pause between them is shorter than the same-speaker merge gap (1 s).
  5. Transcribes each turn, so every word inherits its turn's speaker.

The app saves each speaker's voiceprint with the transcript. That is what lets the Tuning tab regroup speakers instantly. It can merge speakers and turns, but it can never split them, because splitting needs information that only a new run can produce.

The four controls

The Tuning panel with same-speaker merge gap, minimum turn duration, speaker re-link threshold slider, expected speakers with Auto-tune, tips, and Reset to pipeline defaults
The Tuning tab.

Changes show in the Output tab about a quarter of a second after you stop typing or dragging. They're applied in this order:

Speaker re-link threshold
How alike two voiceprints must be to share a label. Higher merges more, so fewer speakers; lower keeps more apart. The help text gives the useful range for the speaker model that was used (0.3 to 0.4 for the default, CAM++). Above it, different people start to merge.
Expected speakers
Keeps this many speakers: the ones who talk most stay, and each other label is folded into whichever of them sounds most alike. Blank or 0 means no target.
Same-speaker merge gap
Joins consecutive turns by the same speaker when the silence between them is no longer than this (up to 5 s).
Minimum turn duration
Turns shorter than this (up to 2 s) are absorbed into the nearer neighbouring turn. No words are lost, but they take on the neighbour's speaker label.

The merge gap and minimum turn duration start at the values used while transcribing and can only go up from there. To go lower, change them in Settings → Advanced and run the file again.

Auto-tune resets the re-link threshold to the transcript's own value and applies the number in Expected speakers. If that box is blank, it fills in a suggestion: the number of speakers who have at least 2% of the speech. Reset to pipeline defaults puts all four back to how the transcript came out.

Export before you switch. Tuning isn't stored per transcript. Showing a different transcript resets the controls to that transcript's own values, so save each one when it looks right.

A worked example

A recorded interview: one interviewer, S_01, and two participants. The app found four speakers. The fourth, S_03, is a single "mm-hm" from the interviewer, too short for its voiceprint to match anyone.

Segments before tuning: the mm-hm is labelled S_03 and the second participant S_04
Before: a stray S_03 for one "mm-hm".
Segments after Auto-tune: the mm-hm is labelled S_01, the interviewer, and the second participant S_03
After Auto-tune: three speakers, and the "mm-hm" belongs to S_01.

Clicking Auto-tune suggested 3 speakers and folded the stray label into the closest-sounding voice, the interviewer's. Labels are renumbered afterwards, so the second participant is now S_03.

The second participant's answer was still split into two cards with a pause of 1.3 seconds between them. Raising the same-speaker merge gap from 1 to 1.5 joined them:

After raising the merge gap to 1.5 seconds, S_03's two consecutive turns are one segment from 00:28.094 to 00:58.064
Merge gap 1.5 s: one turn from 00:28 to 00:58.

Finally, the minimum turn duration, to show what it does and why to be careful. Raised to 1 second, the two backchannels (the interviewer's "mm-hm" and a participant's "yeah") are absorbed into the neighbouring turns:

With minimum turn duration at 1 second, M M H M is now at the start of an S_02 turn and Yeah. at the start of an S_03 turn
Tidier, but "mm-hm" is now credited to S_02 and "yeah" to S_03.

The words are still there, but they now belong to the wrong people. That's fine for a readable content summary, and wrong if you study how listeners respond. Choose based on what your analysis needs.

Problem → fix

What you see What to do
More speakers than took part, and the extras say very little Enter the real number in Expected speakers, or click Auto-tune. Don't raise the re-link threshold to get there: it merges real people before it clears small fragments.
One person's speech is spread across two labels, both with plenty of speech Set Expected speakers. If that picks the wrong pair, raise the re-link threshold a step at a time, staying within the range the help text gives.
Two different people share one label Tuning can't split them. First try lowering the re-link threshold, in case the app joined them in its re-link step. If they stay together, run the file again with a lower cluster threshold.
One speaker's talk is chopped into many short cards Raise the same-speaker merge gap. 1.5 to 2 seconds suits slow or hesitant speech.
Backchannels ("mm-hm", "yeah") clutter the transcript as their own turns Raise the minimum turn duration to about 0.5 to 1 s. Remember that absorbed words take on the neighbouring speaker's label.
Backchannels and short answers matter to your analysis Keep the minimum turn duration low, and set Expected speakers so short turns are folded into the right voice rather than the nearest turn. Check them by ear.
A question and its answer land in one card The change of speaker was missed, which tuning can't undo. This happens most with quick exchanges and similar voices. Correct it by hand in your transcription tool, or try a run with a lower cluster threshold or a different speaker model.
A few words from a TV, a passer-by or the next room Leave Expected speakers blank. Those voices then keep their own label, which makes off-topic speech easy to spot and delete, rather than hiding it inside a participant's turn.

When to run the file again

Some fixes need a new run, because they change what the app finds in the first place. Change the setting in Settings → Advanced, then click beside the file and Start Transcription.

Goal Change
Find more speakers (people merged together) Lower the per-chunk cluster threshold by 0.1 to 0.2. You'll usually get a few extra small labels too, which Expected speakers then tidies up.
Keep very short utterances Lower the minimum turn duration towards 0.1 s. Speech shorter than this setting is left out of the transcript altogether.
Keep every pause as a turn boundary Lower the same-speaker merge gap, down to 0.
Many speakers in English, or persistent mix-ups Try a different speaker embedding model. TitaNet-L did best on English recordings with many speakers.
These settings apply to every recording you transcribe afterwards. When you've finished with a difficult file, set them back, or use Reset to defaults at the bottom of Settings.

10 · Save & export

The three buttons above the transcript save it as it looks on screen, tuning included.

Format Use it for
JSON The complete transcript: every word with its times, confidence and speaker. Import it into DOTE, or read it with your own scripts. Saved as <name>.transcript.json.
SRT Subtitles for video players and editors, one cue per segment with a [S_01]: prefix.
VTT WebVTT subtitles for web video, with speakers as voice tags (<v S_01>).

A trimmed example of one segment in the JSON:

{
  "generator": "dote-whisper",
  "language": "en",
  "text": "Thanks for coming in today. …",
  "segments": [
    {
      "id": 3,
      "start": 21.58,
      "end": 27.318,
      "text": "But the flip side is that I barely see my team. …",
      "speaker": "S_02",
      "confidence": 0.966,
      "words": [
        { "id": 62, "start": 21.68, "end": 21.68, "text": "But",
          "confidence": 0.817, "speaker": "S_02" },
        …
      ]
    }
  ],
  "post_processing_settings": {
    "tuning": { "same_speaker_merge_sec": 1.5, "min_turn_duration_sec": 0.3,
                "relink_threshold": 0.35, "expected_speakers": 3 },
    "pipeline_floors": { … }
  }
}

Times are in seconds from the start of the recording. post_processing_settings records the tuning you applied and the values the transcript came out with, so you, a colleague or a reviewer can see exactly how it was produced. It also keeps each speaker's voiceprint, under diarizationMetadata.

Auto-save and tuning. Auto-saved files use the tuning values active when each file finished, with Expected speakers taken from the Configuration tab. If you tune a transcript afterwards, click JSON and save over the auto-saved file to keep your changes.

The app labels speakers S_01, S_02 and so on. To use names or pseudonyms, replace the labels in your transcription tool, or with find-and-replace in the exported file.

11 · Settings

Open Settings with in the top-right corner. Changes save automatically and apply to the next run.

Settings dialog: processing pipeline, Whisper models directory, word-confidence visualization, speaker embedding model, Advanced settings and Reset preferences
Settings on a Mac. On a Windows PC with an NVIDIA card there's also a GPU/CPU choice.
Setting What it does
Processing pipeline Windows with an NVIDIA card only. Auto uses the GPU when possible, and switches to the CPU by itself if the GPU fails. Macs choose automatically.
Whisper Models Directory Where models are stored. Change points the app at another folder, such as an external drive. Existing models aren't moved; move them yourself, or they download again.
Word-confidence visualization How confidence is shown: underline (default), tint, a per-segment strip, tooltip only, or off.
Speaker embedding model The model that makes voiceprints. Affects speaker labels only, never the words. See below.
Reset to defaults Puts every setting back. Downloaded models are kept.

Choosing a speaker embedding model. In the developer's tests on meetings, broadcasts and conversations in nine languages, recording quality mattered more than language:

Model Good for
CAM++ (default) Any language. Fastest, lightest, and best choice for Mandarin and phone recordings.
TitaNet-L English-trained with many speakers. About twice as slow.
TitaNet-S Faster than TitaNet-L.
ERes2Net May work better than CAM++ on some languages, and in particularly noisy environments or poor accoustics. Slower and more resource-intensive than CAM++.

Each model downloads on first use (27 to 100 MB) and keeps its own threshold settings.

Advanced settings

These are the pipeline settings described in How turns are made. They take effect on the next run. The defaults suit most recordings.

Setting Default Effect
Diarization chunk threshold 40 min Recordings longer than this are processed in overlapping pieces. Longer pieces label speakers more consistently but use much more memory. Computers with little memory get shorter pieces automatically.
Per-chunk cluster threshold 0.7 (CAM++) Lower finds more speakers; higher finds fewer. The main setting for people who were merged together.
Minimum turn duration 0.3 s Speech shorter than this is ignored. Raising it removes breaths and clicks, and also loses short answers like "yes".
Same-speaker merge gap 1.0 s Joins a speaker's consecutive turns across pauses up to this long. 0 turns it off.
Speaker re-link threshold 0.35 (CAM++) Higher merges more voices after clustering. The recommended range for the chosen model is shown, and turns amber outside it.
Advanced settings with a warning and fields for chunk threshold 40, cluster threshold 0.7, minimum turn duration 0.3, merge gap 1 and re-link threshold 0.35
Advanced settings, at their defaults.

12 · Better results

  • Start with the best recording. If you recorded on several devices, use the clearest one. A microphone close to the speakers beats any setting. Where a camera and a separate audio recorder captured the same event, transcribe the audio recorder's file.
  • Set the language when you know it.
  • Use Large V3 unless speed matters more than accuracy. For a language Whisper handles less well, look for a custom model.
  • Tell it how many people speak. An Expected speakers number removes most stray labels.
  • Listen where it's red. Low confidence clusters are where corrections pay off.
  • Try settings on a short excerpt before running a whole project. Speaker settings that suit one room or group usually suit the rest.

13 · Troubleshooting

The same phrase repeats again and again

A known Whisper habit after long silences or music. Dote Whisper no longer carries words from one 30-second window into the next, which was the main cause, so it should be rare. If it still happens, turn on speaker labels: silences are then mostly skipped.

The text is in the wrong language, or nonsense

Set the Language instead of auto-detect. If you chose an English Only model for a recording in another language, the file stops with an error that explains this: pick a multilingual model.

The Tuning tab says "unavailable for this transcript"

The transcript was made without speaker labels, or the voiceprints couldn't be saved. Tick Enable Speaker Diarization, click and transcribe again. The Console shows any error.

It's slow

Check the Progress panel: CPU means no GPU is being used, which is expected on Intel Macs and on PCs without an NVIDIA card. Try Medium or Base, and leave long batches to run overnight.

The memory badge turns red, or a run fails with a memory error

Close other programs. Long recordings with speaker labels need the most memory. The app retries pieces that run out in smaller parts, and uses shorter pieces on computers with less memory. If it still fails, lower the diarization chunk threshold in Advanced settings, or use a smaller Whisper model.

A custom model won't install

"Not found" usually means a typo, or a private model: only public models can be installed. "Incompatible" means it isn't a Whisper speech-to-text model, so try another. A network error is worth retrying. Some converted models give slightly less precise word times (about 20 ms instead of 10); the words themselves are unaffected.

Windows says "Windows protected your PC"

The installer isn't code-signed yet. Click More info, then Run anyway.

The NVIDIA GPU isn't used on Windows

Update the NVIDIA driver. If the GPU fails mid-run, the app finishes on the CPU instead of stopping. Remote-desktop and virtual-machine sessions usually can't use the GPU.

Where are my files?

What Where
Transcripts Beside each recording (auto-save), or where you saved them. Use in the file list.
Models The folder in Settings → Whisper Models Directory; its folder button opens it.
Settings (macOS) ~/Library/Application Support/DOTE Whisper/
Settings (Windows) %APPDATA%\DOTE Whisper\

How do I update?

When a new version is out, the app tells you as it starts, or when you choose File → Check for updates…. Download it from the download page and install it over the old one. Your settings and models are kept.

How do I cite it?

If you use DOTE Whisper in published work, please cite it along with the projects it builds on: OpenAI Whisper (Radford et al., 2022, Robust Speech Recognition via Large-Scale Weak Supervision), whisper.cpp (Georgi Gerganov) and sherpa-onnx (the k2-fsa project). Reporting the model, the speaker model and the post_processing_settings from your JSON files makes your method reproducible.

Ready to try it?

Download DOTE Whisper and transcribe your first recording.

Download free