TLDR: which model and engine options?

If you have a discrete GPU with less than 4 GB VRAM, sadly none of the Voxtral models will run well. They usually fall back to system RAM causing 100% PCIe bus load, and slower than real-time performance. Or sometimes they fail with an out of memory error. Try Whisper Turbo BCML instead – it uses under 1 GB of VRAM, runs much faster, but in my tests the output text quality was noticeably worse than Voxtral.

If you have a discrete GPU with 4 GB VRAM, or integrated graphics with at least 8 GB system memory, download the 3.4 GB “Voxtral Mini BCML” model and select the low preset on the first screen.

If you have a discrete GPU with at least 6 GB VRAM, or current generation integrated graphics with at least 16 GB system memory, download the 4.6 GB “Voxtral Mini BCML text” model and select the medium preset on the first screen.

If you have a discrete GPU with at least 10 GB VRAM, you’re welcome to try the uncompressed 8.5 GB “Voxtral Mini FP16” model with the medium or high preset. That said, in my tests the word error rates were extremely close to the compressed “BCML Text” model across all 12 languages I benchmarked. IMO the uncompressed model is not worth the extra electricity bill.

Quality Benchmarks

All transcriptions used a 2400 ms delay: best quality, and the only option the software supports for file transcription. See Performance.pdf for the longer tables which include different transcription delays.

Model Keys

KeyModelPreset
PythonVoxtral Mini 4B Realtime 2602, 8.25 GBReference
HighVoxtral Mini FP16, 8.5 GBHigh
MediumVoxtral Mini BCML text, 4.6 GBMedium
LowVoxtral Mini BCML, 3.4 GBLow

The tables also include Mistral rows with the figures published on Hugging Face by Mistral AI. Note that Mistral’s numbers likely use a different text normalisation algorithm, as well as a different subset of the FLEURS dataset.

Long Form English

ModelPythonHighMediumLow
Error rate2.81%2.82%2.85%2.83%

Fleurs

The tables below contain error rates for Fleurs dataset, test subset, 2400 milliseconds delay.

Latin Script Languages

ModelGermanEnglishSpanishFrenchItalianDutchPortuguese
Mistral4.15%4.05%2.71%5.23%2.37%5.91%3.93%
Python4.11%4.07%2.67%5.20%2.25%5.57%3.86%
High4.08%4.09%2.66%5.27%2.18%5.59%3.85%
Medium4.18%4.14%2.72%5.37%2.25%5.87%3.97%
Low4.22%4.16%2.72%5.56%2.31%6.03%3.99%

Non-Latin Script Languages

ModelArabicHindiChineseJapaneseKorean
Mistral14.71%10.73%8.48%5.50%14.30%
Python10.46%13.37%5.65%5.33%8.10%
High10.70%13.35%5.67%5.30%8.12%
Medium10.84%13.40%5.92%5.40%8.40%
Low10.47%13.52%5.92%5.47%8.29%

Performance Benchmarks

Memory Usage

The table below lists memory usage of the transcription engine with Voxtral models when selecting a preset in the engine options. The figures reflect memory used by the speech-to-text engine measured in isolation; GUI frontends will consume additional memory on top.

ModelVRAM usage by presets
LowMediumHigh
Voxtral BCML3.3 GB3.4 GB4.3 GB
Voxtral BCML text4.5 GB4.6 GB5.6 GB
Voxtral FP168.4 GB8.5 GB9.4 GB

Transcription Rates

The following table lists transcription rates relative to real-time across several GPUs I have benchmarked. A value of 1.0 means real-time transcription speed; higher is better. I have measured the numbers using the software version 0.9.6.

Take these numbers with a grain of salt, and make sure to evaluate the software on your specific machine before purchasing. I provide 30 days evaluation licences partly for that reason.

The Dictation application has a real-time indicator labelled “GPU load”, which displays the inverse of this value expressed as a percentage. If the figure stays below 90% even whilst speaking quickly, the GPU is capable enough for your chosen combination of model and settings. Note that if you are using a laptop with two GPUs, you may need to select the faster discrete GPU on the application’s first screen.

GPUModelPresetRate
Ryzen 5 5600U iGPU Voxtral BCML low 1.36
Voxtral BCML textmedium1.09
Voxtral FP16high0.81
Ryzen 7 8845HS iGPU Voxtral BCML low 3.21
Voxtral BCML textmedium3.28
Voxtral FP16high1.55
Ryzen 7 8700G iGPU Voxtral BCML low 3.58
Voxtral BCML textmedium3.57
Voxtral FP16high1.54
GeForce 1650 Voxtral BCML low 2.88
Voxtral BCML textmedium1.39
Voxtral FP16highOoM
GeForce 1080 Ti Voxtral BCML low 6.67
Voxtral BCML textmedium6.6
Voxtral FP16high5.63
GeForce 4070 Ti Super Voxtral BCML low 10.9
Voxtral BCML textmedium10.58
Voxtral FP16high9.48

The laptop with the 5600U was running on AC power during the benchmark: it is a 2021 machine and the battery has degraded noticeably over the years. The newer laptop with the 8845HS was running on battery power during the benchmark.

OoM means “Out of Memory”: an error message which says “Insufficient memory to continue the execution”