If you have a discrete GPU with less than 4 GB VRAM, sadly none of the Voxtral models will run well. They usually fall back to system RAM causing 100% PCIe bus load, and slower than real-time performance. Or sometimes they fail with an out of memory error. Try Whisper Turbo BCML instead – it uses under 1 GB of VRAM, runs much faster, but in my tests the output text quality was noticeably worse than Voxtral.
If you have a discrete GPU with 4 GB VRAM, or integrated graphics with at least 8 GB system memory, download the 3.4 GB “Voxtral Mini BCML” model and select the low preset on the first screen.
If you have a discrete GPU with at least 6 GB VRAM, or current generation integrated graphics with at least 16 GB system memory, download the 4.6 GB “Voxtral Mini BCML text” model and select the medium preset on the first screen.
If you have a discrete GPU with at least 10 GB VRAM, you’re welcome to try the uncompressed 8.5 GB “Voxtral Mini FP16” model with the medium or high preset. That said, in my tests the word error rates were extremely close to the compressed “BCML Text” model across all 12 languages I benchmarked. IMO the uncompressed model is not worth the extra electricity bill.
All transcriptions used a 2400 ms delay: best quality, and the only option the software supports for file transcription. See Performance.pdf for the longer tables which include different transcription delays.
| Key | Model | Preset |
|---|---|---|
| Python | Voxtral Mini 4B Realtime 2602, 8.25 GB | Reference |
| High | Voxtral Mini FP16, 8.5 GB | High |
| Medium | Voxtral Mini BCML text, 4.6 GB | Medium |
| Low | Voxtral Mini BCML, 3.4 GB | Low |
The tables also include Mistral rows with the figures published on Hugging Face by Mistral AI. Note that Mistral’s numbers likely use a different text normalisation algorithm, as well as a different subset of the FLEURS dataset.
| Model | Python | High | Medium | Low |
|---|---|---|---|---|
| Error rate | 2.81% | 2.82% | 2.85% | 2.83% |
The tables below contain error rates for Fleurs dataset, test subset, 2400 milliseconds delay.
| Model | German | English | Spanish | French | Italian | Dutch | Portuguese |
|---|---|---|---|---|---|---|---|
| Mistral | 4.15% | 4.05% | 2.71% | 5.23% | 2.37% | 5.91% | 3.93% |
| Python | 4.11% | 4.07% | 2.67% | 5.20% | 2.25% | 5.57% | 3.86% |
| High | 4.08% | 4.09% | 2.66% | 5.27% | 2.18% | 5.59% | 3.85% |
| Medium | 4.18% | 4.14% | 2.72% | 5.37% | 2.25% | 5.87% | 3.97% |
| Low | 4.22% | 4.16% | 2.72% | 5.56% | 2.31% | 6.03% | 3.99% |
| Model | Arabic | Hindi | Chinese | Japanese | Korean |
|---|---|---|---|---|---|
| Mistral | 14.71% | 10.73% | 8.48% | 5.50% | 14.30% |
| Python | 10.46% | 13.37% | 5.65% | 5.33% | 8.10% |
| High | 10.70% | 13.35% | 5.67% | 5.30% | 8.12% |
| Medium | 10.84% | 13.40% | 5.92% | 5.40% | 8.40% |
| Low | 10.47% | 13.52% | 5.92% | 5.47% | 8.29% |
The table below lists memory usage of the transcription engine with Voxtral models when selecting a preset in the engine options. The figures reflect memory used by the speech-to-text engine measured in isolation; GUI frontends will consume additional memory on top.
| Model | VRAM usage by presets | ||
|---|---|---|---|
| Low | Medium | High | |
| Voxtral BCML | 3.3 GB | 3.4 GB | 4.3 GB |
| Voxtral BCML text | 4.5 GB | 4.6 GB | 5.6 GB |
| Voxtral FP16 | 8.4 GB | 8.5 GB | 9.4 GB |
The following table lists transcription rates relative to real-time across several GPUs I have benchmarked. A value of 1.0 means real-time transcription speed; higher is better. I have measured the numbers using the software version 0.9.6.
Take these numbers with a grain of salt, and make sure to evaluate the software on your specific machine before purchasing. I provide 30 days evaluation licences partly for that reason.
The Dictation application has a real-time indicator labelled “GPU load”, which displays the inverse of this value expressed as a percentage. If the figure stays below 90% even whilst speaking quickly, the GPU is capable enough for your chosen combination of model and settings. Note that if you are using a laptop with two GPUs, you may need to select the faster discrete GPU on the application’s first screen.
| GPU | Model | Preset | Rate |
|---|---|---|---|
| Ryzen 5 5600U iGPU | Voxtral BCML | low | 1.36 |
| Voxtral BCML text | medium | 1.09 | |
| Voxtral FP16 | high | 0.81 | |
| Ryzen 7 8845HS iGPU | Voxtral BCML | low | 3.21 |
| Voxtral BCML text | medium | 3.28 | |
| Voxtral FP16 | high | 1.55 | |
| Ryzen 7 8700G iGPU | Voxtral BCML | low | 3.58 |
| Voxtral BCML text | medium | 3.57 | |
| Voxtral FP16 | high | 1.54 | |
| GeForce 1650 | Voxtral BCML | low | 2.88 |
| Voxtral BCML text | medium | 1.39 | |
| Voxtral FP16 | high | OoM | |
| GeForce 1080 Ti | Voxtral BCML | low | 6.67 |
| Voxtral BCML text | medium | 6.6 | |
| Voxtral FP16 | high | 5.63 | |
| GeForce 4070 Ti Super | Voxtral BCML | low | 10.9 |
| Voxtral BCML text | medium | 10.58 | |
| Voxtral FP16 | high | 9.48 |
The laptop with the 5600U was running on AC power during the benchmark: it is a 2021 machine and the battery has degraded noticeably over the years. The newer laptop with the 8845HS was running on battery power during the benchmark.
OoM means “Out of Memory”: an error message which says “Insufficient memory to continue the execution”