From a research perspective, modern AI development is quite interesting. Some of the AI systems in global use are arguably science fiction material. That said, the way AI is currently deployed is, in my opinion, questionable on several fronts.
As a PC enthusiast, I find the current state of GPU and memory prices depressing. As a consumer, I dislike ads and subscription services: both are hostile to the people they’re supposedly serving. As a software developer, I don’t like cloud computing or micro-services, these are only appropriate for a narrow class of applications, big tech companies who genuinely need the complexity to handle their scale.
Another thing about the status quo, from a software engineer’s perspective the ecosystem is not ideal. Most AI frameworks are developed in close collaboration with NVIDIA which uses them to push sales of CUDA GPUs, despite modern operating systems implementing vendor-agnostic APIs to leverage the same hardware. The community is heavily focused on Python: a fine choice for research, poor fit for complex desktop software.
The speech-to-text software I’ve built reflects these priorities directly: 17 MB full installer with no runtime dependencies, no subscription.
After a one-time licence activation and model download, it requires no internet connection whatsoever: no telemetry, no licence checks, no analytics. For fully air-gapped machines, offline activation is also supported: the UX is less smooth, involving some copy-pasting between computers, but it works.
This project traces back to late 2022. I was impressed by the performance of whisper.cpp by Georgi Gerganov, though the initial versions only supported CPU compute. Having spent years doing GPU development, I had a daft idea: port whisper.cpp from CPU-based SIMD to a GPU based DirectCompute stack.
You might be wondering, why DirectCompute specifically? CUDA is locked to nVidia, and nVidia GPUs are a minority of machines in global use. I could have used Vulkan, as whisper.cpp eventually did. However, in my experience, non-default GPU APIs and multi-platform abstraction layers tend to cause headaches. Modern GPUs are incredibly complicated, much of that complexity leaks across the PCIe bus and inflates the drivers accordingly. The fact that a GPU is a single hardware device while all modern OSes are multi-user and multitasking makes this worse still. My experience has been that first-party GPU APIs (Direct3D on Windows, Metal on Apple platforms) are simply more reliable.
Unlike most of my mad ideas, I actually followed through on this one – Christmas holidays were well-timed. The result is open source.
The project validated the high-level architecture: D3D11 compute shaders proved reliably cross-hardware in practice, and substantially faster than CPU SIMD even on integrated graphics.
I also discovered C++ language sucks for this sort of work. GPGPU pipelines do very little on the CPU – they mostly make kernel calls to manage VRAM resources and dispatch shaders. A higher-level language that trades raw CPU performance for usability would help considerably.
The next time I had the inclination to tinker with desktop ML was a year later. This time I shifted much of the complexity from C++ into C# JIT-compiled with the .NET 6 runtime.
The project validated a few more pieces of the puzzle. The .NET for higher-level code with thin C++ layer underneath approach worked well. I wired them together using the ComLightInterop library I’d written previously.
Unlike the original Whisper port – which was adapted from whisper.cpp – the
Mistral inference was ported directly from the OG Python/PyTorch/CUDA stack.
Some infrastructure from that project survived into the next iteration.
For instance, the *.cgml model serialisation format only needed a few adjustments.
Shortly after, a company acquisition landed an unexpected full-time job in my lap. For the next couple of years, free time for hobby projects was scarce. The idea, however, stayed with me.
Eventually, another corporate event outside my control left me looking for my next role.
By that point, the original open-source speech-to-text project had accumulated over 10k stars on github. 10k stars on a free tool confirmed the problem was real, but said nothing about whether people would pay to solve it.
Still, I took my chances. Rather than looking for another job, I decided to become an independent software vendor, and began work on another incarnation of the speech-to-text software.
The goal: offline, commercial-quality speech-to-text for both audio files and real-time use on the desktop.
My performance target was exceeding real-time transcription rate on a median Windows machine. Since I don’t actually know where the median sits, I used my own laptop as a proxy: a Ryzen 5 5600U APU machine, four years old, no discrete graphics, bought cheap.
Here’s what I have built so far.
I don’t have a rigid long-term strategy. The plan is to launch and see what happens.