Hi, I'm a solo iOS developer.
About a year ago, when I first released my transcription app, I wrote this:
"The moment you go subscription, you can't stop."
"Even if only one person has ever paid, you have to keep the servers running — even at a loss."
"A single traffic spike can degrade your service."
While every AI app was racing toward monthly subscriptions, I realized that for a solo developer, running a backend long-term is just too much risk. Then it hit me:
What if I just build the "vessel" — and let the user bring their own OpenAI API key?
A year later, WhisperDirect has grown into something with transcription, summarization, meeting minutes, and external automation built in.
This post is the story of that year. Why an "API-direct, fully one-time-purchase" transcription app ended up here.
WhisperDirect started with a very simple idea:
In other words, WhisperDirect was, from day one, a vessel for using OpenAI via your own API key. No backend to maintain, no middleman markup. You pay OpenAI for what you use. That part of the design has never changed.
On the other hand, I had another app: WhisText.
Think of it like Fedora vs. Red Hat in the Linux world.
Cutting-edge tech goes into WhisText first. Whatever survives real-world use gets ported into WhisperDirect.
Over the past year, that cycle has accelerated dramatically.
WhisText's transcription originally ran on my own GPU server. The model was Whisper large-v3-turbo. My plan was: let people use it for free, watch the usage, then figure out pricing.
Reality was less generous.
Usage never justified keeping a GPU running 24/7, and the bills kept coming. So I switched to CPU and started testing every ASR engine I could find.
My dev server's asr.sherpa directory still holds the scars:
I tried different language combinations, split audio into chunks for parallel processing, tuned for speed. A lot of unglamorous trial and error.
Eventually, Apple Speech became the center of WhisText's final version.
I had actually tried Apple's native Apple Speech pretty early on.
At the time, it was unusable. The biggest blocker was the 1-minute limit. For an app that records meetings and produces minutes, cutting off every 60 seconds is fatal. So I went back to tuning my own server.
Then iOS 26 changed everything.
Apple Speech's accuracy jumped, and the old limits and instability were gone. In my own benchmarks, it was matching or beating engines I'd been running on my own server — Parakeet, ReazonSpeech, all of them.
So WhisText's final version put Apple Speech at the core.
No server round-trip means: the instant you speak into the mic, the text appears. I added live preview as a bonus.
That said, live preview is a bonus. If you want accurate meeting minutes, batch-processing the whole recording with Whisper API afterward still produces cleaner results. Live preview is more about the feeling — "it's recording," "AI is working right now."
But if Apple Speech on iOS 26 is this good, there must be more we can do on-device.
So the modules that moved to on-device in WhisText got ported into WhisperDirect:
These became the foundation that works without an API key, for free. The accuracy improvements in iOS 26's Apple Speech are what made this foundation actually usable.
The latest version — currently in App Store review — packs in a lot more.
WhisperDirect is a voice transcription and summarization app built around Whisper API's accuracy. Alongside Whisper API (with your own key), it ships with on-device features: Apple Speech, speaker diarization, and OCR. For summarization and minutes, you can choose your LLM from OpenAI, Gemini, or any OpenAI-compatible endpoint.
The app is a one-time purchase. No subscriptions.
Cost reference: Whisper API is about $0.006/min.
That's roughly $0.36/hour — about 8.7 hours of transcription for around $3.
With a local LLM, you can bring the summarization / minutes cost down to zero.
For a few thousand characters, cloud APIs cost just a few cents. That kind of freedom is only possible for a solo developer.
I built this for myself, honestly.
You can POST transcription / summary / minutes results to an external endpoint, each at the moment it completes. From there, n8n handles the rest. With n8n, you can do whatever LLM processing you want, then route to Notion, Google Sheets, Slack, or anything else — just configure the URL. Header auth is supported too.
The payload looks like this:
{
"text": "...",
"type": "transcript",
"recorded_at": "2026-10-09T10:00:00Z",
"sent_at": "2026-10-09T10:05:00Z",
"source": "whisperdirect",
"version": "1.0"
}
For a company-run app, you need a monthly subscription just to cover server costs, salaries, and ad spend.
But WhisperDirect has no backend to maintain.
The app is a one-time purchase.
When an API is used, the user pays the provider directly — at cost.
I can't spend money on ads.
So instead of competing on ad spend, I compete on cost-performance and freedom.
What I can do as a solo developer, and what I can't.
I've been carrying both, the whole way here.
The latest version is currently in App Store review (live preview, etc.).
There's a 7-day free trial.
If this sounds interesting, search for WhisperDirect on the App Store.
Once it's approved, give the new transcription experience a try.
My recommended setup:
The point is: you don't have to choose between "cheap" and "accurate." You decide per recording, per segment.
WhisperDirect: https://apps.apple.com/app/id6748595475