{"slug": "1-year-of-building-a-bring-your-own-api-key-transcription-app-and-why-it-s-now", "title": "1 Year of Building a \"Bring Your Own API Key\" Transcription App — and Why It's Now an Automation Hub", "summary": "A solo iOS developer built WhisperDirect, a one-time-purchase transcription app that has no backend and instead has users supply their own OpenAI API key, paying roughly $0.006 per minute of Whisper API transcription. Over a year of development, the app added on-device Apple Speech, speaker diarization, OCR, and summarization via OpenAI, Gemini, or any OpenAI-compatible endpoint, with the developer reporting that Apple Speech on iOS 26 matched or beat self-hosted engines like Parakeet and ReazonSpeech in his own benchmarks.", "body_md": "Hi, I'm a solo iOS developer.\n\nAbout a year ago, when I first released my transcription app, I wrote this:\n\n\"The moment you go subscription, you can't stop.\"\n\n\"Even if only one person has ever paid, you have to keep the servers running — even at a loss.\"\n\n\"A single traffic spike can degrade your service.\"\n\nWhile every AI app was racing toward monthly subscriptions, I realized that for a solo developer, running a backend long-term is just too much risk. Then it hit me:\n\n**What if I just build the \"vessel\" — and let the user bring their own OpenAI API key?**\n\nA year later, WhisperDirect has grown into something with transcription, summarization, meeting minutes, and external automation built in.\n\nThis post is the story of that year. Why an \"API-direct, fully one-time-purchase\" transcription app ended up here.\n\nWhisperDirect started with a very simple idea:\n\nIn other words, WhisperDirect was, from day one, **a vessel for using OpenAI via your own API key**. No backend to maintain, no middleman markup. You pay OpenAI for what you use. That part of the design has never changed.\n\nOn the other hand, I had another app: **WhisText**.\n\nThink of it like Fedora vs. Red Hat in the Linux world.\n\nCutting-edge tech goes into **WhisText** first. Whatever survives real-world use gets ported into **WhisperDirect**.\n\nOver the past year, that cycle has accelerated dramatically.\n\nWhisText's transcription originally ran on my own GPU server. The model was Whisper large-v3-turbo. My plan was: let people use it for free, watch the usage, then figure out pricing.\n\nReality was less generous.\n\nUsage never justified keeping a GPU running 24/7, and the bills kept coming. So I switched to CPU and started testing every ASR engine I could find.\n\nMy dev server's `asr.sherpa` directory still holds the scars:\n\nI tried different language combinations, split audio into chunks for parallel processing, tuned for speed. A lot of unglamorous trial and error.\n\nEventually, **Apple Speech became the center of WhisText's final version.**\n\nI had actually tried Apple's native Apple Speech pretty early on.\n\nAt the time, it was unusable. The biggest blocker was the **1-minute limit**. For an app that records meetings and produces minutes, cutting off every 60 seconds is fatal. So I went back to tuning my own server.\n\nThen iOS 26 changed everything.\n\nApple Speech's accuracy jumped, and the old limits and instability were gone. In my own benchmarks, it was **matching or beating** engines I'd been running on my own server — Parakeet, ReazonSpeech, all of them.\n\nSo WhisText's final version put Apple Speech at the core.\n\nNo server round-trip means: **the instant you speak into the mic, the text appears.** I added live preview as a bonus.\n\nThat said, live preview is a bonus. If you want accurate meeting minutes, batch-processing the whole recording with Whisper API afterward still produces cleaner results. Live preview is more about the *feeling* — \"it's recording,\" \"AI is working right now.\"\n\nBut if Apple Speech on iOS 26 is this good, there must be more we can do on-device.\n\nSo the modules that moved to on-device in WhisText got ported into WhisperDirect:\n\n**These became the foundation that works without an API key, for free.** The accuracy improvements in iOS 26's Apple Speech are what made this foundation actually usable.\n\nThe latest version — currently in App Store review — packs in a lot more.\n\nWhisperDirect is a voice transcription and summarization app built around Whisper API's accuracy. Alongside Whisper API (with your own key), it ships with on-device features: Apple Speech, speaker diarization, and OCR. For summarization and minutes, you can choose your LLM from OpenAI, Gemini, or any OpenAI-compatible endpoint.\n\n**The app is a one-time purchase. No subscriptions.**\n\nCost reference: Whisper API is about **$0.006/min**.\n\nThat's roughly **$0.36/hour** — about **8.7 hours of transcription for around $3**.\n\n**With a local LLM, you can bring the summarization / minutes cost down to zero.**\n\nFor a few thousand characters, cloud APIs cost just a few cents. That kind of freedom is only possible for a solo developer.\n\nI built this for myself, honestly.\n\nYou can POST transcription / summary / minutes results to an external endpoint, each at the moment it completes. From there, **n8n handles the rest.** With n8n, you can do whatever LLM processing you want, then route to Notion, Google Sheets, Slack, or anything else — just configure the URL. Header auth is supported too.\n\nThe payload looks like this:\n\n```\n{\n  \"text\": \"...\",\n  \"type\": \"transcript\",\n  \"recorded_at\": \"2026-10-09T10:00:00Z\",\n  \"sent_at\": \"2026-10-09T10:05:00Z\",\n  \"source\": \"whisperdirect\",\n  \"version\": \"1.0\"\n}\n```\n\nFor a company-run app, you need a monthly subscription just to cover server costs, salaries, and ad spend.\n\nBut WhisperDirect has no backend to maintain.\n\nThe app is a one-time purchase.\n\nWhen an API is used, the user pays the provider directly — at cost.\n\nI can't spend money on ads.\n\nSo instead of competing on ad spend, I compete on cost-performance and freedom.\n\nWhat I can do as a solo developer, and what I can't.\n\nI've been carrying both, the whole way here.\n\nThe latest version is currently in App Store review (live preview, etc.).\n\nThere's a 7-day free trial.\n\nIf this sounds interesting, search for WhisperDirect on the App Store.\n\nOnce it's approved, give the new transcription experience a try.\n\nMy recommended setup:\n\nThe point is: **you don't have to choose between \"cheap\" and \"accurate.\" You decide per recording, per segment.**\n\nWhisperDirect: [https://apps.apple.com/app/id6748595475](https://apps.apple.com/app/id6748595475)", "url": "https://wpnews.pro/news/1-year-of-building-a-bring-your-own-api-key-transcription-app-and-why-it-s-now", "canonical_source": "https://dev.to/whisperdirect/1-year-of-building-a-bring-your-own-api-key-transcription-app-and-why-its-now-an-automation-hub-57jk", "published_at": "2026-10-09 20:48:10+00:00", "updated_at": "2026-10-09 20:54:22.688920+00:00", "lang": "en", "topics": ["ai-products", "ai-tools", "natural-language-processing", "large-language-models"], "entities": ["WhisperDirect", "WhisText", "OpenAI", "Whisper API", "Apple Speech", "Gemini", "Parakeet", "ReazonSpeech"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/1-year-of-building-a-bring-your-own-api-key-transcription-app-and-why-it-s-now", "markdown": "https://wpnews.pro/news/1-year-of-building-a-bring-your-own-api-key-transcription-app-and-why-it-s-now.md", "text": "https://wpnews.pro/news/1-year-of-building-a-bring-your-own-api-key-transcription-app-and-why-it-s-now.txt", "jsonld": "https://wpnews.pro/news/1-year-of-building-a-bring-your-own-api-key-transcription-app-and-why-it-s-now.jsonld"}}