{"slug": "tiny-audio-nanogpt-for-speech-to-text", "title": "tiny-audio, nanoGPT for speech-to-text", "summary": "Tiny Audio, a speech-to-text system from developer mazesmazes, connects a frozen pretrained speech encoder to a pretrained LLM via a small trainable projector and trains only about 80M parameters for roughly $25, achieving 1.8% WER on LibriSpeech test-clean and 7.4% across 12 benchmarks (11,822 samples pooled). The model is published on Hugging Face as mazesmazes/tiny-audio and runs through the transformers pipeline with word-level timestamps and speaker diarization, requiring about 6 GB of GPU or Apple Silicon memory for bf16 weights. A `ta serve` batched HTTP server reaches about 460x real time at 128 concurrent requests on an RTX 4090, and speaker diarization requires transformers installed from main until the next release.", "body_md": "**A speech-to-text system you can train for $25.**\n\nTiny Audio connects a frozen, pretrained speech encoder to a pretrained LLM with a small trainable\nprojector. The model published from this repo gets **1.8% WER on LibriSpeech test-clean and 7.4%\nacross 12 benchmarks** (11,822 samples pooled) while training only ~80M parameters. The codebase is\nsmall enough to read in an afternoon, and you can run a training loop on your laptop in about five\nminutes.\n\n**No install:** open the **[live demo](https://huggingface.co/spaces/mazesmazes/tiny-audio)**,\nrecord yourself or upload a file, and get a transcript.\n\n**In Python:**\n\n```\npip install \"transformers>=5.0\" peft torch torchaudio librosa\npython\nfrom transformers import pipeline\n\npipe = pipeline(\n    \"automatic-speech-recognition\", model=\"mazesmazes/tiny-audio\", trust_remote_code=True\n)\nprint(pipe(\"audio.wav\")[\"text\"])\n# The quarterly revenue grew by 12% according to Dr. Smith.\n```\n\nThe output is punctuated, capitalized, and has numbers formatted, with no post-processing step. The input can be a file path, a URL, or a 16 kHz numpy array. Weights are bf16, so you need roughly 6 GB of GPU or Apple Silicon memory.\n\n```\n# Word-level timestamps (forced alignment)\npipe(\"audio.wav\", return_timestamps=True)\n# {\"text\": \"hello world\", \"words\": [{\"word\": \"hello\", \"start\": 0.0, \"end\": 0.5}, ...]}\n\n# Who spoke when (speaker diarization)\npipe(\"meeting.wav\", return_speakers=True, num_speakers=2)\n```\n\nEach speaker Nemotron-3-Diarization finds is transcribed separately, on a copy of the audio where\neveryone else is silenced, and each word belongs to the stream it came from. This is a zero-shot\nport of NeMo's `masked_asr` recipe; a word two streams both heard at once is kept once. Each speaker\ncosts roughly their own talk time in ASR, and single-speaker audio is transcribed unmasked. Overlap\nis only partly handled: another person's speech inside a speaker's turn stays in that speaker's\nstream.\n\nSpeaker diarization needs `transformers` installed from `main`\n(`pip install git+https://github.com/huggingface/transformers`) until the next release. For\ntoken-by-token streaming output, see [`ASRModel.generate_streaming`](https://github.com/alexkroman/tiny-audio/blob/main/tiny_audio/asr_modeling.py).\nThe [model card](https://huggingface.co/mazesmazes/tiny-audio) covers batching and GPU settings.\n\n`ta serve` puts the model behind a batched HTTP server: requests arriving together share GPU\nbatches, so throughput grows with load (about 460x real time at 128 concurrent requests on an RTX\n4090). To run it on a RunPod GPU:\n\n```\npoetry run ta runpod up --serve                 # create an inference pod; prints <POD_ID>\npoetry run ta runpod wait <POD_ID>              # prints <HOST> <PORT>\npoetry run ta runpod deploy <HOST> <PORT>       # sync the project, install the fast kernels\nTINY_AUDIO_API_KEY=my-secret poetry run ta runpod serve <HOST> <PORT> --no-attach\n# Ready when https://<POD_ID>-8000.proxy.runpod.net/health answers (a few minutes: it compiles first)\n```\n\nWithout `TINY_AUDIO_API_KEY` the server is open to anyone who has the URL. `ta serve` also runs\nlocally, on CUDA, Apple Silicon, or CPU.\n\nSend the audio as the request body, with options in the query string:\n\n```\ncurl -X POST \"https://<POD_ID>-8000.proxy.runpod.net/?return_timestamps=true\" \\\n  -H \"Authorization: Bearer my-secret\" \\\n  -H \"Content-Type: application/octet-stream\" \\\n  --data-binary @audio.wav\npython\nimport httpx\n\nresponse = httpx.post(\n    \"https://<POD_ID>-8000.proxy.runpod.net/\",\n    params={\"return_speakers\": \"true\", \"num_speakers\": \"2\"},\n    content=open(\"meeting.wav\", \"rb\").read(),\n    headers={\"Authorization\": \"Bearer my-secret\"},\n    timeout=600,\n)\nprint(response.json()[\"text\"])\n```\n\n- **Options:**`return_timestamps` ,`return_speakers` ,`num_speakers` and`max_speakers` , as in the\npipeline. The response is the same dict the pipeline returns.\n- **JSON body:** to send JSON instead, use`{\"inputs\": \"<base64 audio>\", \"parameters\": {...}}` .\n- **Audio formats:** anything FFmpeg can read.\n- **Errors:**`400` with`{\"error\": ...}` for bad audio or options, and`401` for a wrong key.\n- **Other endpoints:**`GET /health` and`GET /stats` (batch sizes and GPU time).\n\nRunPod's HTTP proxy rejects request bodies over 500 MiB, and it drops any request that takes more than 100 seconds. For long recordings, send 16 kHz mono FLAC:\n\n```\nffmpeg -i recording.wav -ac 1 -ar 16000 recording.flac\n```\n\nThat's about 1 MB per minute of audio, and it costs nothing in accuracy, because the server converts\neverything to 16 kHz mono anyway. On an RTX 4090, 45 minutes of audio takes about 10 seconds, or 21\nseconds with speaker labels. The [demo Space](https://github.com/alexkroman/tiny-audio/blob/main/demo/app.py) calls the server this way.\n\nWord error rate (%, lower is better) on 11,822 samples (up to 1,000 per dataset), measured with this\nrepo's `ta eval` against the `ta serve` HTTP API on an RTX 4090:\n\n| Dataset | WER | \n|---|---|\n| LibriSpeech test-clean | 1.84 | \n| SPGISpeech | 2.29 | \n| TED-LIUM | 3.71 | \n| LoquaciousSet † | 6.20 | \n| LibriSpeech test-other | 6.38 | \n| VoxPopuli | 7.11 | \n| Common Voice | 7.18 | \n| AMI (IHM) | 8.99 | \n| GigaSpeech | 9.06 | \n| Earnings22 † | 10.58 | \n| People's Speech | 17.59 | \n| AMI (SDM) | 23.53 | \n| **Mean (12 sets)** | **8.71** | \n| **Pooled (11,822 samples)** | **7.42** | \n\n† Held out: no data from this source was used in training.\n\nYou can check these numbers yourself and compare against commercial APIs on the same samples:\n\n```\npoetry run ta eval -m mazesmazes/tiny-audio -d loquacious -n 100\n# Same samples through a commercial API (also: deepgram, elevenlabs, apple-speech)\nASSEMBLYAI_API_KEY=... poetry run ta eval -m assemblyai -d loquacious -n 100\nAudio (16 kHz) → speech encoder (frozen) → MLP projector (trained) → LLM decoder → Text\n```\n\n1. A pretrained **speech encoder** turns audio into a sequence of frame embeddings.\n2. A small **MLP projector** stacks neighbouring frames and maps them into the LLM's embedding\nspace. It is the only part trained from scratch.\n3. The **LLM** reads those projected frames as if they were tokens and writes out the transcript.\n\nEncoder, projector, and decoder are each swappable from config. Two recipes ship with the repo:\n\n| Recipe | Encoder | Decoder | Trained | \n|---|---|---|---|\n| Published model ( `granite_qwen_frozen` ) | Granite Speech 470M | Qwen3.5, frozen + LoRA | Projector + LoRA (~80M) | \n| Default / course recipe ( `stage_1` ) | GLM-ASR-Nano (635M) | Qwen3-0.6B, fine-tuned | Projector + decoder | \n\nStart on your laptop for free, and rent a GPU only once you know the pipeline works.\n\n| Tier | Data | Hardware | Cost | \n|---|---|---|---|\n| **Smoke test** | 73 LibriSpeech clips | Your laptop (CPU, MPS, CUDA) | Free, ~5 minutes | \n| **Course run** | LoquaciousSet `small` (~250 hours) | One rented GPU, a few hours | A few GPU-hours | \n| **Production recipe** | ~3M clips across ten corpora (>1 TB) | One 80 GB GPU, a day or more | Hundreds of dollars | \n\n```\ngit clone https://github.com/alexkroman/tiny-audio.git && cd tiny-audio\npoetry install\n\n# 1. Smoke test: a real training loop on your laptop\npoetry run python scripts/train.py +experiments=mps_smoke\n\n# 2. Before renting hardware, estimate the VRAM and disk a config needs\npoetry run ta runpod plan -e stage_1\n\n# 3. Full run\npoetry run python scripts/train.py +experiments=stage_1\n```\n\nEvery setting is a [Hydra](https://hydra.cc/) override, for example\n`model.projector_hidden_dim=2048` or `training.use_lora=true`. When you're happy with a model,\n`ta push` publishes it to the Hugging Face Hub and `ta deploy` puts a demo like the one above on a\nSpace.\n\n```\npoetry run ta runpod up -e stage_1              # create a pod with enough GPU for the config\npoetry run ta runpod wait <POD_ID>              # prints <HOST> <PORT>\npoetry run ta runpod deploy <HOST> <PORT>       # sync the project and install dependencies\nHF_TOKEN=hf_... poetry run ta runpod train <HOST> <PORT> -e stage_1\npoetry run ta runpod attach <HOST> <PORT>       # watch the run in tmux\n```\n\nThe **[free 3.5-hour course](https://github.com/alexkroman/tiny-audio/blob/main/docs/course/0-course-overview.md)** walks you through the full loop:\nhow the encoder, projector, and decoder fit together (with real tensor shapes), training a model,\nevaluating it against commercial APIs, and publishing it with a live demo. You need Python, the\ncommand line, and git. The course trains the smaller `stage_1` recipe, not the published model, so\nyour WER will be higher than the table above.\n\nWant to try a new projector architecture, add a dataset, or change the codebase? See\n[CONTRIBUTING.md](https://github.com/alexkroman/tiny-audio/blob/main/CONTRIBUTING.md) for the CLI reference, config layout, and quality gates.\n\n- [Granite Speech](https://huggingface.co/ibm-granite/granite-speech-5.0-470m-turboctc) and[GLM-ASR](https://huggingface.co/zai-org/GLM-ASR-Nano-2512) for audio encoding\n- [Qwen3.5](https://huggingface.co/Qwen/Qwen3.5-2B) and[Qwen3](https://huggingface.co/Qwen/Qwen3-0.6B) for language modeling\n- [LoquaciousSet](https://huggingface.co/datasets/speechbrain/LoquaciousSet) for the default\nevaluation set\n\nMIT", "url": "https://wpnews.pro/news/tiny-audio-nanogpt-for-speech-to-text", "canonical_source": "https://github.com/alexkroman/tiny-audio", "published_at": "2026-10-09 19:52:07+00:00", "updated_at": "2026-10-09 20:21:56.161318+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "natural-language-processing", "ai-tools", "developer-tools"], "entities": ["Tiny Audio", "mazesmazes", "Hugging Face", "LibriSpeech", "Nemotron-3-Diarization", "NeMo", "transformers", "RTX 4090"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/tiny-audio-nanogpt-for-speech-to-text", "markdown": "https://wpnews.pro/news/tiny-audio-nanogpt-for-speech-to-text.md", "text": "https://wpnews.pro/news/tiny-audio-nanogpt-for-speech-to-text.txt", "jsonld": "https://wpnews.pro/news/tiny-audio-nanogpt-for-speech-to-text.jsonld"}}