tiny-audio, nanoGPT for speech-to-text Tiny Audio, a speech-to-text system from developer mazesmazes, connects a frozen pretrained speech encoder to a pretrained LLM via a small trainable projector and trains only about 80M parameters for roughly $25, achieving 1.8% WER on LibriSpeech test-clean and 7.4% across 12 benchmarks (11,822 samples pooled). The model is published on Hugging Face as mazesmazes/tiny-audio and runs through the transformers pipeline with word-level timestamps and speaker diarization, requiring about 6 GB of GPU or Apple Silicon memory for bf16 weights. A `ta serve` batched HTTP server reaches about 460x real time at 128 concurrent requests on an RTX 4090, and speaker diarization requires transformers installed from main until the next release. A speech-to-text system you can train for $25. Tiny Audio connects a frozen, pretrained speech encoder to a pretrained LLM with a small trainable projector. The model published from this repo gets 1.8% WER on LibriSpeech test-clean and 7.4% across 12 benchmarks 11,822 samples pooled while training only ~80M parameters. The codebase is small enough to read in an afternoon, and you can run a training loop on your laptop in about five minutes. No install: open the live demo https://huggingface.co/spaces/mazesmazes/tiny-audio , record yourself or upload a file, and get a transcript. In Python: pip install "transformers =5.0" peft torch torchaudio librosa python from transformers import pipeline pipe = pipeline "automatic-speech-recognition", model="mazesmazes/tiny-audio", trust remote code=True print pipe "audio.wav" "text" The quarterly revenue grew by 12% according to Dr. Smith. The output is punctuated, capitalized, and has numbers formatted, with no post-processing step. The input can be a file path, a URL, or a 16 kHz numpy array. Weights are bf16, so you need roughly 6 GB of GPU or Apple Silicon memory. Word-level timestamps forced alignment pipe "audio.wav", return timestamps=True {"text": "hello world", "words": {"word": "hello", "start": 0.0, "end": 0.5}, ... } Who spoke when speaker diarization pipe "meeting.wav", return speakers=True, num speakers=2 Each speaker Nemotron-3-Diarization finds is transcribed separately, on a copy of the audio where everyone else is silenced, and each word belongs to the stream it came from. This is a zero-shot port of NeMo's masked asr recipe; a word two streams both heard at once is kept once. Each speaker costs roughly their own talk time in ASR, and single-speaker audio is transcribed unmasked. Overlap is only partly handled: another person's speech inside a speaker's turn stays in that speaker's stream. Speaker diarization needs transformers installed from main pip install git+https://github.com/huggingface/transformers until the next release. For token-by-token streaming output, see ASRModel.generate streaming https://github.com/alexkroman/tiny-audio/blob/main/tiny audio/asr modeling.py . The model card https://huggingface.co/mazesmazes/tiny-audio covers batching and GPU settings. ta serve puts the model behind a batched HTTP server: requests arriving together share GPU batches, so throughput grows with load about 460x real time at 128 concurrent requests on an RTX 4090 . To run it on a RunPod GPU: poetry run ta runpod up --serve create an inference pod; prints