cd /news/artificial-intelligence/parakeet-java-automatic-speech-recog… · home topics artificial-intelligence article
[ARTICLE · art-138277] src=github.com ↗ pub= topic=artificial-intelligence verified=true sentiment=↑ positive

Parakeet.java: Automatic Speech Recognition in Pure Java

A pure-Java implementation of NVIDIA's Parakeet automatic speech recognition model, called Parakeet.java, matches the reference parakeet.cpp engine word-for-word at F16 quantization with 0.00% WER difference on LibriSpeech test-clean, according to benchmarks published by developer qxoticai. On an AMD Ryzen 7 PRO 8840U laptop with 8 threads, the tdt-0.6b-v3 Q4_K model reached 25.1x real-time factor versus parakeet.cpp's 12.6, while the smaller tdt_ctc-110m Q4_K model hit 81.2x versus 45.8x, transcribing an hour of speech in 44 seconds. The implementation runs without ONNX Runtime, whisper.cpp or a Python runtime, supports streaming transcription, word timings and confidence scores, and offers GraalVM Native Image support that loads a 600M model in about 0.2 seconds.

read5 min views4 publishedSep 23, 2026
Parakeet.java: Automatic Speech Recognition in Pure Java
Image: Michielbdejong (auto-discovered)

Automatic Speech Recognition for the JVM

An implementation of NVIDIA Parakeet for the JVM: blazing fast on ordinary CPUs, competitive with the native implementations. No ONNX Runtime, no whisper.cpp, no Python runtime. Just the JVM.

<sub>Live transcription of JFK's 1961 inaugural address. Play it with sound.</sub>

  • Competitive with the reference engines. Keeps pace withparakeet.cpp andsherpa-onnx , word for word identical.
  • Matches the reference with 100% accuracy. AtF16 the transcript matchesparakeet.cpp exactly; the rest is quantization noise, not drift.
  • Streaming support. Audio in as it arrives, text out in pieces, with a draft tail that keeps up with the speaker.
  • Word timings and confidence. Every token carries its audio span and how sure the decoder was.
  • Multilingual. Parakeet v3 transcribes and punctuates without a language flag.
  • First-class support for GraalVM's Native Image. Low memory footprint, standalone images, fast startup e.g. 600M model is loaded and ready to transcribe in about 0.2 s.

LibriSpeech test-clean, first 100 utterances (901 s), 8 threads on an AMD Ryzen 7 PRO 8840U laptop, median of 3 runs. The last column scores each run against parakeet.cpp's transcript rather than against the reference, so 0% means both engines heard exactly the same words. How to reproduce this, step by step.

Model Engine WER RTFx WER vs. parakeet.cpp
tdt-0.6b-v3 Q4_K jinfer 2.04% 25.1 0.17%
parakeet.cpp 2.04% 12.6 -
tdt-0.6b-v3 Q8_0 jinfer 2.04% 21.6 0.04%
parakeet.cpp 2.04% 14.9 -
tdt-0.6b-v3 F16 jinfer 2.04% 13.6 0.00%
parakeet.cpp 2.04% 15.3 -
tdt_ctc-110m Q4_K jinfer 1.99% 81.2 0.34%
parakeet.cpp 2.12% 45.8 -
v3 int8 ONNX sherpa-onnx 2.16% 18.6 1.27%

RTFx is seconds of audio transcribed per wall second, model load excluded. The small model transcribes an hour of speech in 44 seconds.

For the full matrix - both models, three engines, five quantizations, 1 to 16 threads and three JVMs, on a second machine - see BENCHMARKS.md.

Checkpoint Languages Parameters GGUF ( Q8_0 )
parakeet-tdt-0.6b-v3 multilingual 600M tdt-0.6b-v3-q8_0.gguf
parakeet-tdt-0.6b-v2 English 600M tdt-0.6b-v2-q8_0.gguf
parakeet-tdt-1.1b English 1.1B tdt-1.1b-q8_0.gguf
parakeet-tdt_ctc-1.1b English 1.1B tdt_ctc-1.1b-q8_0.gguf
parakeet-tdt_ctc-110m English 110M tdt_ctc-110m-q8_0.gguf

Every checkpoint also ships Q4_K, Q5_K, Q6_K and F16 in mudler/parakeet-cpp-gguf, named the same way: swap the quantization in the file name. Q8_0 balances quality and size, Q4_K runs quickest. The CLI downloads and caches them by reference, so no manual download is needed:

jinfer pull mudler/parakeet-cpp-gguf/tdt-0.6b-v3-q8_0.gguf
jinfer -m mudler/parakeet-cpp-gguf/tdt-0.6b-v3-q8_0.gguf --transcribe speech.wav

Any format ffmpeg reads, resampled by jinfer-codecs. The transcript goes to stdout, so it pipes.

Raw 16 kHz mono PCM on stdin, transcribed as it arrives. The microphone comes from ffmpeg, so the capture flags are the only part that differs:

ffmpeg -nostats -loglevel error -f pulse -i default -ar 16000 -ac 1 -f s16le - \
  | jinfer -m mudler/parakeet-cpp-gguf/tdt-0.6b-v3-q4_k.gguf --transcribe -
ffmpeg -nostats -loglevel error -f avfoundation -i ":0" -ar 16000 -ac 1 -f s16le - \
  | jinfer -m mudler/parakeet-cpp-gguf/tdt-0.6b-v3-q4_k.gguf --transcribe -
ffmpeg -nostats -loglevel error -f dshow -i audio="Microphone" -ar 16000 -ac 1 -f s16le - ^
  | jinfer -m mudler/parakeet-cpp-gguf/tdt-0.6b-v3-q4_k.gguf --transcribe -

On a terminal, stderr shows the live view: final words settle into the scrollback, each committed piece lands in color and fades into the text, the provisional tail follows in grey italics with a band of light sweeping through it, and a waveform dances with the voice. Doubtful words are flagged, so a name the model is unsure of stands out. --theme mint|nord|catppuccin|ember|frost|mono picks the palette; colors degrade to 256, to 16, and to plain attributes under NO_COLOR.

Redirected, the same run logs final text and partials as plain lines, so it scripts.

Through LangChain4j, which resolves the model ref and downloads it on first use:

try (var transcriber = JinferTranscriptionModel.builder()
        .model("mudler/parakeet-cpp-gguf/tdt-0.6b-v3-q8_0.gguf")
        .build()) {
    Transcription transcription = transcriber.transcribe(Path.of("speech.wav"));
    System.out.println(transcription.text());
    for (Transcription.Word word : transcription.words())
        System.out.printf("%s  %.2f  %s%n", word.start(), word.confidence(), word.text());
}

Or through Spring AI, with jinfer-spring-ai on the classpath:

try (var transcriber = JinferTranscriptionModel.builder()
        .model("mudler/parakeet-cpp-gguf/tdt-0.6b-v3-q8_0.gguf")
        .build()) {
    AudioTranscriptionResponse response = transcriber.call(
            new AudioTranscriptionPrompt(new FileSystemResource("speech.wav")));
    System.out.println(response.getResult().getOutput());
}

Live audio needs jinfer-parakeet directly: feed samples as they arrive, print the final pieces, and poll the provisional tail for a responsive UI.

static <S extends RuntimeState> void live(TranscriptionModel<?, ?, S> model, float[] pcm) {
    try (S state = model.newState();
            TranscriptionStream stream = model.stream(state)) {
        int chunk = model.sampleRate() / 2; // feed half a second at a time
        for (int at = 0; at < pcm.length; at += chunk) {
            Transcription piece = stream.feed(pcm, at, Math.min(chunk, pcm.length - at));
            System.out.print(piece.text()); // final text, often empty, never revised
            System.err.print("\r" + stream.partial().text()); // provisional, replaced next time
        }
        System.out.println(stream.finish().text()); // the rest, final
    }
}

Final pieces never change and join, in order, into the whole transcript. Token times are offsets from the start of the stream.

Audio is decoded in chunks with context on both sides, as NeMo streams Parakeet: 10 s of left context, a 2 s chunk, 2 s of right context. Only the chunk's tokens become final, while the decoder carries its state across chunks, so words continue across boundaries. Final text trails the speaker by 2 to 4 s; the provisional tail, decoded over a shorter window, refreshes twice a second while speech comes in.

Offline transcription runs the same path with 46 s chunks, so a two hour recording never holds a two hour attention matrix.

Parakeet v3 emits cased, punctuated text and one timing per token. It consumes mono 16 kHz audio; anything else is refused rather than silently resampled, since a rate mismatch quietly degrades recognition. Short windows around digital silence can blank (NVIDIA-NeMo/Speech#15757), so live recording with a room noise floor transcribes better than padded silence.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @parakeet.java 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/parakeet-java-automa…] indexed:0 read:5min 2026-09-23 ·