cd /news/artificial-intelligence/birdbuddy-an-offline-bird-call-ident… · home › topics › artificial-intelligence › article
[ARTICLE · art-147051] src=dev.to ↗ pub= topic=artificial-intelligence verified=true sentiment=↑ positive

BirdBuddy: an offline bird call identifier that runs entirely on your device

A developer built BirdBuddy, an offline bird call identifier that runs entirely on-device, combining Cornell's BirdNET ONNX model for species classification with Piper neural text-to-speech for spoken results. The project also fine-tuned Qwen3.5-4B with a LoRA adapter (r=16) via Thinking Machines' Tinker on a 165-example dataset covering 60 North American species, cutting training loss from 202.64 to 68.43 over three epochs. BirdNET classifies 3-second, 48 kHz audio into 6,522 species with sub-two-second CPU inference and no GPU required.

by read5 min views2 publishedOct 7, 2026

This is a submission for the Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass BirdBuddy is an offline bird call identifier. You're on a trail with no signal, you hear a bird, you record three seconds, and BirdBuddy tells you what it is — with everything running locally on the device in your pocket.

The screen is the shortest part of the experience. Open the app, tap once, put the phone away. The identification finishes in about two seconds. No cloud round-trip, no account, no upload. Just you, the bird, and a 52 MB model that never leaves your machine.

Who it's for: hikers, birders, and anyone who's ever heard something beautiful in the woods and had no idea what it was.

Live demo: https://birdbuddy.onrender.com Note: hosted on Render's free tier (512 MB RAM). The first request after idle may take 30–60 seconds to wake up. See the video below for the real, snappy experience.

https://www.youtube.com/watch?v=7b3iXirhnTQ What you're seeing:

GitHub: https://github.com/vinksgoyal/birdbuddy Repo layout:

src/inference.py — BirdNET ONNX inference (audio → species)src/voice.py — Piper TTS (text → speech, offline)app.py — Gradio UIscripts/tinker_finetune.py — Tinker fine-tune of Qwen3.5-4B for bird expertiseDockerfile + render.yaml — one-click deploymentdocs/ — Copilot usage screenshots BirdBuddy is built entirely on open-source AI. Three open pieces make it work.

BirdNET is a bioacoustics model released by the Cornell Lab of Ornithology and Chemnitz University. I used the FP32 truncated-DFT ONNX variant (~52 MB) running through ONNX Runtime on CPU. It classifies 3-second, 48 kHz audio into 6,522 species.

The model runs on a laptop, a phone, or a Raspberry Pi — no GPU required. Inference latency is under two seconds on a modern laptop CPU.

Piper is an MIT-licensed neural TTS that runs on CPU. I used the en_US-lessac-medium voice (~63 MB). When BirdNET returns its top match, Piper speaks it out loud: "I heard a Northern Cardinal."

No ElevenLabs, no Google Cloud TTS — the voice is a local ONNX model, just like the classifier.

The audio model gives us a species ID. To make BirdBuddy useful and not just a label, I fine-tuned a small LLM to act as a bird expert — field notes, habitat, how to identify by sound.

I built a 165-example chat dataset covering 60 common North American species, then fine-tuned Qwen3.5-4B with a LoRA adapter (r=16) using Thinking Machines' Tinker.

Training loss over 3 epochs:

Epoch Avg loss
1 202.64
2 95.09
3 68.43

Baseline output on a held-out prompt (before fine-tune): "I don't know how to use the API. ... Give me a field note for American Kestrel..."

Fine-tuned output (after training): "American Kestrel (Falco sparverius) is a small falcon with brown upperparts and blue wings. Its shape, plumage, and behavior are useful clues when recording a field observation..."

That's the improvement the Tinker category asks for: a specific task, a clear baseline, a measurable gain.

Copilot was the workbench for the entire project:

I used Backboard's R-CLI as a coding agent inside the repo for dataset curation and pipeline review — it summarized the inference pipeline, flagged missing error handling, and produced scripts/agent_review.md.

Docker + Render. The Dockerfile downloads both models (BirdNET and Piper) at build time so the container is fully self-contained. No API keys, no external services at runtime.

This project only exists because of open-source AI. Every design decision traces back to it.

The model is open-weight. BirdNET was released specifically for bioacoustics research. A closed API would have cost per call, required internet, and — critically — wouldn't have let me swap models. During development I tried three different classifiers. With a closed API, I'd be stuck with whatever the vendor offered.

The training data is open. The species library is a community-curated dataset of bird recordings. Thousands of birders contributed. A closed model trained on proprietary data would be a black box. I can inspect exactly what my model learned.

The runtime is open. ONNX Runtime means BirdBuddy runs on Android, iOS, Linux, macOS — even a Raspberry Pi. A closed inference service would lock me into one platform and one pricing tier.

The voice is open. Piper is MIT-licensed. Its voice models are open-weight. The TTS runs on the same laptop as the classifier.

The cost is zero. No API keys. No usage limits. No "your free tier has expired." A kid in a rural area with a hand-me-down phone can use this. A researcher in a country with unreliable internet can use this.

The agent harness is open. Backboard's R-CLI let me automate the boring parts — dataset cleaning, pipeline review — without locking my workflow into a proprietary IDE. I could inspect every step the agent took.

The fine-tune is open. Tinker let me fine-tune Qwen3.5-4B to a specific task and prove, with numbers, that the fine-tune beat the baseline. A closed model endpoint would not have given me the loss curve, the LoRA weights, or the freedom to swap base models.

The closed alternative would be: a subscription app that uploads your audio to a server, identifies birds you could have identified yourself, and charges $9.99/month. That's not innovation — that's rent-seeking.

This is innovation: a 52 MB classifier, a 63 MB voice model, a fine-tuned LoRA — all running on your device, in the middle of nowhere, for free, forever.

Backboard R-CLI session: sess_3f07cb2e — model openai/gpt-5.5, profile coding The session covered:

requirements.txt scripts/agent_review.md Entering:

Built with open-source AI, for the birds, and for everyone who'd rather be outside.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @birdbuddy 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/birdbuddy-an-offline…] indexed:0 read:5min 2026-10-07 · —