This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend
Somewhere on Sam's laptop, there's a riff Sam swears is the best thing they've written all year.
It's in a file called recording_027_final_v2.wav. Or maybe heavyyy_enough.wav. Or idea_actually_good.m4a. He has 100s of these, and not many filenames say what's inside.
Here's the cruel part: you don't remember a riff by its name. You remember how it felt. Ask Sam about the lost one and you get something like:
"heavy, Drop C, around 145, I think it was a chorus idea"
That's a perfectly good search query. It's also one that no folder on earth can answer.
So I built something that can.
RiffSalad (yes, that's what the recordings folder looks like) is a local AI vault for guitarists. You drop riffs in, or record them straight from the browser, and it listens: tempo, key, energy, how busy the playing is, even a MIDI transcription. It writes a short description and a few tags for each one. Then you search the way you actually think:
heavy drop C riff around 140 BPM`` clean mellow fingerpicked thing in E minor``something like my last blues riff
You can talk to your riffs, too. Record a voice note ("this was the bridge, needs a key change") and it's transcribed on your own machine and becomes searchable. There's also a waveform player, inline editing, and a "find similar riffs" button.
Sam is an intermediate guitarist with creative ideas popping off his head all the time. Last time I saw him, he was in front his laptop humming to the tune, me being the overachieving automator.I wanted to build the thing that lets them stop scrolling and just type what they remember. Mostly because i wanted him to stop humming. lol!
Finallyyy, I can have something organized in my life and easy to find.
Your guitar ideas, searchable.
A local-first AI vault for guitarists. Record a riff, import the file β and find it later by describing it in plain English.
Every guitarist knows this feeling: you recorded a cool riff a few months ago, but you can't find it because you saved it as recording_027_final_v2.wav. You remember it was a heavy, slow, drop-tuned thing around 80 BPM β but searching a folder of audio files for that is impossible.
RiffSalad solves this. Import your recordings and search them like you'd describe them to a bandmate.
"heavy drop C riffs around 140 BPM"
"that mellow fingerpicking thing in E minor"
"fast aggressive shredding, lots of notes"
ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β Import WAV / MP3 / M4A β
β β β
β βΌ β
β librosa analyzes: BPM, Key, Energy, Note Density, Duration β
β β β
β ββββΊ Ollama LLM
β¦ Everything runs locally and needs no API keys:
git clone https://github.com/jesi318/Riff-Salad
cd riffsalad
cp .env.example .env
docker compose up
Then open localhost:3000. The first run pulls the models. After that, it works with no internet.
One decision shaped everything: the small model does as little as possible.
I wanted this to run on a normal machine, so the language model is a 1.5B-parameter qwen2.5 served by Ollama. That's tiny. A model that size is good at turning "heavy drop C riff around 140" into structured filters, and bad at deciding whether a riff actually is heavy. So I stopped asking it to decide.
The rest of the stack is deliberately plain:
librosa measures BPM (harmonic/percussive separation, then tempo tracking), key (matching against tonal profiles), energy, and note density. A Basic Pitch-style ONNX model extracts MIDI.faster-whisper (the tiny model) transcribes voice notes locally.nomic-embed-text embeddings via Ollama, stored alongside a SQLite database.
Search runs in four stages, and the chat model only gets the first one:
The rule I'm proudest of: the tagger isn't allowed to say "heavy," "metal," or "doom" unless the numbers back it up. Small models love to over-claim, and "heavy" is the easiest word in music to overuse. So the measured facts get the first vote, and the model just gets to phrase things nicely. "Find similar" follows the same philosophy, blending embedding similarity with plain feature distance (BPM, key, energy), so "sounds like this" is something you can sanity-check.
It isn't magic. Key and BPM detection are heuristics, and sometimes they're wrong, which is why both can be overridden inline.
A riff is a diary entry you can hum. It's half-finished, sometimes embarrassing, often unreleased. I didn't want Sam to upload their unfinished ideas to someone else's server just to find them again later. With open models running locally, they don't have to: the audio, the transcripts, and the database all live in a data/ folder on their own machine.
Open also gave me things a closed API wouldn't have:
LLM_MODEL, EMBEDDING_MODEL, and WHISPER_MODEL are environment variables, so swapping a model never means touching code. [Optional: what you tried and what changed.]
Sam still has 103 files called recording-something. They just don't need the filenames anymore.