Hi all,
I’m building a local job-search assistant and I’m looking for the best
embedding model for it. Everything runs on my own machine, so speed and
memory matter as much as quality.
The setup
What I measured (56 postings, two questions each, pool of 400 postings;
“top 5” = the posting the question was written from is in the top 5)
| Model | Natural top 5 | Paraphrased top 5 | Speed |
|---|---|---|---|
| bge-small-en-v1.5 (current) | 77% | 32% | 366 chunks/s |
| bge-large-en-v1.5 | 79% | 45% | 34 chunks/s |
| Qwen3-Embedding-4B (Q4 GGUF) | 91% | 71% | 3.3 chunks/s |
Qwen3-Embedding-4B is clearly the best, but it would take about 12 hours
to embed my corpus, against about 6 minutes for bge-small. I’m testing
Qwen3-Embedding-0.6B now (about 19 chunks/s so far). What I’m looking for
Constraints: English only, runs locally in under about 3 GB of memory,
free to use.
Thanks for any suggestions!
Gopi Krishna Reddy Katkuri