Best local embedding model for job-posting search on a 16 GB M4 Mac? A developer's benchmark of three local embedding models on a 16 GB M4 Mac found Qwen3-Embedding-4B (Q4 GGUF) reached 91% natural and 71% paraphrased top-5 retrieval accuracy on 56 job-posting questions against a 400-posting pool, versus 77% and 32% for bge-small-en-v1.5 and 79% and 45% for bge-large-en-v1.5. The trade-off is speed: Qwen3-Embedding-4B ran at 3.3 chunks/s (roughly 12 hours for the corpus), bge-small at 366 chunks/s (about 6 minutes), and bge-large at 34 chunks/s, with Qwen3-Embedding-0.6B now testing at about 19 chunks/s. Gopi Krishna Reddy Katkuri is seeking an English-only, locally runnable embedding model under about 3 GB of memory that is free to use. Hi all, I’m building a local job-search assistant and I’m looking for the best embedding model for it. Everything runs on my own machine, so speed and memory matter as much as quality. The setup What I measured 56 postings, two questions each, pool of 400 postings; “top 5” = the posting the question was written from is in the top 5 | Model | Natural top 5 | Paraphrased top 5 | Speed | |---|---|---|---| | bge-small-en-v1.5 current | 77% | 32% | 366 chunks/s | | bge-large-en-v1.5 | 79% | 45% | 34 chunks/s | | Qwen3-Embedding-4B Q4 GGUF | 91% | 71% | 3.3 chunks/s | Qwen3-Embedding-4B is clearly the best, but it would take about 12 hours to embed my corpus, against about 6 minutes for bge-small. I’m testing Qwen3-Embedding-0.6B now about 19 chunks/s so far . What I’m looking for Constraints: English only, runs locally in under about 3 GB of memory, free to use. Thanks for any suggestions Gopi Krishna Reddy Katkuri