{"slug": "best-local-embedding-model-for-job-posting-search-on-a-16-gb-m4-mac", "title": "Best local embedding model for job-posting search on a 16 GB M4 Mac?", "summary": "A developer's benchmark of three local embedding models on a 16 GB M4 Mac found Qwen3-Embedding-4B (Q4 GGUF) reached 91% natural and 71% paraphrased top-5 retrieval accuracy on 56 job-posting questions against a 400-posting pool, versus 77% and 32% for bge-small-en-v1.5 and 79% and 45% for bge-large-en-v1.5. The trade-off is speed: Qwen3-Embedding-4B ran at 3.3 chunks/s (roughly 12 hours for the corpus), bge-small at 366 chunks/s (about 6 minutes), and bge-large at 34 chunks/s, with Qwen3-Embedding-0.6B now testing at about 19 chunks/s. Gopi Krishna Reddy Katkuri is seeking an English-only, locally runnable embedding model under about 3 GB of memory that is free to use.", "body_md": "Hi all,\n\nI’m building a local job-search assistant and I’m looking for the best\n\nembedding model for it. Everything runs on my own machine, so speed and\n\nmemory matter as much as quality.\n\nThe setup\n\nWhat I measured (56 postings, two questions each, pool of 400 postings;\n\n“top 5” = the posting the question was written from is in the top 5)\n\n| Model | Natural top 5 | Paraphrased top 5 | Speed | \n|---|---|---|---|\n| bge-small-en-v1.5 (current) | 77% | 32% | 366 chunks/s | \n| bge-large-en-v1.5 | 79% | 45% | 34 chunks/s | \n| Qwen3-Embedding-4B (Q4 GGUF) | 91% | 71% | 3.3 chunks/s | \n\nQwen3-Embedding-4B is clearly the best, but it would take about 12 hours\n\nto embed my corpus, against about 6 minutes for bge-small. I’m testing\n\nQwen3-Embedding-0.6B now (about 19 chunks/s so far).\n\nWhat I’m looking for\n\nConstraints: English only, runs locally in under about 3 GB of memory,\n\nfree to use.\n\nThanks for any suggestions!\n\nGopi Krishna Reddy Katkuri", "url": "https://wpnews.pro/news/best-local-embedding-model-for-job-posting-search-on-a-16-gb-m4-mac", "canonical_source": "https://discuss.huggingface.co/t/best-local-embedding-model-for-job-posting-search-on-a-16-gb-m4-mac/183201#post_1", "published_at": "2026-10-06 00:13:01+00:00", "updated_at": "2026-10-06 00:17:01.315230+00:00", "lang": "en", "topics": ["ai-search", "machine-learning", "natural-language-processing", "ai-tools"], "entities": ["Qwen3-Embedding-4B", "Qwen3-Embedding-0.6B", "bge-small-en-v1.5", "bge-large-en-v1.5", "Gopi Krishna Reddy Katkuri", "Apple M4 Mac"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/best-local-embedding-model-for-job-posting-search-on-a-16-gb-m4-mac", "markdown": "https://wpnews.pro/news/best-local-embedding-model-for-job-posting-search-on-a-16-gb-m4-mac.md", "text": "https://wpnews.pro/news/best-local-embedding-model-for-job-posting-search-on-a-16-gb-m4-mac.txt", "jsonld": "https://wpnews.pro/news/best-local-embedding-model-for-job-posting-search-on-a-16-gb-m4-mac.jsonld"}}