18:18
2026-08-21
dev.to
large-language-models
Small local models: what actually holds up when you re-run the measurement
A developer testing Qwen3.5 4B via Ollama found that enabling chain-of-thought can return an empty response while the thinking field holds content, and that self-critique can degrade accuracy, citing โฆ