Small local models: what actually holds up when you re-run the measurement
A developer testing Qwen3.5 4B via Ollama found that enabling chain-of-thought can return an empty response while the thinking field holds content, and that self-critique can degrade accuracy, citing …