Cracking the Code: How SPARK Enhances AI Reasoning
SPARK, a new approach to diagnosing hidden-state failures in AI reasoning, boosted accuracy on the MATH-500 benchmark for Qwen3-4B from 82.0% to 84.6% and for Qwen3-8B from 82.4% to 85.6%, according t…