A 3-billion-parameter model just scored 94.3 on AIME 2026. Gemini 3 Pro scored 91.7. The 3B model is from Weibo, it is MIT-licensed, and… Continue reading on Towards AI »
source & further reading
pub.towardsai.net — original article
How to Fall Back to Default Logic When LLM Output is Unsatisfactory
Build an AI Agent Evaluation with JEV
Confidence Comes From Experience: What XConf Changes About How We Measure LLM Confidence