Kimi K3 w. context tree beats GPT5.6 SOL
- awesome work @Kimi_Moonshot - Using one real issue and publishing the PRs makes this more useful than a model-only benchmark. The remaining trust boundary is the grader. Did a human independently review Claude's scoring against the frozen rubric?
source & further reading
twitter.com — original article