Kimi K3 with First Tree Beats GPT 5.6 Sol on a Real Engineering Task Kimi K3, developed by Moonshot AI, outperformed GPT-5.6 SOL on a real engineering task by using a context tree, according to a post on X by @Kimi_Moonshot. The post highlights that publishing pull requests (PRs) for a real issue makes the result more useful than a model-only benchmark, though it notes the remaining trust boundary is the grader, questioning whether a human independently reviewed Claude's scoring against the frozen rubric. Kimi K3 w. context tree beats GPT5.6 SOL - awesome work @Kimi Moonshot https://x.com/Kimi Moonshot - Using one real issue and publishing the PRs makes this more useful than a model-only benchmark. The remaining trust boundary is the grader. Did a human independently review Claude's scoring against the frozen rubric?