Last test Sep 25 · next test Sep 26
Every model takes the same test every day and is only ever compared with its own first week. Then it's your turn: say how it feels, and see if the numbers back you up.
Get an alert when a model gets nerfed #
One email when a verdict changes. Nothing else.
GPT-6 Astragpt-6-astra · low Benchmark
93% Too soon 1 run so far Community
– Too few votes Not enough data on either side yet.
What do you think?
Grok 4.7grok-4.7 · low Benchmark
91% Too soon 1 run so far Community
– Too few votes Not enough data on either side yet.
What do you think?
Muse Spark 1.3muse-spark-1.3 · low Benchmark
89% Too soon 1 run so far Community
– Too few votes Not enough data on either side yet.
What do you think?
Claude Opus 5.5claude-opus-5.5 · low Benchmark
85% Too soon 1 run so far Community
– Too few votes Not enough data on either side yet.
What do you think?