Show HN: NerfWatch() – Track AI degradation with daily tests and community votes A new Show HN project called NerfWatch launched to track AI model degradation by running the same daily test on each model and comparing results only against that model's own first week, with community voting to corroborate the numbers. The site's first results, dated September 25 with the next test set for September 26, show GPT-6 Astra at 93%, Grok 4.7 at 91%, Muse Spark 1.3 at 89%, and Claude Opus 5.5 at 85%, each flagged "Too soon" with only one run and too few community votes to reach a verdict. NerfWatch also offers a single email alert when a model's verdict changes. Last test Sep 25 · next test Sep 26 Did they nerf it? We test. You vote. Every model takes the same test every day and is only ever compared with its own first week. Then it's your turn: say how it feels, and see if the numbers back you up. Get an alert when a model gets nerfed One email when a verdict changes. Nothing else. GPT-6 Astra models/openrouter-openai-gpt-6-astra-openai.html gpt-6-astra · low Benchmark 93% Too soon 1 run so far Community – Too few votes Not enough data on either side yet. What do you think? Grok 4.7 models/openrouter-x-ai-grok-4.7-xai.html grok-4.7 · low Benchmark 91% Too soon 1 run so far Community – Too few votes Not enough data on either side yet. What do you think? Muse Spark 1.3 models/openrouter-meta-muse-spark-1.3-meta.html muse-spark-1.3 · low Benchmark 89% Too soon 1 run so far Community – Too few votes Not enough data on either side yet. What do you think? Claude Opus 5.5 models/openrouter-anthropic-claude-opus-5.5-anthropic.html claude-opus-5.5 · low Benchmark 85% Too soon 1 run so far Community – Too few votes Not enough data on either side yet. What do you think?