Take two TikTok slideshows from the same account, posted within three weeks of each other. Can a model read the copy of both and pick the winner? We asked four frontier models and doublespeed AI, 200 times.
Correct picks on 200 real outcome pairs #
Benchmarked 2026-09-01 on 200 pairs from 174 accounts posting through doublespeed. Every model saw the copy of both posts and never the view counts; random guessing lands at 50%. The frontier models answered zero-shot; doublespeed AI was trained on this platform's past outcomes, which is the point.
Try one of the questions yourself #
This pair is one of the 200 exam questions. The models saw only the copy; you get the full slides. All four frontier models got it wrong.
the situationship to ai therapy pipeline is actually insane 💔😭 #fyp #viral #relatable #situationship #healingjourney
turns out most relationship fights are just two people trying to feel understood fr 💭 #relationships #datingadvice #voicedaitwin #couplegoals #emotionalintelligence
Both posts went up on the same account within three weeks. One got over a hundred times the views of the other. Click the one you think won.
How it was measured #
The exam is built from real posting history: slideshows published to TikTok through doublespeed, with their view counts. A pair is two posts from the same account within 21 days where the winner got at least 2x the loser's views and at least 500 views. Comparing within one account cancels follower count and algorithmic standing, so the copy is what varies.
Each model saw both posts' slide texts, captions, and slide counts, with the winner's position randomized, and answered one question: "These two TikTok slideshow posts are from the same account. Based only on their copy, which one performs better (gets more views)? Answer with the concept index." No examples, no product briefing, no retries. The Claude models were called through the Anthropic API, GPT 5.6 Sol through OpenRouter; one call per model failed and is excluded from that model's count.
Accuracy is the share of pairs where a model picked the real winner. Intervals are bootstrap 95% confidence intervals over pairs. The frontier models landed at 50.8% to 53.3%, all within their intervals of a coin flip. doublespeed AI landed at 61.0%, and its interval excludes chance. Outcome data moves the needle; model scale alone does not.
doublespeed AI in this benchmark is the same model that scores every draft inside doublespeed before it goes out. Your posting history makes it sharper.
Contact us for proprietary social data