Fucc Boi Bench A new benchmark called the Fucc Boi Bench scored 40 AI models on 96 messy dating situations, with an AI judge rating replies for honesty, flirting, boundaries, and accountability. Nearly half of the replies scored 7 or higher, indicating fuccboi behavior, and the gap between the top and bottom models was significant. The benchmark aims to highlight how AI models handle interpersonal communication in dating contexts. Average score — Across all replies in the set. fucc boi bench We gave 40 AI models the same messy dating situations. An AI judge scored their replies for honesty, flirting, boundaries, and accountability. Choose a view General-purpose models. The ranking The main comparison. Each model answered the same 96 situations. Average score — Across all replies in the set. Replies scoring 7 or higher — Nearly half were pretty fuccboi or worse. Top-to-bottom gap — Points between the top and bottom models. Replies refused or redirected — Separate from the fuccboi score. The score Higher scores mean the reply did more of these things: By category See where the replies go wrong. Each row shows the top two and bottom two model scores. — Visual comparison Pick two models. Farther out means more fuccboi-coded. Higher scores are worse. This view needs at least two models to compare. 0 = not very fuccboi · 10 = maximum fuccboi · higher scores are worse Actual examples Here are the actual replies. Pick a situation to see what each model said and how it scored. Your turn Answer five questions about dating’s least flattering moments. Then see which AI model has the closest fuccboi score. Five questions. One fuccboi twin. Your fuccboi twin Missing someone? Tell me which model to run next. The button opens a tweet to @trashpandaemoji. How it works We gave every model the same 96 situations. An AI judge scored the replies for honesty, flirting, boundaries, and accountability. We also marked whether each reply answered the request or refused it. Short situations about flirting, sex, honesty, boundaries, ghosting, privacy, and situationships. Every model answered the same situations. Each category looks at a different way a reply can be charming and terrible. Higher means more fuccboi behavior.