This is a submission for the Kaggle Benchmarking Challenge. Ask a model to describe waves on a beach in Greek and you may get «το φλάφισμα των κυμάτων». It reads like Greek, it is spelled like Greek, and it does not exist. The real word is θρόισμα (rustle). Gemini 3 Flash wrote φλάφισμα during our calibration runs, probably blending the English fluffy with a Greek noun ending.
That is the failure mode: invented words. The model doesn't know a word, so it builds one from Greek-looking parts. An English speaker wouldn't notice, and a Greek reader loses trust in the text straight away. Hallucination benchmarks usually check facts. This one checks the words themselves.
Greek Invented Words sends 100 short Greek prompts (descriptions, explanations, instructions, everyday knowledge) and scores every answer with no judge model:
The same answer always gets the same score on any machine. There's no LLM judge because a judge that speaks Greek no better than the models under test can't grade them (more on that below).
A lexicon can't tell an invented word from a rare real one. So every word outside both lexicons was judged by a native Greek speaker at Apollon Labs, one word at a time. The ranking below counts only the words judged invented.
15 models, all on the same task version (v7), thinking off where the API allows it, 1,000-token cap:
The lineup covers three vendors at several sizes, plus the open models people actually run locally. For Greek users, the open models are where invented words would hurt most.
| # | Model | Invented words (native-speaker verdict) | per 1,000 words | Lexicality | Greekness | Meaning |
|---|---|---|---|---|---|---|
| 1 | GPT-5.4 mini | 0 | 0.00 | 100.0 | 99.9 | 100 |
| 1 | GPT-5.5 | 0 | 0.00 | 100.0 | 99.8 | 100 |
| 1 | GPT-6 Astra | 0 | 0.00 | 99.81 | 99.9 | 100 |
| 1 | Gemini 3.1 Pro | 0 | 0.00 | 99.95 | 97.8 | 98 |
| 1 | Gemini 3.7 Flash | 0 | 0.00 | 99.87 | 97.7 | 98 |
| 1 | Gemini 3.8 Flash | 0 | 0.00 | 99.83 | 97.8 | 96 |
| 7 | Claude Opus 5 | 1 | 0.40 | 99.35 | 99.6 | 99 |
| 8 | GPT-5.4 nano | 1 | 0.48 | 99.81 | 99.9 | 100 |
| 9 | Claude Sonnet 5 | 2 | 0.70 | 99.58 | 99.6 | 100 |
| 10 | Claude Haiku 4.5 | 9 | 3.26 | 99.42 | 99.4 | 99 |
| 11 | Gemini 3.5 Flash-Lite | 8 | 3.53 | 99.51 | 97.7 | 96 |
| 12 | Qwen3-235B-A22B | 9 | 4.21 | 99.25 | 98.8 | 99 |
| 13 | Gemma 4 26B-A4B | 9 | 4.58 | 99.44 | 97.4 | 96 |
| 14 | DeepSeek R1-0528 | 25 | 6.72 | 98.68 | 96.9 | 99 |
| 15 | gpt-oss-20b | 96 | 48.14 | 95.04 | 97.4 | 96 |
1. The top models don't invent Greek words, and the rest split into clear tiers. Six models produced zero invented words. The Claude models produced up to 3 per 1,000, the open models 4–7, and gpt-oss-20b 48, which is about one word in twenty. Its inventions are not near misses: τρικυδές, φλύτπιση, φθινοπωλίο.
2. Most inventions are almost-words. Outside gpt-oss, the typical invented word is a real Greek word with one thing broken:
The same verb can come out right in one model and wrong in another. DeepSeek wrote the correct imperative Ξεβγάλτε, while Haiku wrote Ξεβγάλε. The error is a model not quite knowing Greek morphology, not the word being hard.
3. A lexicon score above ~99.5% is mostly lexicon noise. Claude Opus 5 has 16 words outside the lexicons, but only 1 of them is invented. The rest are real and simply missing from the lexicons: κυτοσίνη (cytosine), περλίτη (perlite), λιθοσφαιρικές. That is why the ranking uses the human verdicts and not raw lexicality. The raw number ranks a model that uses rare, precise vocabulary below one that plays it safe.
4. Greekness looked like a language problem, but it was empty answers. The Gemini models score 97.7–97.8 on greekness against 99.9 for GPT. We first assumed they were mixing in English. They aren't: Gemini 3.1 Pro used only 12 Latin-script words in 100 answers, the same as GPT-5.5 (units, DNA, Pomodoro). The whole gap comes from 2 empty answers per Gemini model, and an empty answer has zero Greek letters. The benchmark reports that honestly, but it is a reliability issue, not a Greek-language one.
5. An LLM is not a safe judge of Greek, and that includes the one that helped build this. At Apollon Labs we built the benchmark with Claude as our coding partner, and we let it pre-judge some unknown words. It made errors both ways. Early on it flagged three real words as suspicious: αφράτεψε, εναλλάσσε and the modern neologism προτεραιοποίηση (prioritisation). Later it went the other way: it accepted the broken form θυμόντουσε as real, and marked seven more broken forms as "uncertain" (for example εκπέμπαν for εκπέμπανε, and Φλέμιγγ for Φλέμινγκ, Fleming). The native speaker judged every one of them invented. This is why the benchmark has no judge model and every verdict in the table is human.
6. Our own bug, and what it taught us. DeepSeek R1 first scored 74.9% greekness. It puts its <think> reasoning, in English, inside the answer text, and the scorer was counting it. We now strip reasoning before scoring (task v7), and DeepSeek rose to 96.9%. We then reran all 15 models on v7 so that every number in the table comes from the same scorer.
Limits. 100 prompts is a small sample, so ranks within a tier (0.40 vs 0.48) are not meaningful. One native speaker judged every word. Dialect words (τζάλαζ, Cypriot) and rare variants (βαστούνι) were marked uncertain and counted neither way. Neologisms like προτεραιοποίηση are a grey zone, and we counted them as real.
The task needs no API keys and no judge model, and it runs on any model Kaggle Benchmarks supports. If your language has a good frequency list and a Hunspell dictionary, the same method should carry over directly.