When LLMs don't know a Greek word, they make one up A developer built Greek Invented Words, a Kaggle benchmarking submission that sends 100 short Greek prompts to 15 language models and scores answers without an LLM judge, using native-speaker verdicts from Apollon Labs to count invented words. Six models produced zero invented Greek words, while gpt-oss-20b produced 96, roughly one word in twenty, and most other inventions were real Greek words with a single morphological error. The ranking relies on human judgments rather than raw lexicon scores because lexicality above about 99.5% is mostly lexicon noise. This is a submission for the Kaggle Benchmarking Challenge https://dev.to/challenges/kaggle-2026-09-23 . Ask a model to describe waves on a beach in Greek and you may get «το φλάφισμα των κυμάτων» . It reads like Greek, it is spelled like Greek, and it does not exist. The real word is θρόισμα rustle . Gemini 3 Flash wrote φλάφισμα during our calibration runs, probably blending the English fluffy with a Greek noun ending. That is the failure mode: invented words . The model doesn't know a word, so it builds one from Greek-looking parts. An English speaker wouldn't notice, and a Greek reader loses trust in the text straight away. Hallucination benchmarks usually check facts. This one checks the words themselves. Greek Invented Words sends 100 short Greek prompts descriptions, explanations, instructions, everyday knowledge and scores every answer with no judge model : The same answer always gets the same score on any machine. There's no LLM judge because a judge that speaks Greek no better than the models under test can't grade them more on that below . A lexicon can't tell an invented word from a rare real one. So every word outside both lexicons was judged by a native Greek speaker at Apollon Labs , one word at a time. The ranking below counts only the words judged invented. 15 models, all on the same task version v7 , thinking off where the API allows it, 1,000-token cap: The lineup covers three vendors at several sizes, plus the open models people actually run locally. For Greek users, the open models are where invented words would hurt most. | | Model | Invented words native-speaker verdict | per 1,000 words | Lexicality | Greekness | Meaning | |---|---|---|---|---|---|---| | 1 | GPT-5.4 mini | 0 | 0.00 | 100.0 | 99.9 | 100 | | 1 | GPT-5.5 | 0 | 0.00 | 100.0 | 99.8 | 100 | | 1 | GPT-6 Astra | 0 | 0.00 | 99.81 | 99.9 | 100 | | 1 | Gemini 3.1 Pro | 0 | 0.00 | 99.95 | 97.8 | 98 | | 1 | Gemini 3.7 Flash | 0 | 0.00 | 99.87 | 97.7 | 98 | | 1 | Gemini 3.8 Flash | 0 | 0.00 | 99.83 | 97.8 | 96 | | 7 | Claude Opus 5 | 1 | 0.40 | 99.35 | 99.6 | 99 | | 8 | GPT-5.4 nano | 1 | 0.48 | 99.81 | 99.9 | 100 | | 9 | Claude Sonnet 5 | 2 | 0.70 | 99.58 | 99.6 | 100 | | 10 | Claude Haiku 4.5 | 9 | 3.26 | 99.42 | 99.4 | 99 | | 11 | Gemini 3.5 Flash-Lite | 8 | 3.53 | 99.51 | 97.7 | 96 | | 12 | Qwen3-235B-A22B | 9 | 4.21 | 99.25 | 98.8 | 99 | | 13 | Gemma 4 26B-A4B | 9 | 4.58 | 99.44 | 97.4 | 96 | | 14 | DeepSeek R1-0528 | 25 | 6.72 | 98.68 | 96.9 | 99 | | 15 | gpt-oss-20b | 96 | 48.14 | 95.04 | 97.4 | 96 | 1. The top models don't invent Greek words, and the rest split into clear tiers. Six models produced zero invented words. The Claude models produced up to 3 per 1,000, the open models 4–7, and gpt-oss-20b 48, which is about one word in twenty . Its inventions are not near misses: τρικυδές , φλύτπιση , φθινοπωλίο . 2. Most inventions are almost-words. Outside gpt-oss, the typical invented word is a real Greek word with one thing broken: The same verb can come out right in one model and wrong in another. DeepSeek wrote the correct imperative Ξεβγάλτε , while Haiku wrote Ξεβγάλε . The error is a model not quite knowing Greek morphology, not the word being hard. 3. A lexicon score above ~99.5% is mostly lexicon noise. Claude Opus 5 has 16 words outside the lexicons, but only 1 of them is invented. The rest are real and simply missing from the lexicons: κυτοσίνη cytosine , περλίτη perlite , λιθοσφαιρικές . That is why the ranking uses the human verdicts and not raw lexicality. The raw number ranks a model that uses rare, precise vocabulary below one that plays it safe. 4. Greekness looked like a language problem, but it was empty answers. The Gemini models score 97.7–97.8 on greekness against 99.9 for GPT. We first assumed they were mixing in English. They aren't: Gemini 3.1 Pro used only 12 Latin-script words in 100 answers, the same as GPT-5.5 units, DNA , Pomodoro . The whole gap comes from 2 empty answers per Gemini model , and an empty answer has zero Greek letters. The benchmark reports that honestly, but it is a reliability issue, not a Greek-language one. 5. An LLM is not a safe judge of Greek, and that includes the one that helped build this. At Apollon Labs we built the benchmark with Claude as our coding partner, and we let it pre-judge some unknown words. It made errors both ways. Early on it flagged three real words as suspicious: αφράτεψε , εναλλάσσε and the modern neologism προτεραιοποίηση prioritisation . Later it went the other way: it accepted the broken form θυμόντουσε as real, and marked seven more broken forms as "uncertain" for example εκπέμπαν for εκπέμπανε , and Φλέμιγγ for Φλέμινγκ , Fleming . The native speaker judged every one of them invented. This is why the benchmark has no judge model and every verdict in the table is human. 6. Our own bug, and what it taught us. DeepSeek R1 first scored 74.9% greekness. It puts its