{"slug": "when-llms-don-t-know-a-greek-word-they-make-one-up", "title": "When LLMs don't know a Greek word, they make one up", "summary": "A developer built Greek Invented Words, a Kaggle benchmarking submission that sends 100 short Greek prompts to 15 language models and scores answers without an LLM judge, using native-speaker verdicts from Apollon Labs to count invented words. Six models produced zero invented Greek words, while gpt-oss-20b produced 96, roughly one word in twenty, and most other inventions were real Greek words with a single morphological error. The ranking relies on human judgments rather than raw lexicon scores because lexicality above about 99.5% is mostly lexicon noise.", "body_md": "*This is a submission for the [Kaggle Benchmarking Challenge](https://dev.to/challenges/kaggle-2026-09-23).*\n\nAsk a model to describe waves on a beach in Greek and you may get *«το φλάφισμα των κυμάτων»*. It reads like Greek, it is spelled like Greek, and it does not exist. The real word is *θρόισμα* (rustle). Gemini 3 Flash wrote *φλάφισμα* during our calibration runs, probably blending the English *fluffy* with a Greek noun ending.\n\nThat is the failure mode: **invented words**. The model doesn't know a word, so it builds one from Greek-looking parts. An English speaker wouldn't notice, and a Greek reader loses trust in the text straight away. Hallucination benchmarks usually check facts. This one checks the words themselves.\n\n**Greek Invented Words** sends 100 short Greek prompts (descriptions, explanations, instructions, everyday knowledge) and scores every answer with **no judge model**:\n\nThe same answer always gets the same score on any machine. There's no LLM judge because a judge that speaks Greek no better than the models under test can't grade them (more on that below).\n\nA lexicon can't tell an invented word from a rare real one. So every word outside both lexicons was **judged by a native Greek speaker at Apollon Labs**, one word at a time. The ranking below counts only the words judged invented.\n\n15 models, all on the same task version (v7), thinking off where the API allows it, 1,000-token cap:\n\nThe lineup covers three vendors at several sizes, plus the open models people actually run locally. For Greek users, the open models are where invented words would hurt most.\n\n| # | Model | Invented words (native-speaker verdict) | per 1,000 words | Lexicality | Greekness | Meaning | \n|---|---|---|---|---|---|---|\n| 1 | GPT-5.4 mini | 0 | 0.00 | 100.0 | 99.9 | 100 | \n| 1 | GPT-5.5 | 0 | 0.00 | 100.0 | 99.8 | 100 | \n| 1 | GPT-6 Astra | 0 | 0.00 | 99.81 | 99.9 | 100 | \n| 1 | Gemini 3.1 Pro | 0 | 0.00 | 99.95 | 97.8 | 98 | \n| 1 | Gemini 3.7 Flash | 0 | 0.00 | 99.87 | 97.7 | 98 | \n| 1 | Gemini 3.8 Flash | 0 | 0.00 | 99.83 | 97.8 | 96 | \n| 7 | Claude Opus 5 | 1 | 0.40 | 99.35 | 99.6 | 99 | \n| 8 | GPT-5.4 nano | 1 | 0.48 | 99.81 | 99.9 | 100 | \n| 9 | Claude Sonnet 5 | 2 | 0.70 | 99.58 | 99.6 | 100 | \n| 10 | Claude Haiku 4.5 | 9 | 3.26 | 99.42 | 99.4 | 99 | \n| 11 | Gemini 3.5 Flash-Lite | 8 | 3.53 | 99.51 | 97.7 | 96 | \n| 12 | Qwen3-235B-A22B | 9 | 4.21 | 99.25 | 98.8 | 99 | \n| 13 | Gemma 4 26B-A4B | 9 | 4.58 | 99.44 | 97.4 | 96 | \n| 14 | DeepSeek R1-0528 | 25 | 6.72 | 98.68 | 96.9 | 99 | \n| 15 | gpt-oss-20b | 96 | 48.14 | 95.04 | 97.4 | 96 | \n\n**1. The top models don't invent Greek words, and the rest split into clear tiers.** Six models produced zero invented words. The Claude models produced up to 3 per 1,000, the open models 4–7, and gpt-oss-20b 48, which is **about one word in twenty**. Its inventions are not near misses: *τρικυδές*, *φλύτπιση*, *φθινοπωλίο*.\n\n**2. Most inventions are almost-words.** Outside gpt-oss, the typical invented word is a real Greek word with one thing broken:\n\nThe same verb can come out right in one model and wrong in another. DeepSeek wrote the correct imperative *Ξεβγάλτε*, while Haiku wrote *Ξεβγάλε*. The error is a model not quite knowing Greek morphology, not the word being hard.\n\n**3. A lexicon score above ~99.5% is mostly lexicon noise.** Claude Opus 5 has 16 words outside the lexicons, but only 1 of them is invented. The rest are real and simply missing from the lexicons: *κυτοσίνη* (cytosine), *περλίτη* (perlite), *λιθοσφαιρικές*. That is why the ranking uses the human verdicts and not raw lexicality. The raw number ranks a model that uses rare, precise vocabulary below one that plays it safe.\n\n**4. Greekness looked like a language problem, but it was empty answers.** The Gemini models score 97.7–97.8 on greekness against 99.9 for GPT. We first assumed they were mixing in English. They aren't: Gemini 3.1 Pro used only 12 Latin-script words in 100 answers, the same as GPT-5.5 (units, *DNA*, *Pomodoro*). The whole gap comes from **2 empty answers per Gemini model**, and an empty answer has zero Greek letters. The benchmark reports that honestly, but it is a reliability issue, not a Greek-language one.\n\n**5. An LLM is not a safe judge of Greek, and that includes the one that helped build this.** At Apollon Labs we built the benchmark with Claude as our coding partner, and we let it pre-judge some unknown words. It made errors both ways. Early on it flagged three real words as suspicious: *αφράτεψε*, *εναλλάσσε* and the modern neologism *προτεραιοποίηση* (prioritisation). Later it went the other way: it accepted the broken form *θυμόντουσε* as real, and marked seven more broken forms as \"uncertain\" (for example *εκπέμπαν* for *εκπέμπανε*, and *Φλέμιγγ* for *Φλέμινγκ*, Fleming). The native speaker judged every one of them invented. This is why the benchmark has no judge model and every verdict in the table is human.\n\n**6. Our own bug, and what it taught us.** DeepSeek R1 first scored 74.9% greekness. It puts its `<think>` reasoning, in English, inside the answer text, and the scorer was counting it. We now strip reasoning before scoring (task v7), and DeepSeek rose to 96.9%. We then reran all 15 models on v7 so that every number in the table comes from the same scorer.\n\n**Limits.** 100 prompts is a small sample, so ranks within a tier (0.40 vs 0.48) are not meaningful. One native speaker judged every word. Dialect words (*τζάλαζ*, Cypriot) and rare variants (*βαστούνι*) were marked uncertain and counted neither way. Neologisms like *προτεραιοποίηση* are a grey zone, and we counted them as real.\n\nThe task needs no API keys and no judge model, and it runs on any model Kaggle Benchmarks supports. If your language has a good frequency list and a Hunspell dictionary, the same method should carry over directly.", "url": "https://wpnews.pro/news/when-llms-don-t-know-a-greek-word-they-make-one-up", "canonical_source": "https://dev.to/apollonlabsai/when-llms-dont-know-a-greek-word-they-make-one-up-4ef2", "published_at": "2026-09-30 11:34:13+00:00", "updated_at": "2026-09-30 11:47:39.226321+00:00", "lang": "en", "topics": ["large-language-models", "natural-language-processing", "ai-research", "ai-tools"], "entities": ["Kaggle", "Apollon Labs", "GPT-5.4 mini", "GPT-5.5", "Gemini 3.1 Pro", "Claude Opus 5", "DeepSeek R1-0528", "gpt-oss-20b"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/when-llms-don-t-know-a-greek-word-they-make-one-up", "markdown": "https://wpnews.pro/news/when-llms-don-t-know-a-greek-word-they-make-one-up.md", "text": "https://wpnews.pro/news/when-llms-don-t-know-a-greek-word-they-make-one-up.txt", "jsonld": "https://wpnews.pro/news/when-llms-don-t-know-a-greek-word-they-make-one-up.jsonld"}}