cd /news/large-language-models/when-llms-don-t-know-a-greek-word-th… · home › topics › large-language-models › article
[ARTICLE · art-142460] src=dev.to ↗ pub= topic=large-language-models verified=true sentiment=· neutral

When LLMs don't know a Greek word, they make one up

A developer built Greek Invented Words, a Kaggle benchmarking submission that sends 100 short Greek prompts to 15 language models and scores answers without an LLM judge, using native-speaker verdicts from Apollon Labs to count invented words. Six models produced zero invented Greek words, while gpt-oss-20b produced 96, roughly one word in twenty, and most other inventions were real Greek words with a single morphological error. The ranking relies on human judgments rather than raw lexicon scores because lexicality above about 99.5% is mostly lexicon noise.

read5 min views1 publishedSep 30, 2026

This is a submission for the Kaggle Benchmarking Challenge. Ask a model to describe waves on a beach in Greek and you may get «το φλάφισμα των κυμάτων». It reads like Greek, it is spelled like Greek, and it does not exist. The real word is θρόισμα (rustle). Gemini 3 Flash wrote φλάφισμα during our calibration runs, probably blending the English fluffy with a Greek noun ending.

That is the failure mode: invented words. The model doesn't know a word, so it builds one from Greek-looking parts. An English speaker wouldn't notice, and a Greek reader loses trust in the text straight away. Hallucination benchmarks usually check facts. This one checks the words themselves.

Greek Invented Words sends 100 short Greek prompts (descriptions, explanations, instructions, everyday knowledge) and scores every answer with no judge model:

The same answer always gets the same score on any machine. There's no LLM judge because a judge that speaks Greek no better than the models under test can't grade them (more on that below).

A lexicon can't tell an invented word from a rare real one. So every word outside both lexicons was judged by a native Greek speaker at Apollon Labs, one word at a time. The ranking below counts only the words judged invented.

15 models, all on the same task version (v7), thinking off where the API allows it, 1,000-token cap:

The lineup covers three vendors at several sizes, plus the open models people actually run locally. For Greek users, the open models are where invented words would hurt most.

# Model Invented words (native-speaker verdict) per 1,000 words Lexicality Greekness Meaning
1 GPT-5.4 mini 0 0.00 100.0 99.9 100
1 GPT-5.5 0 0.00 100.0 99.8 100
1 GPT-6 Astra 0 0.00 99.81 99.9 100
1 Gemini 3.1 Pro 0 0.00 99.95 97.8 98
1 Gemini 3.7 Flash 0 0.00 99.87 97.7 98
1 Gemini 3.8 Flash 0 0.00 99.83 97.8 96
7 Claude Opus 5 1 0.40 99.35 99.6 99
8 GPT-5.4 nano 1 0.48 99.81 99.9 100
9 Claude Sonnet 5 2 0.70 99.58 99.6 100
10 Claude Haiku 4.5 9 3.26 99.42 99.4 99
11 Gemini 3.5 Flash-Lite 8 3.53 99.51 97.7 96
12 Qwen3-235B-A22B 9 4.21 99.25 98.8 99
13 Gemma 4 26B-A4B 9 4.58 99.44 97.4 96
14 DeepSeek R1-0528 25 6.72 98.68 96.9 99
15 gpt-oss-20b 96 48.14 95.04 97.4 96

1. The top models don't invent Greek words, and the rest split into clear tiers. Six models produced zero invented words. The Claude models produced up to 3 per 1,000, the open models 4–7, and gpt-oss-20b 48, which is about one word in twenty. Its inventions are not near misses: τρικυδές, φλύτπιση, φθινοπωλίο.

2. Most inventions are almost-words. Outside gpt-oss, the typical invented word is a real Greek word with one thing broken:

The same verb can come out right in one model and wrong in another. DeepSeek wrote the correct imperative Ξεβγάλτε, while Haiku wrote Ξεβγάλε. The error is a model not quite knowing Greek morphology, not the word being hard.

3. A lexicon score above ~99.5% is mostly lexicon noise. Claude Opus 5 has 16 words outside the lexicons, but only 1 of them is invented. The rest are real and simply missing from the lexicons: κυτοσίνη (cytosine), περλίτη (perlite), λιθοσφαιρικές. That is why the ranking uses the human verdicts and not raw lexicality. The raw number ranks a model that uses rare, precise vocabulary below one that plays it safe.

4. Greekness looked like a language problem, but it was empty answers. The Gemini models score 97.7–97.8 on greekness against 99.9 for GPT. We first assumed they were mixing in English. They aren't: Gemini 3.1 Pro used only 12 Latin-script words in 100 answers, the same as GPT-5.5 (units, DNA, Pomodoro). The whole gap comes from 2 empty answers per Gemini model, and an empty answer has zero Greek letters. The benchmark reports that honestly, but it is a reliability issue, not a Greek-language one.

5. An LLM is not a safe judge of Greek, and that includes the one that helped build this. At Apollon Labs we built the benchmark with Claude as our coding partner, and we let it pre-judge some unknown words. It made errors both ways. Early on it flagged three real words as suspicious: αφράτεψε, εναλλάσσε and the modern neologism προτεραιοποίηση (prioritisation). Later it went the other way: it accepted the broken form θυμόντουσε as real, and marked seven more broken forms as "uncertain" (for example εκπέμπαν for εκπέμπανε, and Φλέμιγγ for Φλέμινγκ, Fleming). The native speaker judged every one of them invented. This is why the benchmark has no judge model and every verdict in the table is human.

6. Our own bug, and what it taught us. DeepSeek R1 first scored 74.9% greekness. It puts its <think> reasoning, in English, inside the answer text, and the scorer was counting it. We now strip reasoning before scoring (task v7), and DeepSeek rose to 96.9%. We then reran all 15 models on v7 so that every number in the table comes from the same scorer.

Limits. 100 prompts is a small sample, so ranks within a tier (0.40 vs 0.48) are not meaningful. One native speaker judged every word. Dialect words (τζάλαζ, Cypriot) and rare variants (βαστούνι) were marked uncertain and counted neither way. Neologisms like προτεραιοποίηση are a grey zone, and we counted them as real.

The task needs no API keys and no judge model, and it runs on any model Kaggle Benchmarks supports. If your language has a good frequency list and a Hunspell dictionary, the same method should carry over directly.

── more in #large-language-models 4 stories · sorted by recency
── more on @kaggle 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/when-llms-don-t-know…] indexed:0 read:5min 2026-09-30 · —