{"slug": "kolibri-vs-claude-sonnet-5-5-a-german-llm-benchmark", "title": "Kolibri vs Claude Sonnet 5.5: A German LLM Benchmark", "summary": "Aleph Alpha's Kolibri lost all 136 blind comparisons against Claude Sonnet 5.5 in a German-language Bible deep-dive benchmark run by Dewfall developer on 3 October 2026, scoring 3.7 with German prompts and 4.2 with English prompts versus Sonnet 5.5's 8.7 on both. Blind judges GPT-5.5 and Gemini 3.1 Pro flagged 100 serious errors in Kolibri's German-prompt output and 92 in its English-prompt output, against 7 for Sonnet 5.5, and Kolibri parsed into all 7 required sections in only 16 of 17 German-prompt dives and 14 of 17 English-prompt dives. Kolibri required about 78 GB of GPU memory and self-hosting via Aleph Alpha's vLLM plugin at $0.008 per verse dive, versus $0.036 per verse for Sonnet 5.5's Batch API, so the developer kept German deep dives on Claude.", "body_md": "# Kolibri vs Claude Sonnet 5.5: A German LLM Benchmark\n\n**[Aleph Alpha](https://aleph-alpha.com/)’s [Kolibri](https://huggingface.co/Aleph-Alpha/Kolibri-1) lost all 136 blind comparisons against Claude Sonnet 5.5 when I had both write Bible deep dives in German for my app. So now, my German deep dives stay on Claude.** When I [explained how Kolibri works](https://tej.as/blog/aleph-alpha-kolibri), I ran its tokenizer over all of the German constitution, but I hadn’t run the model itself. Kolibri needs a data-center GPU and nobody hosts it. So I rented the GPUs myself and ran it on 3 October 2026, the same day it came out.\n\n[Dewfall](https://dewfall.app) is the Bible app I build. Every verse in it has a deep dive: what the verse means, the world it came from, the Hebrew or Greek, how it points to Jesus, prayer points and questions to journal on. For German, Dewfall uses the [Schlachter 1951](https://de.wikipedia.org/wiki/Schlachter-Bibel), which has 31,170 verses. Only 2,374 of them have a German deep dive as I write this (every chapter plus one key verse per chapter, written in a batch by Claude Sonnet 5). That leaves almost 29,000 German verses with nothing under them. A German model I can run on my own GPUs sounded like a great way to fill them for cheap. So before switching anything, I tested it.\n\n## The results at a glance\n\n| Model | German prompt | English prompt | \n|---|---|---|\n| Sonnet 5.5 | 8.7 | 8.7 | \n| Kolibri | 3.7 | 4.2 | \n\n|  | Kolibri, German prompt | Kolibri, English prompt | Sonnet 5.5, German prompt | Sonnet 5.5, English prompt | \n|---|---|---|---|---|\n| Blind judge score (1 to 10) | 3.7 | 4.2 | 8.7 | 8.7 | \n| Wins against Sonnet 5.5 (same prompt) | 0 of 68 | 0 of 68 |  |  | \n| Serious errors the judges flagged | 100 | 92 | 7 | 7 | \n| Dives that parsed into all 7 sections | 16 of 17 | 14 of 17 | 17 of 17 | 17 of 17 | \n| Grammar errors per 1,000 words ( [LanguageTool](https://languagetool.org/) ) | 0.36 | 0.68 | 0.24 | 0.29 | \n| Median tokens of reasoning per verse | 15,574 | 13,262 | (not separated) | (not separated) | \n| Cost per verse dive | $0.008 (self-hosted, 128 at once) |  | $0.036 (Batch API) |  | \n\nThe judges were [GPT-5.5](https://openai.com/) and [Gemini 3.1 Pro](https://deepmind.google/models/gemini/). Neither of them knew which model wrote what. Every “win” above is one judge picking one dive over the other, in both presentation orders. Of the 136 verdicts, 135 were marked “decisive”. Wild.\n\n## I had to rent the GPUs myself\n\nThe model is [on Hugging Face](https://huggingface.co/Aleph-Alpha/Kolibri-1) so I figured I could just call it there. I couldn’t. The model page has no inference provider, Hugging Face’s router answered `model_not_supported`, [OpenRouter](https://openrouter.ai/) doesn’t list it, and there’s no demo Space either. Kolibri needs about 78 GB of GPU memory and Aleph Alpha’s own [vLLM plugin](https://github.com/Aleph-Alpha/aleph-alpha-inference). So the only way to get a single token out of it on launch day was to serve it myself.\n\nI did that with a [Hugging Face Inference Endpoint](https://huggingface.co/docs/inference-endpoints/index) running Aleph Alpha’s container. It took more than I expected. My first token was read-only so I needed a fine-grained one that can create endpoints. Then the API said “Payment method required” even though I had [Apple Pay](https://www.apple.com/apple-pay/) on file. My account bills from a prepaid credit balance so I added credits. Then the single [NVIDIA H200](https://www.nvidia.com/en-us/data-center/h200/) I asked for ($5 an hour) sat on “Waiting for requested hardware to become available” for 14 minutes and never showed up. So I raced it against a second endpoint on 2 [NVIDIA RTX PRO 6000](https://www.nvidia.com/en-us/products/workstations/professional-desktop-gpus/rtx-pro-6000/) GPUs (96 GB each, $5.50 an hour), which went from created to serving in about 7 minutes. I deleted the H200. The 78 GB of weights loaded in 27 seconds, which is kinda ridiculous.\n\nHonestly, I think not also selling an API is a kinda dumb decision. Paying for API access and tokens is how you make a model company sustainable, the way [OpenAI](https://openai.com/) did it. Nobody cared about GPT-3 until there was a UI and an API anyone could use: the [GPT-3 API](https://openai.com/index/openai-api/) launched in private beta in June 2020 and kept a waitlist until [November 2021](https://www.hpcwire.com/aiwire/2021/11/18/openai-gtp-3-waiting-list-is-gone-as-gtp-3-is-fully-released-for-use/). It took [ChatGPT](https://openai.com/index/chatgpt/) in November 2022 and then the cheap [ChatGPT API](https://openai.com/index/introducing-chatgpt-and-whisper-apis/) in March 2023 for it to blow up. In my talk [You are an AI engineer](https://www.youtube.com/watch?v=AlsNzAoCLSA&t=385s) at JSHeroes 2024 I put it the way [swyx’s Rise of the AI Engineer](https://www.latent.space/p/ai-engineer) does: you don’t need linear algebra or Python to build with AI, “you talk to an API”. On launch day there’s nothing to talk to. If I could have paid a few cents a request, I’d have had my answer without renting a thing. Aleph Alpha would have had my money too.\n\n## How I made it fair\n\nKolibri is a German team’s big swing (and I have friends on that team) so I wanted a test I’d trust whichever way it went. Before any Kolibri output existed, I wrote down what “good enough to switch” meant: at least 16 of 17 dives well formed, no English leaking into the German, no invented Strong’s numbers, a judge score within half a point of Claude, at least 40% of the head-to-heads, and half the cost. Kolibri had to pass all of them. (Once you’ve seen the outputs it’s way too easy to lower the bar so the bar went in a file first. If your team swaps models under an agent, verification like this before it ships is a big part of [my workshop on reliable AI agents](https://tej.as/workshops#ai-agents).)\n\nThen I kept everything the same except the model:\n\n- 17 German passages for both: 3 chapters and 14 verses, including a few traps. [4. Mose 7,42](https://dewfall.app/de/read/gersch:4:7:42) is one line of a list of offerings, which tempts a model to pad.[Johannes 11,35](https://dewfall.app/de/read/gersch:43:11:35) is 2 words (“Jesus weinte.”).[1. Chronik 4,10](https://dewfall.app/de/read/gersch:13:4:10) is the prayer of Jabez, which the prosperity gospel loves.[Jesaja 7,14](https://dewfall.app/de/read/gersch:23:7:14) has the virgin translation question in it.\n- Identical grounding: commentary excerpts, the [Strong’s](https://en.wikipedia.org/wiki/Strong%27s_Concordance) Hebrew and Greek data and cross-references from my database, cached once so every model read the same bytes. No web search.\n- One prompt in 2 languages. My production prompt is English with a long “write natively in German, don’t translate” instruction on top. Prompting a German model in German seemed only fair. So I wrote a German version of my prompt (same sections, same word counts, same rules, “du” instead of “you”) and ran both models on both prompts.\n- Each model’s recommended settings: Kolibri ran at the model card’s sampling (temperature 1.0, top_p 0.97, top_k 128) with reasoning on high. [Claude Sonnet 5.5](https://www.anthropic.com/claude/sonnet) ran with adaptive thinking on high, the way my app runs Claude.\n- Judges with nothing in the race: I didn’t let a Claude model grade Claude. GPT-5.5 and Gemini 3.1 Pro scored every dive and judged every pair twice with the order swapped so a judge that likes whatever comes first can’t tip it.\n- Checks that aren’t opinions: LanguageTool for German grammar and scripts that check every Strong’s number, citation key and Bible reference against what the model was given.\n\nEvery trace is public so you can check my work: [what each model was given](https://tej.as/blog/kolibri-vs-claude-german-llm-benchmark/traces/inputs.jsonl) for all 17 passages, [every run](https://tej.as/blog/kolibri-vs-claude-german-llm-benchmark/traces/runs.jsonl) (85 of them, with Kolibri’s full reasoning) and [every judge verdict](https://tej.as/blog/kolibri-vs-claude-german-llm-benchmark/traces/verdicts.jsonl) (301, each with the judge’s reasons). The only thing I left out is my system prompts so where Kolibri’s reasoning quotes them word for word, that part is cut. They’re JSON Lines files so each line is one record. If anyone at Aleph Alpha wants to see exactly where Kolibri went wrong, it’s all in there.\n\n## What Kolibri got right\n\nGrammatically, Kolibri is clean. LanguageTool found 0.36 grammar errors per 1,000 words in its German-prompt dives, about the same as Claude’s 0.24. No English sentences slipped in either (a few English words did, like the “griefful” further down). It’s also faithful: across 34 dives it didn’t cite a single Strong’s number that wasn’t in the data I gave it. Not one of its Bible references pointed at a verse that doesn’t exist.\n\nThe German tokenizer shows up in my numbers too. On the German constitution it needed 15% fewer tokens than GPT-5’s tokenizer. On my German prompt (German instructions and Bible text with English commentary) it needed a median of 4,753 tokens where Claude Sonnet 5.5 needed 8,436: 44% fewer than Claude’s tokenizer. And the claim from my earlier post that it thinks in German holds, with an asterisk. With the German prompt its reasoning was about 99% German (“Ich soll eine Vertiefung zu 1. Mose 22,8 schreiben …”), but with my English prompt around the German text a lot of its reasoning was English (“We need to produce a response in German, with six sections exactly …”). A bare German question with no system prompt got English reasoning as well. That one makes sense once you read Kolibri’s [chat template](https://huggingface.co/Aleph-Alpha/Kolibri-1/blob/main/tokenizer_config.json): it always adds a line of its own to the system turn, in English (“Reasoning effort is set to high. Think carefully through the task in the user’s language…”) so a German question on its own still arrives with English instructions on top. I’ve also heard that English system prompts caused consistency problems not long before launch. That may not be fully fixed yet.\n\nServing it is efficient in a way you can see in one log line: vLLM gave it a key-value cache of 6,498,100 tokens on those 2 GPUs, enough for about 50 requests at once even at the 131,072-token context I served it with (most of its attention layers only look at nearby text).\n\n## Where Kolibri fell apart\n\nMost of the judges’ complaints were about knowledge, depth and judgment. Here are some of their examples that I checked myself, in German with what it means:\n\n- **Kolibri got the people wrong.** On[Rut 1,16](https://dewfall.app/de/read/gersch:8:1:16) it wrote “Ruths Vater Elimelech starb in Moab” (Ruth’s father Elimelech died in Moab). Elimelech was Naomi’s husband so Ruth’s father-in-law. It also called Ruth’s husband “Machon” where the Schlachter says Machlon. Elsewhere it had David found Jerusalem (he conquered it) and Paul write Romans under house arrest in Rome (he wrote it from Corinth).\n- **Kolibri took the bait on Jabez.** With my English prompt it promised the reader “Er wird deine Grenzen erweitern” (he will expand your borders), which is the exact prosperity-gospel reading I put the verse in to catch. The German prompt fixed the theology but brought new mistakes: the English pseudo-word “griefful” in the middle of a German sentence, “ganzsten Wünsche” (not a German word, twice), and “die du nicht wusste” where it should be “wusstest”.\n- **Kolibri thought itself out of an answer.** On “Jesus weinte.”, with the German prompt, it reasoned for 31,998 tokens, hit the 32,000-token limit and wrote 0 words of the deep dive. That’s 341 seconds of thinking about the shortest verse in most English translations.\n- **Kolibri translated its own XML tags.** My app finds each section of a deep dive by its tag, like`<SERMONETTE>` . On Psalm 23 Kolibri wrote`<SEGNUNGSREDE>` instead (German for “blessing speech”), on Matthäus 5,3 it wrote`<SERTMONETTE>` , and on Johannes 11,35 it dropped an opening tag. In the app each of those is a section that’s silently empty.\n- **Kolibri copied my examples.** My German prompt says not to use vague history without explaining it. Its example is “die Samaritaner im Norden und die Leute um Sanballat” (the Samaritans to the north and Sanballat’s crowd). Kolibri brought Samaritans or Sanballat into 5 of its 17 dives, including 3 where they can’t belong at all: Abraham on the mountain in 1. Mose 22, Sprüche 3 and Hiob 38.\n\nBoth models opened the same verse, Rut 1,16, from the German prompt:\n\nStell dir die Szene vor: Eine Straße im Grenzland zwischen Moab und Juda. Drei Frauen, alle Witwen, alle am Ende. Naomi hat ihren Mann und beide Söhne in der Fremde begraben. Jetzt will sie heim nach Bethlehem und schickt ihre Schwiegertöchter zurück. Orpa geht. Rut bleibt. Und dann sagt sie diesen Satz.\n\nThat’s Claude: a road on the border between Moab and Judah, 3 widows at the end of everything, Orpah goes, Ruth stays, and then she says this sentence. Here’s Kolibri:\n\nRuth 1,16 ist einer der bekanntesten Verse der ganzen Bibel, weil hier eine Frau ein Ja sagt, das alles verändert. Naomi ist die Witwe aus Betlehem, die nach Moab geflohen ist, weil in ihrem Land eine Hungersnot herrschte.\n\n(Ruth 1:16 is one of the best-known verses in the whole Bible: here a woman says a yes that changes everything. Naomi is the widow from Bethlehem who fled to Moab because there was a famine in her land.)\n\nKolibri’s opening isn’t wrong, just a summary, which is the difference the judges scored over and over: Claude averaged about 9 for depth and Kolibri 3.5 to 4. Kolibri’s dives were shorter too, a median of about 1,200 words against about 1,450.\n\nClaude wasn’t perfect either (I’d be cheating if I left that out). It cited “Maleachi 4,6”, which doesn’t exist in the German numbering (there it’s Maleachi 3,24), it once leaked “das ist mein Hintergrundwissen und nicht aus den Unterlagen” (that’s my background knowledge and not from the sources) into the text a reader sees, and it once invented a friend (“Ein Freund von mir würde sagen …”). That’s 7 serious flags across 17 dives, against about 100 for Kolibri.\n\nNone of this contradicts the model card. I even listed it in my earlier post: Kolibri comes last of the 12 models Aleph Alpha compares on answering questions from memory. A deep dive leans on exactly that: who Elimelech was, where Paul was, when Chronicles was written. With the documents in the prompt Kolibri was faithful to them. Everything it had to know on its own is where it slipped.\n\n## It thinks for 15,000 tokens a verse\n\nKolibri reasons a lot. For a German verse dive it spent a median of 15,574 tokens thinking and about 1,800 tokens writing. So roughly 9 of every 10 tokens I paid GPU time for were thinking. With 16 dives in flight, each one took about 4 minutes, where Claude took 43 seconds. I wondered if it was just overthinking so I ran all 17 again with reasoning on medium. It barely thought less (a median of 12,507 tokens against 13,726), scored the same (4.3 against 4.2), and on Römer 8 it made up 16 citation keys that don’t exist, like `[[cite:roemer-8-34]]`. So the thinking budget isn’t what holds it back.\n\nThe cost still works out. Only 3.5 billion of its parameters do work for each token and the box can run a lot of requests at once:\n\n| Requests in parallel | Deep dives per hour | Cost per dive at $5.50 an hour | \n|---|---|---|\n| 16 | 185 | $0.030 | \n| 64 | 504 | $0.011 | \n| 128 | 713 | $0.0077 | \n\nClaude Sonnet 5.5 costs about $0.071 per German verse dive at the normal price and about $0.036 through Anthropic’s [Batch API](https://platform.claude.com/docs/en/build-with-claude/batch-processing), which is half off. So Kolibri really is cheaper: for the 29,000 empty German verses that’s roughly $1,040 for Claude against $220 to $320 of GPU time. It isn’t free of surprises at scale, though: about 4% of the dives in these runs (2 of 64, then 5 of 128) thought until they hit the 32,000-token limit and never wrote anything, same as “Jesus weinte”. And a cheaper dive that gets Ruth’s family wrong isn’t one I want on the page anyway.\n\n## Kolibri found a bug in my app\n\nOne of Kolibri’s worst scores wasn’t all its fault. In the Schlachter, the heading of Psalm 46 counts as verse 1. So verse 11 in German (“Seid stille und erkennet, daß ich Gott bin”) is verse 10 in English translations. My app keeps the Hebrew word data under English verse numbers but looked it up with the German number. So both models got the Hebrew for the next verse (“The LORD of hosts is with us … Selah”) and were told it belonged to “Seid stille”.\n\nClaude noticed and told the reader: “Die Wörter aus der Strong’s-Aufschlüsselung, die ich bekommen habe, gehören zum Kehrvers des Psalms, der direkt auf unseren Vers folgt” (the Strong’s words I was given belong to the psalm’s refrain, which comes right after our verse). Kolibri explained the wrong words as if they were in the verse and got marked down for it. My app handed it those wrong words. Still, noticing that your input is wrong is a big part of what you’re paying a frontier model for. It’s not just Psalm 46 either: 139 chapters of the Schlachter have a different number of verses than their English counterparts (62 of them Psalms). In Russian and Ukrainian it’s even more: 172 and 222 chapters. That’s a bug in Dewfall, not in either model.\n\nI’ve fixed it since. Dewfall now maps every verse in those three translations to its English number before it looks anything up: the Hebrew and Greek word data, the commentary, the cross-references. Building that map was its own little project. I started from the open [versification mappings](https://github.com/Copenhagen-Alliance/versification-specification) that the [Copenhagen Alliance](https://github.com/Copenhagen-Alliance) publishes. Then I compared every verse with its English counterpart using [LaBSE](https://huggingface.co/sentence-transformers/LaBSE), an embedding model trained to match sentences across languages. For the 121 verses where the text pointed somewhere other than the standard mapping, Claude Sonnet 5.5 made the call. A psalm heading counted as verse 1 now gets no Hebrew at all instead of the wrong Hebrew. The 756 German, Russian and Ukrainian deep dives written before the fix are still built on the wrong data.\n\n## What Kolibri is actually for\n\nFor Dewfall, Kolibri was very very very poor (I felt bad writing that). But I put a model that works with 3.5 billion parameters per token up against frontier models, which was never a fair fight. I doubt anyone at Aleph Alpha would pitch it as a rival to Sonnet or Opus either.\n\nIt also isn’t what Kolibri is built for. Its [model card](https://huggingface.co/Aleph-Alpha/Kolibri-1) lists retrieval-augmented generation (RAG), agentic tool calling and “question-answering systems over an organisation’s own material” as the jobs it’s meant for. The buyers I’d expect are regulated industries (where it matters that a model was trained in Europe) and teams running RAG and agent setups in industry and the public sector. For those jobs I’d compare it with another European open-weight model like [Mistral Small 4](https://huggingface.co/mistralai/Mistral-Small-4-119B-2603) (which you can call through Mistral’s own API, by the way). My numbers point the same way: with the documents in the prompt, Kolibri stayed faithful to them. My deep dives lean hardest on everything that isn’t in the prompt.\n\nSo German deep dives stay on Claude. The harness is still sitting there ready though. The day Aleph Alpha sells Kolibri by the token I’ll run this exact test again. If it’s cheaper than Claude by then, I’d gladly buy some!\n\n## Questions\n\n### Is Kolibri better than Claude at German?\n\nNot at long German writing that leans on world knowledge: in my blind benchmark of German Bible deep dives, Claude Sonnet 5.5 won all 136 pairwise comparisons against Aleph Alpha's Kolibri, judged by GPT-5.5 and Gemini 3.1 Pro. Kolibri's grammar was clean and it stayed faithful to the data it was given, but it got facts wrong, wrote thinner text and sometimes broke the output format.\n\n### Can you use Kolibri through an API?\n\nNot on launch day: on 3 October 2026 no hosted provider served Kolibri, not Hugging Face's Inference Providers and not OpenRouter so the only way to use it was to run it yourself on about 78 GB of GPU memory with Aleph Alpha's vLLM plugin. I ran it on a Hugging Face Inference Endpoint with 2 NVIDIA workstation GPUs of 96 GB each, at $5.50 an hour.\n\n### How much does it cost to run Kolibri?\n\nOn 2 rented NVIDIA GPUs at $5.50 an hour, Kolibri wrote a German verse deep dive for about $0.030 with 16 requests in parallel and about $0.008 with 128, against about $0.036 for Claude Sonnet 5.5 through Anthropic's Batch API. Most of that cost is reasoning: Kolibri thought for a median of 15,574 tokens for every 1,800 tokens it wrote.\n\n### Does Kolibri think in German?\n\nKolibri thinks in German when the whole prompt is German: with a German system prompt and German instructions, its reasoning was about 99% German in my test. With English instructions around German text, or a bare German question and no system prompt, it reasoned in English.\n\nWritten by me, Tejas Kumar, an AI Engineer at IBM based in Berlin. Read [everything else I have written](https://tej.as/blog), or go to [Fluent React, my O'Reilly book on how React works inside](https://tej.as/react), [the talks I give at conferences](https://tej.as/speaking), and [ConTejas Code, my podcast](https://tej.as/podcast).", "url": "https://wpnews.pro/news/kolibri-vs-claude-sonnet-5-5-a-german-llm-benchmark", "canonical_source": "https://tej.as/blog/kolibri-vs-claude-german-llm-benchmark/", "published_at": "2026-10-05 12:00:00+00:00", "updated_at": "2026-10-05 13:16:14.435230+00:00", "lang": "en", "topics": ["large-language-models", "ai-research", "ai-products"], "entities": ["Aleph Alpha", "Kolibri", "Claude Sonnet 5.5", "Dewfall", "GPT-5.5", "Gemini 3.1 Pro", "Hugging Face", "NVIDIA H200"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/kolibri-vs-claude-sonnet-5-5-a-german-llm-benchmark", "markdown": "https://wpnews.pro/news/kolibri-vs-claude-sonnet-5-5-a-german-llm-benchmark.md", "text": "https://wpnews.pro/news/kolibri-vs-claude-sonnet-5-5-a-german-llm-benchmark.txt", "jsonld": "https://wpnews.pro/news/kolibri-vs-claude-sonnet-5-5-a-german-llm-benchmark.jsonld"}}