Kolibri vs Claude Sonnet 5.5: A German LLM Benchmark Aleph Alpha's Kolibri lost all 136 blind comparisons against Claude Sonnet 5.5 in a German-language Bible deep-dive benchmark run by Dewfall developer on 3 October 2026, scoring 3.7 with German prompts and 4.2 with English prompts versus Sonnet 5.5's 8.7 on both. Blind judges GPT-5.5 and Gemini 3.1 Pro flagged 100 serious errors in Kolibri's German-prompt output and 92 in its English-prompt output, against 7 for Sonnet 5.5, and Kolibri parsed into all 7 required sections in only 16 of 17 German-prompt dives and 14 of 17 English-prompt dives. Kolibri required about 78 GB of GPU memory and self-hosting via Aleph Alpha's vLLM plugin at $0.008 per verse dive, versus $0.036 per verse for Sonnet 5.5's Batch API, so the developer kept German deep dives on Claude. Kolibri vs Claude Sonnet 5.5: A German LLM Benchmark Aleph Alpha https://aleph-alpha.com/ ’s Kolibri https://huggingface.co/Aleph-Alpha/Kolibri-1 lost all 136 blind comparisons against Claude Sonnet 5.5 when I had both write Bible deep dives in German for my app. So now, my German deep dives stay on Claude. When I explained how Kolibri works https://tej.as/blog/aleph-alpha-kolibri , I ran its tokenizer over all of the German constitution, but I hadn’t run the model itself. Kolibri needs a data-center GPU and nobody hosts it. So I rented the GPUs myself and ran it on 3 October 2026, the same day it came out. Dewfall https://dewfall.app is the Bible app I build. Every verse in it has a deep dive: what the verse means, the world it came from, the Hebrew or Greek, how it points to Jesus, prayer points and questions to journal on. For German, Dewfall uses the Schlachter 1951 https://de.wikipedia.org/wiki/Schlachter-Bibel , which has 31,170 verses. Only 2,374 of them have a German deep dive as I write this every chapter plus one key verse per chapter, written in a batch by Claude Sonnet 5 . That leaves almost 29,000 German verses with nothing under them. A German model I can run on my own GPUs sounded like a great way to fill them for cheap. So before switching anything, I tested it. The results at a glance | Model | German prompt | English prompt | |---|---|---| | Sonnet 5.5 | 8.7 | 8.7 | | Kolibri | 3.7 | 4.2 | | | Kolibri, German prompt | Kolibri, English prompt | Sonnet 5.5, German prompt | Sonnet 5.5, English prompt | |---|---|---|---|---| | Blind judge score 1 to 10 | 3.7 | 4.2 | 8.7 | 8.7 | | Wins against Sonnet 5.5 same prompt | 0 of 68 | 0 of 68 | | | | Serious errors the judges flagged | 100 | 92 | 7 | 7 | | Dives that parsed into all 7 sections | 16 of 17 | 14 of 17 | 17 of 17 | 17 of 17 | | Grammar errors per 1,000 words LanguageTool https://languagetool.org/ | 0.36 | 0.68 | 0.24 | 0.29 | | Median tokens of reasoning per verse | 15,574 | 13,262 | not separated | not separated | | Cost per verse dive | $0.008 self-hosted, 128 at once | | $0.036 Batch API | | The judges were GPT-5.5 https://openai.com/ and Gemini 3.1 Pro https://deepmind.google/models/gemini/ . Neither of them knew which model wrote what. Every “win” above is one judge picking one dive over the other, in both presentation orders. Of the 136 verdicts, 135 were marked “decisive”. Wild. I had to rent the GPUs myself The model is on Hugging Face https://huggingface.co/Aleph-Alpha/Kolibri-1 so I figured I could just call it there. I couldn’t. The model page has no inference provider, Hugging Face’s router answered model not supported , OpenRouter https://openrouter.ai/ doesn’t list it, and there’s no demo Space either. Kolibri needs about 78 GB of GPU memory and Aleph Alpha’s own vLLM plugin https://github.com/Aleph-Alpha/aleph-alpha-inference . So the only way to get a single token out of it on launch day was to serve it myself. I did that with a Hugging Face Inference Endpoint https://huggingface.co/docs/inference-endpoints/index running Aleph Alpha’s container. It took more than I expected. My first token was read-only so I needed a fine-grained one that can create endpoints. Then the API said “Payment method required” even though I had Apple Pay https://www.apple.com/apple-pay/ on file. My account bills from a prepaid credit balance so I added credits. Then the single NVIDIA H200 https://www.nvidia.com/en-us/data-center/h200/ I asked for $5 an hour sat on “Waiting for requested hardware to become available” for 14 minutes and never showed up. So I raced it against a second endpoint on 2 NVIDIA RTX PRO 6000 https://www.nvidia.com/en-us/products/workstations/professional-desktop-gpus/rtx-pro-6000/ GPUs 96 GB each, $5.50 an hour , which went from created to serving in about 7 minutes. I deleted the H200. The 78 GB of weights loaded in 27 seconds, which is kinda ridiculous. Honestly, I think not also selling an API is a kinda dumb decision. Paying for API access and tokens is how you make a model company sustainable, the way OpenAI https://openai.com/ did it. Nobody cared about GPT-3 until there was a UI and an API anyone could use: the GPT-3 API https://openai.com/index/openai-api/ launched in private beta in June 2020 and kept a waitlist until November 2021 https://www.hpcwire.com/aiwire/2021/11/18/openai-gtp-3-waiting-list-is-gone-as-gtp-3-is-fully-released-for-use/ . It took ChatGPT https://openai.com/index/chatgpt/ in November 2022 and then the cheap ChatGPT API https://openai.com/index/introducing-chatgpt-and-whisper-apis/ in March 2023 for it to blow up. In my talk You are an AI engineer https://www.youtube.com/watch?v=AlsNzAoCLSA&t=385s at JSHeroes 2024 I put it the way swyx’s Rise of the AI Engineer https://www.latent.space/p/ai-engineer does: you don’t need linear algebra or Python to build with AI, “you talk to an API”. On launch day there’s nothing to talk to. If I could have paid a few cents a request, I’d have had my answer without renting a thing. Aleph Alpha would have had my money too. How I made it fair Kolibri is a German team’s big swing and I have friends on that team so I wanted a test I’d trust whichever way it went. Before any Kolibri output existed, I wrote down what “good enough to switch” meant: at least 16 of 17 dives well formed, no English leaking into the German, no invented Strong’s numbers, a judge score within half a point of Claude, at least 40% of the head-to-heads, and half the cost. Kolibri had to pass all of them. Once you’ve seen the outputs it’s way too easy to lower the bar so the bar went in a file first. If your team swaps models under an agent, verification like this before it ships is a big part of my workshop on reliable AI agents https://tej.as/workshops ai-agents . Then I kept everything the same except the model: - 17 German passages for both: 3 chapters and 14 verses, including a few traps. 4. Mose 7,42 https://dewfall.app/de/read/gersch:4:7:42 is one line of a list of offerings, which tempts a model to pad. Johannes 11,35 https://dewfall.app/de/read/gersch:43:11:35 is 2 words “Jesus weinte.” . 1. Chronik 4,10 https://dewfall.app/de/read/gersch:13:4:10 is the prayer of Jabez, which the prosperity gospel loves. Jesaja 7,14 https://dewfall.app/de/read/gersch:23:7:14 has the virgin translation question in it. - Identical grounding: commentary excerpts, the Strong’s https://en.wikipedia.org/wiki/Strong%27s Concordance Hebrew and Greek data and cross-references from my database, cached once so every model read the same bytes. No web search. - One prompt in 2 languages. My production prompt is English with a long “write natively in German, don’t translate” instruction on top. Prompting a German model in German seemed only fair. So I wrote a German version of my prompt same sections, same word counts, same rules, “du” instead of “you” and ran both models on both prompts. - Each model’s recommended settings: Kolibri ran at the model card’s sampling temperature 1.0, top p 0.97, top k 128 with reasoning on high. Claude Sonnet 5.5 https://www.anthropic.com/claude/sonnet ran with adaptive thinking on high, the way my app runs Claude. - Judges with nothing in the race: I didn’t let a Claude model grade Claude. GPT-5.5 and Gemini 3.1 Pro scored every dive and judged every pair twice with the order swapped so a judge that likes whatever comes first can’t tip it. - Checks that aren’t opinions: LanguageTool for German grammar and scripts that check every Strong’s number, citation key and Bible reference against what the model was given. Every trace is public so you can check my work: what each model was given https://tej.as/blog/kolibri-vs-claude-german-llm-benchmark/traces/inputs.jsonl for all 17 passages, every run https://tej.as/blog/kolibri-vs-claude-german-llm-benchmark/traces/runs.jsonl 85 of them, with Kolibri’s full reasoning and every judge verdict https://tej.as/blog/kolibri-vs-claude-german-llm-benchmark/traces/verdicts.jsonl 301, each with the judge’s reasons . The only thing I left out is my system prompts so where Kolibri’s reasoning quotes them word for word, that part is cut. They’re JSON Lines files so each line is one record. If anyone at Aleph Alpha wants to see exactly where Kolibri went wrong, it’s all in there. What Kolibri got right Grammatically, Kolibri is clean. LanguageTool found 0.36 grammar errors per 1,000 words in its German-prompt dives, about the same as Claude’s 0.24. No English sentences slipped in either a few English words did, like the “griefful” further down . It’s also faithful: across 34 dives it didn’t cite a single Strong’s number that wasn’t in the data I gave it. Not one of its Bible references pointed at a verse that doesn’t exist. The German tokenizer shows up in my numbers too. On the German constitution it needed 15% fewer tokens than GPT-5’s tokenizer. On my German prompt German instructions and Bible text with English commentary it needed a median of 4,753 tokens where Claude Sonnet 5.5 needed 8,436: 44% fewer than Claude’s tokenizer. And the claim from my earlier post that it thinks in German holds, with an asterisk. With the German prompt its reasoning was about 99% German “Ich soll eine Vertiefung zu 1. Mose 22,8 schreiben …” , but with my English prompt around the German text a lot of its reasoning was English “We need to produce a response in German, with six sections exactly …” . A bare German question with no system prompt got English reasoning as well. That one makes sense once you read Kolibri’s chat template https://huggingface.co/Aleph-Alpha/Kolibri-1/blob/main/tokenizer config.json : it always adds a line of its own to the system turn, in English “Reasoning effort is set to high. Think carefully through the task in the user’s language…” so a German question on its own still arrives with English instructions on top. I’ve also heard that English system prompts caused consistency problems not long before launch. That may not be fully fixed yet. Serving it is efficient in a way you can see in one log line: vLLM gave it a key-value cache of 6,498,100 tokens on those 2 GPUs, enough for about 50 requests at once even at the 131,072-token context I served it with most of its attention layers only look at nearby text . Where Kolibri fell apart Most of the judges’ complaints were about knowledge, depth and judgment. Here are some of their examples that I checked myself, in German with what it means: - Kolibri got the people wrong. On Rut 1,16 https://dewfall.app/de/read/gersch:8:1:16 it wrote “Ruths Vater Elimelech starb in Moab” Ruth’s father Elimelech died in Moab . Elimelech was Naomi’s husband so Ruth’s father-in-law. It also called Ruth’s husband “Machon” where the Schlachter says Machlon. Elsewhere it had David found Jerusalem he conquered it and Paul write Romans under house arrest in Rome he wrote it from Corinth . - Kolibri took the bait on Jabez. With my English prompt it promised the reader “Er wird deine Grenzen erweitern” he will expand your borders , which is the exact prosperity-gospel reading I put the verse in to catch. The German prompt fixed the theology but brought new mistakes: the English pseudo-word “griefful” in the middle of a German sentence, “ganzsten Wünsche” not a German word, twice , and “die du nicht wusste” where it should be “wusstest”. - Kolibri thought itself out of an answer. On “Jesus weinte.”, with the German prompt, it reasoned for 31,998 tokens, hit the 32,000-token limit and wrote 0 words of the deep dive. That’s 341 seconds of thinking about the shortest verse in most English translations. - Kolibri translated its own XML tags. My app finds each section of a deep dive by its tag, like