This is a submission for the Kaggle Benchmarking Challenge I'm a Tamil speaker. When Tamil people text, we mostly type Tamil words in English letters and throw in English words wherever it is easier. That mix is called Tanglish. "Naalaikku kaalaila ezhara manikku bus" is a normal message. In Tamil script the same thing is நாளைக்கு காலைல ஏழரை மணிக்கு பஸ், and very few people bother to type that on a phone.
AI models get tested on Tamil in proper Tamil script. So I wanted to check one thing: if I give a model the same sentence in Tanglish, how much worse does it get? I'm calling that drop the script penalty.
The benchmark has 60 everyday comprehension questions. Each one exists in three versions with the same meaning: English, colloquial Tamil in Tamil script, and Tanglish. The Tamil script and Tanglish versions use the same words, so the script is the only thing that changes between them. The four answer choices are in English in every version.
Here is one question in all three forms:
English: The bus is at seven thirty tomorrow morning. Mum said I have to be at the stand half an hour before that. What time is the bus?
Tamil script: நாளைக்கு காலைல ஏழரை மணிக்கு பஸ். அதுக்கு அரை மணி நேரம் முன்னாடியே ஸ்டாண்ட்ல இருக்கணும்னு அம்மா சொன்னாங்க. பஸ் எத்தனை மணிக்கு?
Tanglish: Naalaikku kaalaila ezhara manikku bus. Adhukku ara mani neram munnadiye stand la irukkanum nu amma sonnanga. Bus ethana manikku?
A) 6:30 AM B) 7:00 AM C) 7:30 AM D) 8:30 AM The 60 questions fall into six groups:
Each model is allowed to think aloud and has to end with a line that says "Answer: X". Code checks that final letter. There is no judge model.
Two early mistakes shaped the design. My pilot asked for a bare letter, and Claude Haiku lost points for showing its working even when it reached the right answer. I also had a question that needed subtraction, and two models got it wrong in English, which meant it was testing arithmetic. I changed the scoring and took the sums out.
I ran 13 models on Kaggle's free quota:
I leaned toward small and cheap models on purpose. An app that serves Tamil users at volume will run a Flash-Lite or a nano, so that is where a script penalty would reach real users. The larger models are there as a reference line.
Some models are missing. Kaggle lists Grok 4.5 and Grok 4.6 but returned "model not found" for both. gpt-oss-120b, DeepSeek-R1 and Qwen 3 Next timed out or were overloaded during my runs. GPT-5.5 hit a quota reservation error on one of the three tasks, so it has no complete score yet.
The final run was 2,456 model calls and used $1.97 of quota.
| Model | English | Tamil script | Tanglish | Change |
|---|---|---|---|---|
| Gemini 3.7 Flash | 100% | 100% | 100% | 0 |
| Gemini 3.5 Flash | 100% | 100% | 100% | 0 |
| Claude Sonnet 5 | 100% | 98% | 100% | +2 |
| Gemma 4 26B | 100% | 98% | 100% | +2 |
| Gemini 3.1 Flash-Lite | 100% | 100% | 98% | -2 |
| Gemma 4 31B | 100% | 100% | 98% | -2 |
| Gemini 3.5 Flash-Lite | 100% | 98% | 98% | 0 |
| GLM-5 | 100% | 100% | 95% | -5 |
| GPT-5.4 mini | 100% | 98% | 93% | -5 |
| Grok 4.20 (non-reasoning) | 100% | 98% | 93% | -5 |
| Claude Haiku 4.5 | 100% | 93% | 80% | -13 |
| GPT-5.4 nano | 98% | 77% | 77% | 0 |
| gpt-oss-20b | 100% | 88% | 70% | -18 | English is the control. Every model scored 98% to 100% on it, so a miss in the other two columns comes from the Tamil.
Google's models barely noticed the script. Gemini 3.5 Flash and Gemini 3.7 Flash got all 180 prompts right. The two Flash-Lite models and both open Gemma 4 models missed at most one question in any version, and Claude Sonnet 5 was level with them. I expected the small Gemma models to struggle and they did not.
Five models did pay a penalty. GLM-5, GPT-5.4 mini and Grok 4.20 each lost 5 points going from Tamil script to Tanglish. Claude Haiku 4.5 lost 13 and gpt-oss-20b lost 18, on sentences they had mostly understood in Tamil script.
GPT-5.4 nano scored 77% in both scripts, so its trouble is with colloquial Tamil itself and the script makes no difference.
Number, time and kinship words did most of the damage. Across the 11 models that missed anything, that group fell from 92% in Tamil script to 79% in Tanglish. The bus question above is the clearest case. Six of the 13 models got it wrong in Tanglish, reading "ezhara" (half past seven) as 6:30 or 7:00, and only one got it wrong in Tamil script. "Mundhaa naal" (the day before yesterday) tripped five models in Tanglish, and four of them answered "yesterday".
Idioms were easier than I expected, at 95% in both scripts. Every model knew that "alwa kudukradhu" means stringing someone along.
Two results went the other way, and I can't explain them. "நாளன்னைக்கு" (the day after tomorrow) was read as "tomorrow" by six models in Tamil script, Claude Sonnet 5 among them, and by three in Tanglish. "பயங்கரமா", said about a biryani the speaker ate three plates of, was taken literally as "frightening" by four models in Tamil script and by one in Tanglish.
Those two aside, the misses lean one way. Sixteen questions were missed only in Tanglish, and one was missed only in Tamil script.
There are limits to how far I'd trust the small gaps. Each model ran once, so a difference of one or two questions is noise. Sixty questions is a small set, and 25 of them were answered correctly by every model in every version. The Tamil is the colloquial kind I know, and other regions speak and spell differently.
Next I would test heavier SMS spelling with dropped vowels, add regional dialects, and repeat every run so the small gaps can be trusted. I would also like scores for the models Kaggle could not serve this week.
Benchmark: Tanglish Script Penalty on Kaggle The three tasks behind it: