cd /news/large-language-models/mega-ai-battle-benchmarking-6-top-ll… · home › topics › large-language-models › article
[ARTICLE · art-146995] src=dev.to ↗ pub= topic=large-language-models verified=true sentiment=· neutral

Mega AI Battle: Benchmarking 6 Top LLMs with Advanced Bangla Logic Riddles

A developer built a custom Kaggle benchmark of linguistically trapped Bangla logic riddles to test the reasoning of six AI models, publishing the dataset for replication. Across five rounds of advanced evaluation, Google Gemini scored highest at 4/5, while ChatGPT and Blink each managed only 2/5, with the creator concluding that trick logic in a low-resource language can still break modern LLM reasoning.

by read1 min views3 publishedOct 7, 2026

Hi everyone! I am thrilled to share my project for the Kaggle Benchmarking Challenge.

Instead of using standard English datasets, I created a custom evaluation benchmark consisting of highly complex, linguistically trapped Bangla logic riddles to test the actual reasoning capabilities of 6 world-class AI models: Google Gemini, ChatGPT, Claude, Grok, ElevenLabs, and Blink.

I have published the official dataset on Kaggle to allow other developers to replicate this evaluation.

AI Logic Riddles Evaluation Benchmark

While the models easily solved basic linear logic (like matchstick rooms or boiling egg time), the real evaluation happened when I introduced non-linear geometric and linguistic traps. Based on 5 intensive rounds of advanced evaluation, here is the official performance leaderboard:

Rank AI Model Name Score (Out of 5) Performance Verdict
🥇 1 Google Gemini 4 / 5 Exceptional circular logic, but fell for semantic trapping.
🥈 2 Claude 3 / 5 Strong language structure, struggled with non-linear math.
🥈 3 Grok 3 / 5 Good baseline reasoning, lacked linguistic edge.
🥈 4 ElevenLabs 3 / 5 Stable processing, tripped on advanced variables.
🥉 5 ChatGPT 2 / 5 High hallucination on Bangla logic, fell for basic traps.
🥉 6 Blink 2 / 5 Basic semantic pattern matching, failed reasoning.

Here are the verification logs showing the exact chatform responses and dataset generation:

!AI Battle Proof- https://drive.google.com/file/d/1STihSJLsUPlO3QvEa_ktq1C3oT5ATO8G/view?usp=sharing

Creating this benchmark proved that while modern LLMs are great at text generation, specialized local language processing combined with trick logic can still easily break their reasoning frameworks. Thank you to Kaggle and DEV for this outstanding hackathon experience!

── more in #large-language-models 4 stories · sorted by recency
── more on @kaggle 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/mega-ai-battle-bench…] indexed:0 read:1min 2026-10-07 · —