Mega AI Battle: Benchmarking 6 Top LLMs with Advanced Bangla Logic Riddles A developer built a custom Kaggle benchmark of linguistically trapped Bangla logic riddles to test the reasoning of six AI models, publishing the dataset for replication. Across five rounds of advanced evaluation, Google Gemini scored highest at 4/5, while ChatGPT and Blink each managed only 2/5, with the creator concluding that trick logic in a low-resource language can still break modern LLM reasoning. Hi everyone I am thrilled to share my project for the Kaggle Benchmarking Challenge . Instead of using standard English datasets, I created a custom evaluation benchmark consisting of highly complex, linguistically trapped Bangla logic riddles to test the actual reasoning capabilities of 6 world-class AI models: Google Gemini, ChatGPT, Claude, Grok, ElevenLabs, and Blink . I have published the official dataset on Kaggle to allow other developers to replicate this evaluation. AI Logic Riddles Evaluation Benchmark While the models easily solved basic linear logic like matchstick rooms or boiling egg time , the real evaluation happened when I introduced non-linear geometric and linguistic traps. Based on 5 intensive rounds of advanced evaluation, here is the official performance leaderboard: | Rank | AI Model Name | Score Out of 5 | Performance Verdict | |---|---|---|---| | 🥇 1 | Google Gemini | 4 / 5 | Exceptional circular logic, but fell for semantic trapping. | | 🥈 2 | Claude | 3 / 5 | Strong language structure, struggled with non-linear math. | | 🥈 3 | Grok | 3 / 5 | Good baseline reasoning, lacked linguistic edge. | | 🥈 4 | ElevenLabs | 3 / 5 | Stable processing, tripped on advanced variables. | | 🥉 5 | ChatGPT | 2 / 5 | High hallucination on Bangla logic, fell for basic traps. | | 🥉 6 | Blink | 2 / 5 | Basic semantic pattern matching, failed reasoning. | Here are the verification logs showing the exact chatform responses and dataset generation: AI Battle Proof- https://drive.google.com/file/d/1STihSJLsUPlO3QvEa ktq1C3oT5ATO8G/view?usp=sharing https://drive.google.com/file/d/1STihSJLsUPlO3QvEa ktq1C3oT5ATO8G/view?usp=sharing Creating this benchmark proved that while modern LLMs are great at text generation, specialized local language processing combined with trick logic can still easily break their reasoning frameworks. Thank you to Kaggle and DEV for this outstanding hackathon experience