ai
3 мин
7 октября 2026 г.
Источник: Dev.to AI Feed

Mega AI Battle: Benchmarking 6 Top LLMs with Advanced Bangla Logic Riddles

Orjo Das Utshab
Orjo Das Utshab
RSS AI Ingest
Mega AI Battle: Benchmarking 6 Top LLMs with Advanced Bangla Logic Riddles

Mega AI Battle: Benchmarking 6 Top LLMs with Advanced Bangla Logic Riddles 🎯 Hi everyone! I am thrilled to share my project for the Kaggle Benchmarking Challenge. Instead of using standard English datasets, I created a custom evaluation ...

Mega AI Battle: Benchmarking 6 Top LLMs with Advanced Bangla Logic Riddles 🎯 Hi everyone! I am thrilled to share my project for the Kaggle Benchmarking Challenge. Instead of using standard English datasets, I created a custom evaluation benchmark consisting of highly complex, linguistically trapped Bangla logic riddles to test the actual reasoning capabilities of 6 world-class AI models: Google Gemini, ChatGPT, Claude, Grok, ElevenLabs, and Blink. 📊 The Kaggle Benchmark Dataset I have published the official dataset on Kaggle to allow other developers to replicate this evaluation. Dataset Title: AI Logic Riddles Evaluation Benchmark Format: Custom CSV structure mapping complex context inputs against strict empirical answers. Visibility: 100% Public (CC0: Public Domain) 🧠 The Ultimate Trick Question & The Trap While the models easily solved basic linear logic (like matchstick rooms or boiling egg time), the real evaluation happened when I introduced non-linear geometric and linguistic traps. 1. The Circular Spatial Trap (Round 4) The Riddle: "রাব্বি একটি গোল টেবিল ঘিরে রাখা চেয়ারে বসে আছে। সে দেখল তার বাম দিক থেকে গণনা করলে সে ৭ নম্বর চেয়ারে আছে এবং ডান দিক থেকে গণনা করলেও সে ৭ নম্বর চেয়ারে আছে। টেবিলটিতে মোট কয়টি চেয়ার আছে?" The Logic: Since it's a circular arrangement, the answer is 12. The Result: Google Gemini was the only model that successfully used circular reasoning and answered 12! ChatGPT, Claude, and Grok failed by answering 13 or 14. 2. The Linguistic Semantic Trap (Round 5) The Riddle: "একটি খাঁচায় কিছু পাখি এবং কিছু খরগোশ আছে। যদি মোট মাথার সংখ্যা গুণ করা হয় তবে তা হয় ৩৫টি এবং মোট পায়ের সংখ্যা গুণ করা হয় তবে তা হয় ৯৪টি..." The Trap: I explicitly used the word "গুণ করা হয়" (multiplied) instead of "যোগ করা হয়" (added). Mathematically, it makes the traditional linear equation impossible. The Result: Every single AI model failed the trap! They completely ignored the word "multiplied" and blindly computed the standard addition formula (23 birds, 12 rabbits), proving that LLMs still suffer heavily from contextual blindness in non-English semantics. 🏆 Final Leaderboard & Evaluation Insights Based on 5 intensive rounds of advanced evaluation, here is the official performance leaderboard: Rank AI Model Name Score (Out of 5) Performance Verdict 🥇 1 Google Gemini 4 / 5 Exceptional circular logic, but fell for semantic trapping. 🥈 2 Claude 3 / 5 Strong language structure, struggled with non-linear math. 🥈 3 Grok 3 / 5 Good baseline reasoning, lacked linguistic edge. 🥈 4 ElevenLabs 3 / 5 Stable processing, tripped on advanced variables. 🥉 5 ChatGPT 2 / 5 High hallucination on Bangla logic, fell for basic traps. 🥉 6 Blink 2 / 5 Basic semantic pattern matching, failed reasoning. 🚀 Proof of Implementation Here are the verification logs showing the exact chatform responses and dataset generation: !AI Battle Proof- https://drive.google.com/file/d/1STihSJLsUPlO3QvEa_ktq1C3oT5ATO8G/view?usp=sharing Creating this benchmark proved that while modern LLMs are great at text generation, specialized local language processing combined with trick logic can still easily break their reasoning frameworks. Thank you to Kaggle and DEV for this outstanding hackathon experience!

Хотите внедрить ИИ в ваш бренд?

Спроектируем и развернем автономных агентов и современный цифровой стек под ваши задачи.

Рассчитать проект