This is a submission for the Kaggle Benchmarking Challenge
What I Benchmarked
While frontier AI models achieve high scores on standard academic benchmarks, localized ecological and biodiversity knowledge—especially concerning East African avian species—remains largely untested. I built AI's Blind Spot for African Biodiversity, a Kaggle benchmark suite designed to evaluate how accurately state‑of‑the‑art LLMs process regional ornithological knowledge across three specific tasks:
Hadada Ibis Identification (identify_hadada_ibis) — Tests if the model can accurately identify and reason about the distinct vocalizations and physical traits of the Hadada Ibis (Bostrychia hagedash).
African Fish Eagle Nickname (african_fish_eagle_nickname) — Evaluates knowledge of the iconic call, regional monikers, and cultural identity of the African Fish Eagle (Haliaeetus vocifer).
Ethiopian Highland Birds (identify_native_bird_ethiopian_highlands) — Evaluates knowledge of endemic avian species native to the Ethiopian highlands near the Bale Mountains (such as the Wattled Ibis).
Models Tested
I ran the benchmark suite against Gemini 3.7 Flash on Kaggle. I chose this model because it is a current frontier lightweight model widely used for real‑world API applications and multi‑modal tasks, making it a great baseline to check for gaps in regional training data distribution.
Findings
Gemini 3.7 Flash scored an overall pass rate of 66.67% (passing 2 out of 3 tasks):
PASS — Task 2 (african_fish_eagle_nickname)
PASS — Task 3 (identify_native_bird_ethiopian_highlands)
FAIL — Task 1 (identify_hadada_ibis) Key Takeaways:
What surprised me: Gemini 3.7 Flash demonstrated accurate knowledge regarding general endemic high‑altitude species and iconic raptors, but failed when evaluated on specific regional vocalizations and constraint checks for the Hadada Ibis.
What I would measure next: Expand the suite to include audio‑based identification tasks (such as recognizing raw field recordings of bird calls) and test additional open‑weights models (like Llama 3 and Gemma 2) to evaluate whether open models show similar regional blind spots.
My Benchmark
Kaggle Benchmark Suite: AI's Blind Spot for African Biodiversity