{"slug": "tii-trains-a-7b-falcon-model-to-answer-in-emirati-arabic", "title": "TII trains a 7B Falcon model to answer in Emirati Arabic", "summary": "The Technology Innovation Institute (TII) released Falcon-Emirati-7B on October 6, a 7-billion-parameter model adapted to Emirati Arabic that TII reports scores 84.83% on Alyah, an Emirati-dialect benchmark the team helped develop. The model builds on the 7B version of Falcon-H1-Arabic, introduced in January, and was trained on three data streams: native-dialect web content and forums, Modern Standard Arabic material on Emirati culture, and synthetic dialect examples generated with glossaries and style rules. TII says the 7B size balances quality against training and inference cost, since the 3B version lacked sufficient capacity and adapting the 34B model was not pursued.", "body_md": "# TII trains a 7B Falcon model to answer in Emirati Arabic\n\n**The Technology Innovation Institute reports an 84.83% score on its Emirati-language benchmark, which tests dialect, cultural knowledge and poetry.**\n\n        By [Ryan Merket](https://runtimewire.com/author/ryan-merket)\n        · Published \n\nPrimary source: [Hugging Face Newsroom](https://huggingface.co/blog/tiiuae/falcon-emirati)\n\n## Why it matters\n\nDialect performance affects whether a model can respond in the register users expect. Falcon-Emirati pairs a 7B model with a test set focused on cultural meaning and register; its headline scores are TII-reported evaluations.\n\nThe Technology Innovation Institute (TII) released [Falcon-Emirati-7B](https://huggingface.co/blog/tiiuae/falcon-emirati?ref=runtimewire) on October 6th, a language model adapted to understand and generate Emirati Arabic, including its idioms, cultural references and everyday register. TII's researchers report an 84.83% score on Alyah, an Emirati-dialect benchmark they helped develop; that result is a measure of the team's own evaluation, not an independent assessment.\n\nThe [Technology Innovation Institute](https://www.tii.ae/?ref=runtimewire), Abu Dhabi's applied research institute, is part of the emirate's Advanced Technology Research Council. The Hugging Face post credits researchers including [Shaikha Alsuwaidi](https://huggingface.co/Shaikha710?ref=runtimewire), [Omar Alkaabi](https://huggingface.co/Omar-Alkaabi?ref=runtimewire), [Mohammed Alyafeai](https://huggingface.co/Alyafeai?ref=runtimewire) and [Basma Boussaha](https://huggingface.co/basma-b?ref=runtimewire). [Hakim Hacid](https://www.tii.ae/team/dr-hakim-hacid?ref=runtimewire), also named on the release, is TII's chief researcher for its AI and Digital Science Research Center. [Boussaha's earlier work](https://basma-b.github.io/?ref=runtimewire) includes research on neural dialogue systems, question answering and machine translation.\n\nThe team is addressing how models handle Emirati Arabic's local vocabulary, conversational conventions and cultural references. Modern Standard Arabic is common in formal writing, but those features can disappear when a model defaults to formal Arabic. Nabati poetry, proverbs and jokes depend on meaning that a literal translation may miss.\n\n### The adaptation is as much about data as model size\n\nFalcon-Emirati-7B builds on the 7-billion-parameter version of [Falcon-H1-Arabic](https://huggingface.co/blog/tiiuae/falcon-h1-arabic?ref=runtimewire), which TII introduced in January. That family combines Mamba state-space components with Transformer attention in parallel blocks. TII says the hybrid design aims to handle long sequences efficiently while retaining attention to long-range relationships.\n\nFor the Emirati version, TII describes three data streams: web content and forums written natively in the dialect; Modern Standard Arabic material about Emirati culture and identity; and synthetic dialect examples generated with glossaries and style rules. Researchers say the synthetic material helped fill gaps in topics a conversational model needs to cover, while native-speaker reviews helped assess whether responses sounded natural and culturally appropriate.\n\nA central constraint is the limited amount of written Emirati Arabic online. Emirati Arabic is spoken more often than it is written, leaving less text for training than widely represented formal Arabic. A larger general-purpose model may know Arabic vocabulary and grammar, yet still miss local usage or answer in the wrong register. TII says its experiments tested different training stages and data mixes to balance dialect fluency with cultural context.\n\nThe choice of 7 billion parameters reflects a deployment tradeoff, according to TII. The researchers describe it as a balance between model quality and training and inference cost. They say the 3B version did not provide enough capacity for the depth of linguistic and cultural knowledge they wanted, while adapting the 34B model would increase costs for a specialized chat model. Users can try Falcon-Emirati-7B through [TII's Falcon chat platform](https://chat.falconllm.tii.ae/?model=Falcon-Emirati-7B&ref=runtimewire).\n\n### What the benchmark establishes\n\nThe score TII reports comes from [Alyah](https://huggingface.co/blog/tiiuae/emirati-benchmarks?ref=runtimewire), a multiple-choice benchmark of 1,173 questions collected manually from native Emirati speakers. It covers everyday expressions, etiquette, figurative language, heritage and poetry. TII says Falcon-Emirati-7B scored 84.83%, ahead of the other Arabic and multilingual models in its comparison. The score tests specialized dialect knowledge; it does not by itself show how the model performs in every open-ended conversation or real-world application.\n\n[TII's release post](https://huggingface.co/blog/tiiuae/falcon-emirati?ref=runtimewire) says it also ran open-ended tests scored by [Gemini 3.7 Flash](https://runtimewire.com/models/google/gemini-3.7-flash), separating correctness from whether an answer actually used Emirati Arabic. On the judge's partial-credit dialect-fidelity score, TII reports 0.52 for Falcon-Emirati-7B, compared with 0.05 for ALLaM-7B-Instruct-preview, 0.03 for Gemma 3 27B, 0.02 for [Jais-2-8B-Chat](https://runtimewire.com/models/huggingface/inception42-jais-2-8b-chat-750e11ac9bd39e7c) and effectively zero for Fanar-2-27B-Instruct. The comparison tests whether a system can respond in Emirati Arabic when a user expects it, even if it knows the answer.\n\nAlyah's questions were manually collected, and the open-ended comparison relies on an LLM judge's assessment of correctness and dialect fidelity. TII's post describes the comparison and its methods; the figures should be read as the institute's reported results, not a third-party certification. TII also cautions that cultural judgments can be subjective and that the model can make mistakes, especially on rare expressions and localized references.\n\nFalcon-Emirati tests whether an Arabic model can recognize when formal Arabic is the wrong answer. The release makes a 7B system and a dialect-focused test set available to researchers. Whether native speakers and developers find the model useful beyond the benchmark and sample prompts remains untested.", "url": "https://wpnews.pro/news/tii-trains-a-7b-falcon-model-to-answer-in-emirati-arabic", "canonical_source": "https://runtimewire.com/article/tii-falcon-emirati-arabic-model", "published_at": "2026-10-07 00:53:34+00:00", "updated_at": "2026-10-07 01:18:38.058554+00:00", "lang": "en", "topics": ["large-language-models", "natural-language-processing", "ai-research", "generative-ai"], "entities": ["Technology Innovation Institute", "Falcon-Emirati-7B", "Falcon-H1-Arabic", "Alyah", "Shaikha Alsuwaidi", "Omar Alkaabi", "Mohammed Alyafeai", "Hakim Hacid"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/tii-trains-a-7b-falcon-model-to-answer-in-emirati-arabic", "markdown": "https://wpnews.pro/news/tii-trains-a-7b-falcon-model-to-answer-in-emirati-arabic.md", "text": "https://wpnews.pro/news/tii-trains-a-7b-falcon-model-to-answer-in-emirati-arabic.txt", "jsonld": "https://wpnews.pro/news/tii-trains-a-7b-falcon-model-to-answer-in-emirati-arabic.jsonld"}}