TII trains a 7B Falcon model to answer in Emirati Arabic The Technology Innovation Institute (TII) released Falcon-Emirati-7B on October 6, a 7-billion-parameter model adapted to Emirati Arabic that TII reports scores 84.83% on Alyah, an Emirati-dialect benchmark the team helped develop. The model builds on the 7B version of Falcon-H1-Arabic, introduced in January, and was trained on three data streams: native-dialect web content and forums, Modern Standard Arabic material on Emirati culture, and synthetic dialect examples generated with glossaries and style rules. TII says the 7B size balances quality against training and inference cost, since the 3B version lacked sufficient capacity and adapting the 34B model was not pursued. TII trains a 7B Falcon model to answer in Emirati Arabic The Technology Innovation Institute reports an 84.83% score on its Emirati-language benchmark, which tests dialect, cultural knowledge and poetry. By Ryan Merket https://runtimewire.com/author/ryan-merket ยท Published Primary source: Hugging Face Newsroom https://huggingface.co/blog/tiiuae/falcon-emirati Why it matters Dialect performance affects whether a model can respond in the register users expect. Falcon-Emirati pairs a 7B model with a test set focused on cultural meaning and register; its headline scores are TII-reported evaluations. The Technology Innovation Institute TII released Falcon-Emirati-7B https://huggingface.co/blog/tiiuae/falcon-emirati?ref=runtimewire on October 6th, a language model adapted to understand and generate Emirati Arabic, including its idioms, cultural references and everyday register. TII's researchers report an 84.83% score on Alyah, an Emirati-dialect benchmark they helped develop; that result is a measure of the team's own evaluation, not an independent assessment. The Technology Innovation Institute https://www.tii.ae/?ref=runtimewire , Abu Dhabi's applied research institute, is part of the emirate's Advanced Technology Research Council. The Hugging Face post credits researchers including Shaikha Alsuwaidi https://huggingface.co/Shaikha710?ref=runtimewire , Omar Alkaabi https://huggingface.co/Omar-Alkaabi?ref=runtimewire , Mohammed Alyafeai https://huggingface.co/Alyafeai?ref=runtimewire and Basma Boussaha https://huggingface.co/basma-b?ref=runtimewire . Hakim Hacid https://www.tii.ae/team/dr-hakim-hacid?ref=runtimewire , also named on the release, is TII's chief researcher for its AI and Digital Science Research Center. Boussaha's earlier work https://basma-b.github.io/?ref=runtimewire includes research on neural dialogue systems, question answering and machine translation. The team is addressing how models handle Emirati Arabic's local vocabulary, conversational conventions and cultural references. Modern Standard Arabic is common in formal writing, but those features can disappear when a model defaults to formal Arabic. Nabati poetry, proverbs and jokes depend on meaning that a literal translation may miss. The adaptation is as much about data as model size Falcon-Emirati-7B builds on the 7-billion-parameter version of Falcon-H1-Arabic https://huggingface.co/blog/tiiuae/falcon-h1-arabic?ref=runtimewire , which TII introduced in January. That family combines Mamba state-space components with Transformer attention in parallel blocks. TII says the hybrid design aims to handle long sequences efficiently while retaining attention to long-range relationships. For the Emirati version, TII describes three data streams: web content and forums written natively in the dialect; Modern Standard Arabic material about Emirati culture and identity; and synthetic dialect examples generated with glossaries and style rules. Researchers say the synthetic material helped fill gaps in topics a conversational model needs to cover, while native-speaker reviews helped assess whether responses sounded natural and culturally appropriate. A central constraint is the limited amount of written Emirati Arabic online. Emirati Arabic is spoken more often than it is written, leaving less text for training than widely represented formal Arabic. A larger general-purpose model may know Arabic vocabulary and grammar, yet still miss local usage or answer in the wrong register. TII says its experiments tested different training stages and data mixes to balance dialect fluency with cultural context. The choice of 7 billion parameters reflects a deployment tradeoff, according to TII. The researchers describe it as a balance between model quality and training and inference cost. They say the 3B version did not provide enough capacity for the depth of linguistic and cultural knowledge they wanted, while adapting the 34B model would increase costs for a specialized chat model. Users can try Falcon-Emirati-7B through TII's Falcon chat platform https://chat.falconllm.tii.ae/?model=Falcon-Emirati-7B&ref=runtimewire . What the benchmark establishes The score TII reports comes from Alyah https://huggingface.co/blog/tiiuae/emirati-benchmarks?ref=runtimewire , a multiple-choice benchmark of 1,173 questions collected manually from native Emirati speakers. It covers everyday expressions, etiquette, figurative language, heritage and poetry. TII says Falcon-Emirati-7B scored 84.83%, ahead of the other Arabic and multilingual models in its comparison. The score tests specialized dialect knowledge; it does not by itself show how the model performs in every open-ended conversation or real-world application. TII's release post https://huggingface.co/blog/tiiuae/falcon-emirati?ref=runtimewire says it also ran open-ended tests scored by Gemini 3.7 Flash https://runtimewire.com/models/google/gemini-3.7-flash , separating correctness from whether an answer actually used Emirati Arabic. On the judge's partial-credit dialect-fidelity score, TII reports 0.52 for Falcon-Emirati-7B, compared with 0.05 for ALLaM-7B-Instruct-preview, 0.03 for Gemma 3 27B, 0.02 for Jais-2-8B-Chat https://runtimewire.com/models/huggingface/inception42-jais-2-8b-chat-750e11ac9bd39e7c and effectively zero for Fanar-2-27B-Instruct. The comparison tests whether a system can respond in Emirati Arabic when a user expects it, even if it knows the answer. Alyah's questions were manually collected, and the open-ended comparison relies on an LLM judge's assessment of correctness and dialect fidelity. TII's post describes the comparison and its methods; the figures should be read as the institute's reported results, not a third-party certification. TII also cautions that cultural judgments can be subjective and that the model can make mistakes, especially on rare expressions and localized references. Falcon-Emirati tests whether an Arabic model can recognize when formal Arabic is the wrong answer. The release makes a 7B system and a dialect-focused test set available to researchers. Whether native speakers and developers find the model useful beyond the benchmark and sample prompts remains untested.