{"slug": "deepedu-v1-efficient-and-scalable-agentic-llms-for-vietnamese-education", "title": "DeepEdu-v1: Efficient and Scalable Agentic LLMs for Vietnamese Education", "summary": "Researchers submitted DeepEdu-v1, an AI-tutoring system for Vietnamese education built on the SCALE (Self-improving Context-Aware Learning Engine) framework, to arXiv on 25 Sep 2026. In its deployed configuration, DeepEdu achieves a nearly x2 TTFT speedup over standard vLLM serving and lifts agentic accuracy from 70.0% to 79.5% on complex tasks, with the strongest per-track gains on financial-reasoning and interactive-agent benchmarks. The system's long-context inference engine issues x7.7 fewer retrieval calls than a state-of-the-art selective-attention baseline and cuts prefill latency (TTFT) by roughly 35% while matching or improving task accuracy, addressing data-sovereignty constraints such as Vietnam's Decree 53 and hallucination on region-specific material.", "body_md": "# Computer Science > Artificial Intelligence\n\n  [Submitted on 25 Sep 2026]\n\n# Title:DeepEdu-v1: Efficient and Scalable Agentic LLMs for Vietnamese Education\n\n[View PDF](http://arxiv.org/pdf/2609.31568v1)\n\n[HTML (experimental)](https://arxiv.org/html/2609.31568v1)\n\nAbstract:AI tutoring could markedly improve learning outcomes for students in developing regions such as Vietnam, yet the two obvious paths both fall short. Cloud assistants such as ChatGPT route sensitive student data to foreign servers---violating data-sovereignty laws such as Vietnam's Decree 53---and, pre-trained on Western-centric corpora, are not organized around the national textbook curriculum, so their knowledge of local content is unsystematic and frequently hallucinated. Self-hosting an open model keeps data on-premise but hits a two-fold wall: post-training quantization (AWQ, GPTQ) tames the static weight footprint, yet the dynamic KV cache and prefill latency of long tutoring contexts still cause out-of-memory failures and slow responses on consumer GPUs, while the model keeps hallucinating on region-specific material. We present DeepEdu-v1, an AI-tutoring system for Vietnamese education built on SCALE (Self-improving Context-Aware Learning Engine), a framework with two innovations. First, a long-context inference engine amortizes token selection from per-sub-chunk to per-cluster granularity; on long-context retrieval it issues x7.7 fewer retrieval calls than a state-of-the-art selective-attention baseline, cutting prefill latency (TTFT) by roughly 35% while matching or improving task accuracy. Second, a self-improving agentic layer continuously curates a verified playbook from past interactions instead of fine-tuning, a design intended to progressively reduce reliance on dominant-language priors as trustworthy local knowledge accumulates. In its deployed configuration, DeepEdu achieves a nearly x2 TTFT speedup over standard vLLM serving and lifts agentic accuracy from 70.0% to 79.5% on complex tasks, with the strongest per-track gains across financial-reasoning and interactive-agent benchmarks.\n    \n\n### References & Citations\n\nLoading...\n\n# Bibliographic and Citation Tools\n\nBibliographic Explorer \n\n*(*[What is the Explorer?](https://info.arxiv.org/labs/showcase.html#arxiv-bibliographic-explorer))\nConnected Papers \n\n*(*[What is Connected Papers?](https://www.connectedpapers.com/about))\nLitmaps \n\n*(*[What is Litmaps?](https://www.litmaps.co/))\nscite Smart Citations \n\n*(*[What are Smart Citations?](https://www.scite.ai/))\n# Code, Data and Media Associated with this Article\n\nalphaXiv \n\n*(*[What is alphaXiv?](https://alphaxiv.org/))\nCatalyzeX Code Finder for Papers \n\n*(*[What is CatalyzeX?](https://www.catalyzex.com))\nDagsHub \n\n*(*[What is DagsHub?](https://dagshub.com/))\nGotit.pub \n\n*(*[What is GotitPub?](http://gotit.pub/faq))\nHugging Face \n\n*(*[What is Huggingface?](https://huggingface.co/huggingface))\nScienceCast \n\n*(*[What is ScienceCast?](https://sciencecast.org/welcome))\n# Demos\n\n# Recommenders and Search Tools\n\nInfluence Flower \n\n*(*[What are Influence Flowers?](https://influencemap.cmlab.dev/))\nCORE Recommender \n\n*(*[What is CORE?](https://core.ac.uk/services/recommender))\n# arXivLabs: experimental projects with community collaborators\n\narXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.\n\nBoth individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.\n\nHave an idea for a project that will add value for arXiv's community? [**Learn more about arXivLabs**](https://info.arxiv.org/labs/index.html).", "url": "https://wpnews.pro/news/deepedu-v1-efficient-and-scalable-agentic-llms-for-vietnamese-education", "canonical_source": "http://arxiv.org/abs/2609.31568v1", "published_at": "2026-09-28 19:23:23+00:00", "updated_at": "2026-09-28 19:49:55.300187+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-agents", "ai-research", "ai-infrastructure"], "entities": ["DeepEdu-v1", "SCALE", "Vietnam", "Decree 53", "arXiv", "vLLM", "AWQ", "GPTQ"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/deepedu-v1-efficient-and-scalable-agentic-llms-for-vietnamese-education", "markdown": "https://wpnews.pro/news/deepedu-v1-efficient-and-scalable-agentic-llms-for-vietnamese-education.md", "text": "https://wpnews.pro/news/deepedu-v1-efficient-and-scalable-agentic-llms-for-vietnamese-education.txt", "jsonld": "https://wpnews.pro/news/deepedu-v1-efficient-and-scalable-agentic-llms-for-vietnamese-education.jsonld"}}