Shanghai AI Lab's New Model Learns to Predict Concepts, Not Just Words Shanghai AI Laboratory and Shanghai Jiao Tong University's LUMIA Lab released NCP-ArchPreview, an approximately 8.94-billion-parameter model that reaches the final Stage 1 pretraining loss of Allen Institute for AI's OLMo-3-7B after using 51.3% of the training tokens, according to a technical report posted on Hugging Face for arXiv 2609.10715. Trained on 5.73 trillion tokens from Dolma 3 Mix, the model adds Next Concept Prediction to standard next-token training by pooling every four token states into a concept representation, and its Stage 1 weights are available on Hugging Face under an Apache 2.0 license. NCP-ArchPreview scored 49.04 on the report's Stage 1 overall average versus 46.59 for OLMo-3-7B, 45.26 versus 39.27 on GSM8K, and 31.38 versus 27.10 on HumanEval. A Shanghai AI model is making a plain bet: language models learn faster when they predict the next idea, not only the next token. Researchers at Shanghai AI Laboratory and Shanghai Jiao Tong University's LUMIA Lab have put a real checkpoint behind that bet. NCP-ArchPreview, their roughly 8.94-billion-parameter model, reaches the final Stage 1 pretraining loss of Allen Institute for AI's OLMo-3-7B after using 51.3% of the training tokens, according to the technical report posted on Hugging Face's paper hub for arXiv 2609.10715. That is the number to watch. The model was trained on 5.73 trillion tokens from Dolma 3 Mix, and the Stage 1 weights are already on Hugging Face under an Apache 2.0 license. This isn't a slide deck. You can download it. Most modern language models train by guessing the next token, again and again, across enormous piles of text. NCP-ArchPreview keeps that job. But it adds a second one: Next Concept Prediction. A 16-layer token encoder reads the text, then pools every four token states into a single concept representation. An 8-layer concept module predicts the next one of those representations through a product-quantized codebook built from the model's own hidden states. A 16-layer decoder then turns the result back into ordinary next-token predictions. That's the trick. DeepSeek's New AI Model Spooked Samsung and SK Hynix Investors https://startupfortune.com/deepseeks-new-ai-model-spooked-samsung-and-sk-hynix-investors/ DeepSeek's new V4.1 Flash model claims a fourfold cut in KV cache memory use, and Korean investors reacted fast: Samsung fell 3.5% and SK Hynix 2.2% in Seoul on September 11 while Micron and SanDisk held steady in the US. The drop came just two days after Goldman Sachs called a bottom in the memory downturn, complicating the idea that DeepSeek... - deepseek AI model memory footprint reduction impact https://startupfortune.com/deepseeks-new-ai-model-spooked-samsung-and-sk-hynix-investors/ - Samsung SK Hynix stock price drops September https://startupfortune.com/deepseeks-new-ai-model-spooked-samsung-and-sk-hynix-investors/ The model is not literally learning human concepts in the clean, dictionary sense. Hugging Face's model card is careful on that point, noting that the learned codes are latent representations and aren't guaranteed to map neatly to meanings a person would name. Still, the training signal is different from the usual token grind. The concept sequence is one quarter the length of the token sequence, so the model gets pressure to learn a coarser structure alongside the word-by-word task. The benchmark table gives the claim some weight. NCP-ArchPreview scored 49.04 on the report's Stage 1 overall average, compared with 46.59 for OLMo-3-7B. On GSM8K math problems, it posted 45.26 against 39.27. On HumanEval coding tasks, it scored 31.38 against 27.10. None of those numbers makes it a frontier chatbot. They do show that the token saving didn't come from simply stopping early and accepting a weaker model. The Efficiency Claim Needs Care This story fits a broader pattern in Chinese AI research. But be precise about what that pattern is. DeepSeek's Nature paper on R1, published in September 2025, put the cost of its R1 reinforcement-learning and supervised-data creation stages at about $294,000 using Nvidia H800 hardware. That figure did not cover the full cost of training the underlying base model. Treating it as the entire price tag for R1 is wrong. The fair comparison is narrower. DeepSeek showed how much room there still was in training methods after the industry had started talking as if only bigger clusters mattered. NCP-ArchPreview is making a similar point at the architecture level. It doesn't claim to beat the largest US models. It says a mid-sized open model can get more training signal out of the same data by learning over chunks of meaning as well as over tokens. Frankly, that is the useful part. Frontier labs can spend hundreds of millions of dollars on compute, but most companies and research groups can't. If a method like Next Concept Prediction keeps working at larger scales, the gain compounds quickly: fewer tokens to reach a target loss, stronger benchmark scores at the same stage, and a latent interface that the team says can be adapted by updating about 17 million VQ-module parameters instead of retraining the whole backbone. There is an inference angle too. The report says feeding NCP-ArchPreview's concept representations into a speculative-decoding drafter called DFlash2 improved mean accepted length by 4.17%, with negligible extra overhead. That's a modest figure. It still matters, because serving models is where the cost keeps showing up after the training run is over. NCP-ArchPreview is best read as a serious architecture preview, not a finished consumer product. The model card says it is a pretrained base model, not an aligned assistant, and warns that it can produce inaccurate or harmful outputs. That caveat is not boilerplate. It tells you where the release sits: useful for researchers, readable for competitors, and early enough that the next question is whether the same idea holds when the model is much larger. Garry Tan Tells Regulators to Leave AI Model Distillation Alone https://startupfortune.com/garry-tan-tells-regulators-to-leave-ai-model-distillation-alone/ Y Combinator CEO Garry Tan says he'd "do nothing" to stop AI model distillation, even as U.S. agencies accuse six Chinese firms of stripping Claude and GPT for training data. Instead, he wants American open-weight labs to distill frontier models too, building a non-Chinese alternative rather than policing the practice. - AI model distillation regulations Washington should avoid https://startupfortune.com/garry-tan-tells-regulators-to-leave-ai-model-distillation-alone/ - Chinese firms copying Claude and GPT models https://startupfortune.com/garry-tan-tells-regulators-to-leave-ai-model-distillation-alone/ For now, the receipts are public. The paper is online, the weights are online, and the benchmark table is specific enough to argue with. That is exactly how an architecture bet should enter the market. Also read: AI Agents Can Burn 136 Times More Power Than a Standard Chatbot Query https://startupfortune.com/ai-agents-can-burn-136-times-more-power-than-a-standard-chatbot-query/ • France Just Minted Six AI Billionaires From Mistral and Hugging Face https://startupfortune.com/france-just-minted-six-ai-billionaires-from-mistral-and-hugging-face/ • Palantir's Shyam Sankar Calls AI Doomerism a Fundraising Shtick https://startupfortune.com/palantirs-shyam-sankar-calls-ai-doomerism-a-fundraising-shtick/ This article is posted in AI News https://startupfortune.com/category/ai/ , check it out for more related stories. Join the discussion Open in the community → /community/ Almost there. Sign in and your reply posts straight away.