cd /news/artificial-intelligence/deepseek-s-liang-wenfeng-breaks-his-… · home topics artificial-intelligence article
[ARTICLE · art-70820] src=fredgao.com ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

DeepSeek's Liang Wenfeng Breaks His Silence

DeepSeek founder Liang Wenfeng, in a rare four-hour talk with investors on May 20, argued that China's only real gap with the U.S. in AI is compute and that America's lead among top models is cyclical. He described AGI as a tide no single company can own, advocated for open-source models and fair-profit pricing, and predicted that Nvidia's CUDA ecosystem wall is being broken down, creating a historic chance for Chinese chips.

read97 min views1 publishedJul 23, 2026
DeepSeek's Liang Wenfeng Breaks His Silence
Image: source

In a rare four-hour talk, the reclusive founder reveals an almost Daoist philosophy of AI—AGI as a tide no company can own, he argues China's only real gap with America is compute.

Today, a record of a conversation between Liang Wenfeng, the founder of the Chinese leading AI company DeepSeek, and its investors was made public. The talk happened on May 20. In it, Liang explained in detail DeepSeek’s vision, why he thinks open-source models can help businesses, his views on American AI models, how he compares AI in China and the US, and his picture of the future AI world.

Chinese AI founders are often different from American ones. They tend to stay away from the media and put all their energy into building models. They seldom get the chance to fully share their vision for the industry. So this talk, which lasted nearly four hours, gives us a rare look into his thinking.

I don’t usually like ideas that are too abstract, but Liang’s understanding of AI is almost like Daoist philosophy. He does not see AI simply as a tool to build a monopoly. Instead, he sees it as a huge wave that is changing human history, a tide that no single company can ever own. That is why he believes the best attitude is to hold back and be kind: you should set limits on your own gain, earn only a fair profit, and share openly to reduce pushback. In his heart, the real goal is to reach AGI. All other business wins are just side products, natural results as the technology moves forward, not the main goal. One interesting point: he thinks the biggest key to reaching AGI is not finding the very best people, but keeping the team stable.

Liang believes that the lead among top American models, such as OpenAI, Anthropic, and Google, is cyclical and will not last. He says Anthropic’s early edge in coding agents will soon fade. When talking about the gap between China and the US, he says America’s lead comes only from having more computing power. He thinks talent is spread around the world by chance, and China has no real shortage of it. In his eyes, American models keep getting bigger because they have plenty of resources, and their basic business idea is to lock up the market with closed models and high profits. But this pursuit of very high profits is weak in terms of strategy, because it will surely be beaten by those who are happy to earn only a fair profit.

He believes China has built-in advantages in cost and user experience, since American companies have no strong reason to push costs to the very limit. He pictures a future where competition looks like that in other manufacturing fields: China will play the role of offering cheaper services in a systematic way. At the same time, he thinks Nvidia’s CUDA ecosystem wall is being broken down by new technology, and that building a new ecosystem for Chinese chips is a historic chance. Once the limit on how many chips can be made is overcome, the gap in basic computing power between China and the US will go away.

In his view, the future AI world will never be a pyramid controlled by one or a few giants. It will be a system based on holding back, one where everyone gains and lives together. First, the number of companies working on base models will shrink; in the end, only three or four need to compete fully, and any that try to make very high profits will be pushed out. Second, he insists that open source is the base of the ecosystem. He believes a pricing plan that earns only a small profit, such as getting back the hardware costs in ten months, can let open source and business work together without conflict. This will push more partners to build apps on top of open source models. Finally, DeepSeek itself will focus strongly on the main road to AGI: mainly language models, chains of thought, agents, and continual learning. It will actively leave aside areas like video generation and world models, giving those business opportunities to the whole ecosystem and to society. By making only the most important tech advances, the company hopes to help the whole ecosystem grow, with an attitude of not fighting or competing.

For other highlights, I would recommend X.PIN’s post The DeepSeek Doctrine Below is the transcript I made with the help of AI translation:

Audio “deepseek 0520.m4a”, total length approximately 3 hours 44 minutes. This transcript was auto-transcribed from speech recognition and organized by AI, without distinguishing speakers; the content in brackets indicates the audio timestamp position; individual proper nouns and numbers may contain recognition errors, please refer to the original recording as authoritative.

[00:00:01] ...as well as our company’s other colleagues, when we first started this company, our original intention was not that I ultimately wanted to make a certain amount of money, to go to the capital markets, to IPO, or anything like that. So we didn’t have that original intention.

The first few dozen people didn’t think this way at all. If someone thought that way, they wouldn’t have come. So overall, we’re doing this with tremendous goodwill toward the world, and then we feel this is useful for humanity—this is something beyond money.

Of course, later on, after discovering that the benefits of this thing are enormous, there were other temptations—that’s a separate matter. But our original intention when we set out, our vision, and the vision we’ve maintained until now, was not done in a way that maximizes commercial interest. I think this point is quite crucial.

About twenty years ago, the person I admired most in management was Jack Welch, the former CEO of GE. Looking back now, most of what he said may already be wrong, but he got one most important point right: the most important thing about a company is its vision.

Managing a large company relies not on your rules and regulations, but on vision. What is vision? Vision is not a slogan hanging on the wall—vision is how you do things, not how you talk, that is, how you actually operate. Anyway, I forget Jack Welch’s exact words, but that’s roughly the meaning.

So how do we manage so many people, how are we organized? Actually, we have no organization—it’s vision-driven, organized by a vision. We have no organization.

This has advantages and disadvantages. In the future we’ll figure out how to build on strengths and avoid weaknesses, but this is our characteristic. We don’t operate in a way of “I need to achieve some KPI, no assessments”—only vision.

This vision isn’t even written down, it’s not written out, nothing has ever been written down. This vision exists in our methods of doing things, in our attitude toward the world. Perhaps everyone in our company understands this vision differently, perhaps each person’s vision is somewhat different, but on a broad direction we are aligned.

I think it’s still about having tremendous goodwill toward the world, and then wanting to accomplish something. This is what we use to organize ourselves.

Next I’ll speak first, and after I finish everyone can ask questions. I’ll probably talk about the subsequent matters revolving around this vision. This vision is real, not made up—we truly think this way, truly do it this way. Otherwise you can’t explain many of our actions.

Why do we insist so much on open source because of this vision? Because this vision itself requires open source. Without this vision, you can’t organize people.

For example, Zhipu also open-sources, but Zhipu’s open source is different from ours. Zhipu’s open source has a forced feeling—they feel it’s not their true intention, but for us, this is our true intention. Then on the matter of open source, we thought it through very clearly from the start. First, the first reason is vision; second, we believe that to succeed commercially with AI, open source has benefits.

This sounds a bit contradictory, a bit counterintuitive, because historically, open source and commercialization have been in conflict. But I think AI is different from before. Because historically, a software company’s market might be just a few billion dollars a year, and once open-sourced, it’s gone—maybe only tens of millions or a few hundred million left.

But AI is big enough—ultimately it might occupy, say, ten percent of human society’s GDP. That’s actually a very large number. One person monopolizing this thing—you can’t monopolize this thing, you must share with others, otherwise you definitely can’t survive.

This is different from previously open-sourcing a piece of software, because that software’s market wasn’t that big. But AI is simply too big. If we want to monopolize this benefit, then we’re bound to be abandoned by history. I think the main point is that this is an objective law, this is a view of history.

It’s not that if I don’t open-source, I can monopolize this market—that theoretically doesn’t conform to objective reality. You’ll definitely encounter many obstacles, there will definitely be other methods to stop you from achieving this goal.

In this situation, I think you don’t necessarily have to follow traditional business thinking. You need a mechanism to ensure that the benefits you yourself can obtain are limited—only then do you have a chance of succeeding. Restraint is needed, I think restraint is needed.

If we want to succeed at AI in our hands, first of all, I think restraint is needed. You can’t think that some percent of humanity’s GDP all belongs to me, or that some percent of China’s GDP all belongs to me. The more you think this way, the less you’ll succeed. So from the beginning we felt restraint is needed. The more restrained you are, the more likely you are to succeed at this thing. This is a commercial consideration, though it’s a macro-level consideration.

I think this conforms to intuition, at least it conforms to my intuition, or at least I truly think this way. We don’t have many other advantages, we have no special skills, we’re not richer than others, nor are our personnel better than other companies—we really aren’t.

Think about it, when we founded this company two years ago, we didn’t have much money, we didn’t have many cards [GPUs], we had no fame, no influence—we were just a group of very ordinary people.

[00:11:49] We really are just a group of ordinary people. If there’s a narrative I like, it’s a group of ordinary people accomplishing extraordinary things, rather than a group of geniuses accomplishing extraordinary things. This is closely related to our restraint—it’s of one piece with our restraint and our vision.

So, will open source and commercialization conflict? I think on this matter of AI, if you’re not restrained, you won’t rise. Open source is part of restraint, and our restraint is manifested not only in open source but in many other aspects too. But overall, we don’t need to worry about open source, don’t need to worry about restraint.

The more restrained you are, the easier it may be to succeed, or at least so far it’s been borne out, so far it can be explained. Otherwise there’s no way to explain why we could succeed: we had no weapons, our starting point was very low, our resources were very few, and our people are actually just a random group of ordinary people. I myself am just a university graduate, and not from the most top-tier school either.

This restraint is also part of our vision. AI is too big, the benefits too large. We are very restrained—as long as we can succeed, ultimately the benefits will be enormous. However you divide up even a little, the benefits are enormous, so right now there’s simply no need to consider which portion of these benefits to take, or how to take it—I think there’s simply no need to consider this, because the benefit is already big enough.

You only need to take a tiny bit and it’s already more than enough. So as we said before, we only take a reasonable profit—it depends on your willingness, not the size of profit—this is different. This isn’t our API pricing. Our API pricing considers a reasonable profit—roughly, we go to the market and buy a batch of equipment, and recover the cost in ten months. I think this is a reasonable profit.

Under current circumstances, considering you have risk, plus upfront investment and so on—if a server, we financially amortize over three or five years, but commercially, we feel roughly ten months to recover cost, we feel that’s enough, OK, we feel it’s enough. So this is the logic of our current API pricing. Our V3.2 Flash and the others all recover equipment cost in ten months.

This is our standard. It’s actually not profit-maximizing—if it were profit-maximizing, we should set prices higher. Because in this price range, user demand is inelastic, meaning if I cut the price in half, or if I raise the price by double, the token consumption doesn’t differ much.

If I make the price twice as expensive, my total revenue would be close to double. Wait a moment, let me check. Ah, great.

Let me tell you a story—it’s about our DDCP, one of our models. At first we worried demand would be too high, so we set the price relatively high at first, and people on the team weren’t very happy. Later I brought the price back down, down to a quarter, and everyone was happy.

I think this is our true thinking. It’s the vision I mentioned earlier—we still want this thing to be useful to people, rather than us making the most money—rather it’s that under the condition of being able to make a reasonable profit, everyone can afford it. I think our company’s other people’s thinking too—when we cut prices, many people in the company group cheered, everyone felt very happy.

Because this is the purpose for which we spent so much effort, so much care making this model well. The purpose is to be very cheap, very effective, letting everyone fully use it. We just feel this is happy, this is our motivation, this is our vision, this is the consensus that allows our company to come together to do this thing—this is our company’s internal consensus.

This point should be relatively special, because price cuts like this are certainly not a good thing for our competitors—they definitely don’t cheer. Because your revenue, your ARR—if you cut in half, ARR drops by half. Right, this is a place where we’re different.

We feel this is enough. From our company’s internal perspective, I recover cost in ten months—commercially I’m already very satisfied. For the company’s external perspective, we also feel this price is one everyone is happier and more willing to see—it’s win-win for everyone, win-win for the company, society, and everyone.

I think, OK, someone just left a message on screen saying ten-month payback is too high a profit. Indeed there’s still room for price cuts, indeed still room. There’s also room for optimization in the models, so overall there’s quite a lot of room for price cuts.

But this cost, ten-month payback, we ourselves can achieve—other companies can’t. Like maybe Alibaba or Tencent—they don’t have our optimization, their cost should be several times higher than this. There’s still a lot of optimization work here.

I just said why we don’t keep cutting prices—it’s because it’s inelastic. That is, if I cut prices further, demand won’t increase more, or if I cut prices further, demand increases very little. Because at this price everyone can already afford it, everyone feels this price is satisfactory, and won’t stop using it because the price is expensive.

So cutting prices—first, the company won’t have more revenue; for society there’s no more value either, because everyone is already satisfied with this price. Even if you lower the price further, it won’t increase society’s sense of happiness much. Right, OK, but on this issue just now, on this matter of pricing, we’re definitely not...

[00:25:32] ...taking company revenue highest or profit highest as our starting point. This is part of our restraint, because short-term, if your price is a bit higher, you might have a bit more revenue; but long-term, it’s really hard to say. Because I think restraint is a strategy.

For me, restraint is a strategy. It lies in the fact that sometimes you can give up some things to exchange for more of other things. The matter of not open-sourcing is actually the same—it can be seen as our pressure, or as our concession of benefits. First, this concession—for our company internally, we’re very happy, everyone is happy, employees feel a sense of accomplishment, and we’ll have cohesion because of it. And this concession is good for society, society is also happy, other peers or ordinary people will be happy. So this restraint, my understanding is, this restraint from a long-term perspective can increase our probability of achieving AGI.

When considering a matter, I have no doubt that AGI will have enormous commercial value. So on this basis, my priority consideration is not how I add a bit more share, how I take a bit more share—my priority consideration is how I increase my probability of succeeding.

This restraint is perhaps reflected in many other aspects too. For example, last Spring Festival we suddenly had many users, but we didn’t pursue retaining these users, or monetizing these users, or grabbing these commercial benefits, cashing in on the users.

We didn’t grab users, didn’t make money, but we worked very hard to figure out how to serve users well. We wouldn’t have the thought of “I want to become the next super app, and then I want to compete with someone, I want to become the next ByteDance, become the next Tencent”—we have no such thoughts at all.

We could do it this way, but we didn’t. My understanding is, this is also part of restraint. You shouldn’t think you have to profit from everything—it seems like once you have users, it seems you could become the next ByteDance, and then you just gobble that up.

I think this is workable commercially, it’s possible. If last year we had used a large sum of money to grab users from ByteDance, that would also be a way to play. But we chose a very restrained approach—that is, I won’t compete with you over this thing, because there are still watermelons behind, and what’s in front is maybe just sesame seeds.

I shouldn’t grab all the sesame seeds. Of course, maybe these sesame seeds are relatively large, but I think the AI behind might make what’s in front seem small. Looking at it now, our not pushing hard on the C-end last year was probably correct. Because you can see that behind there really are bigger watermelons, and what’s in front really is just some small sesame seeds.

If last year I had a lot of money and made this thing very big, what would be the benefit? You wouldn’t have gained anything. These are my true thoughts, because I think the AGI opportunity behind should be very large, the AGI opportunity behind is always very large. I don’t even need to consider whether I’ll occupy a position in it at that time, or what my business model in it will be—we simply don’t need to consider it. As long as there’s such a big commercial opportunity, you’ll definitely have a way. So for the sesame seeds in front, we’ll also pick some up, but we just pick up a bit casually—we won’t stop, or treat it as an important matter.

So last year’s C-end DAU and such—I think it’s probably a small matter. But we picked it up too, we maintained user usage at a relatively low cost, because it might be useful later. Although right now we don’t know what use these users have, right now it’s a pure cost expenditure, but it might be useful later.

Since it can be obtained casually, we’ll pick it up along the way. Including this year, it’s very likely that our ARR revenue from API or AI also has an opportunity. That is, if this demand can continue to expand—if it keeps expanding, and we can buy more GPUs—then ARR reaching several hundred million dollars is very possible.

If AI can reach a billion dollars, then basically my company’s cash flow could turn positive, could cover my R&D expenses, could cover all my expenses. So this is also possible, but we haven’t made it a priority. This we’ll do, but I think this is an important matter. It’s not our first priority consideration, or not what we truly care about today. The bigger opportunities should still be behind—the opportunities in front, including last year’s C-end and this year’s B-end—I think this needs to be done, needs to be done well, but this is not our goal.

Or rather, most people in our company don’t think this is a very important matter, don’t think this is a matter equally important compared to AGI.

The open source part I can say more about, because many previous questions were about open source. First, I think we will open-source, and then even our strongest model will probably be open-sourced too. Because I don’t see the benefits of closed source, I don’t see the inevitable benefits. ByteDance’s model is closed source—what benefit does it have? I don’t see any benefit.

Even if a model is open-sourced, even if you tell others everything, the threshold is still very high. For others to actually use it, the threshold is very high. For them to use it, it’s very hard; secondly, for them to use it and get the cost very low too, is also very very hard, not that easy.

It’s not that once I open-source, they can easily achieve the same deployment cost as me. There’s still a lot of work to be done here. Although the principles of this work are all understood, not every company is willing to, or has the willingness and ability to organize manpower to reach this goal.

I’m quite used to this point too. It may just not be good at doing this thing, because the resistance is too great. Controlling this cost is very hard—it has many management and physical constraints. This is also the advantage of startups, because if a startup is too small, you don’t have the strength to do this thing; if you’re a big company, you find it hard to organize.

[00:38:36] This matter has difficulties everywhere, so this belongs to the sweet spot of a company of our scale, the sweet point. Then if we were bigger, we might have other problems; if smaller, our friends’ strength would be insufficient.

So, as for open source, I think we should set pricing. Currently we should recognize, we won’t force people. Because for the pricing model, I won’t charge a very high fee—I’ll probably charge based on ten-month payback. At ten-month payback, it’s already enough to keep competitors...

At my ten-month payback, I can make independent third-party deployment unprofitable. The third party can’t do it, can’t achieve this cost—they definitely can’t.

So open source won’t affect my revenue. Of course, if I wanted to earn a hundredfold profit, then open source is...

“I can hear you, but the video seems to have dropped, boss.”

“Just now maybe a phone call came in.”

So open source, I think has no impact whatsoever on our business model. The premise is that we only earn sixfold profit—ten months to recover cost roughly corresponds to sixfold profit. Under the condition that we only earn sixfold profit, open source won’t have any impact.

But if you wanted to earn a hundredfold profit, then open source would indeed affect your earning a hundredfold profit, because third parties would deploy—they might be at twentyfold cost, which would be lower than yours.

Is this model sustainable long-term? I think it is. Under our vision, I think open source is sustainable long-term, or we intend to do it this way. You can consider it as having restraint, which is also giving you longer-term benefit.

This strategy can give us more opportunities technically at the front, and our probability of achieving AGI is also greater. We’re more at ease.

Think about it, we don’t even need to work overtime, because it’s just not that hard. But for other companies, it might be very hard, because you’re thinking about too much.

It’s actually not hard, simply not hard. At first it’s a... From the outside it might look like we chose a very hard model—we do research, we do the hardest things, seemingly a hard mode. But actually we gave up a lot in other areas, which lets us still be very strong, still doing it very easily.

So on this matter of open source, actually my judgment is that it’s sustainable. There’s no conflict between open source and commercial paid services—the premise is that under sixfold profit there’s no conflict.

Sixfold profit looks high, but actually it’s not high. Because now with AI efficiency this high, right now a reasonable profit might just be this much. In the future it might drop to, say, fourfold, threefold—I think that’s already... any lower and it couldn’t drop further. But it will still have very large profit.

Just looking at selling API alone—but I don’t think selling API is that attractive. But on this matter, you can explain, no, I don’t see any conflict.

Then I’m also not worried about others deploying our models and then competing with us—not worried at all. We even hope they can deploy it successfully.

We provide as much help as possible to the open source community, assisting everyone to deploy our models. I’m not worried they’ll steal this business from me, because this market is big enough. I’m only worried they can’t deploy it—that they got some details wrong, made the effect relatively poor, or that their cost is relatively high.

Right, there’s no conflict here.

Then last year, when there was this To B business, a question asked more often was: our C-end—if I open-source, then won’t the C-end conflict with my C-end again? Because I don’t have a traffic advantage, or Tencent itself has lots of traffic—if they deploy our open source model, they’d take over all the C-end users, taking away my C-end users.

But actually it won’t, for many reasons.

Then the question is: is the open source model we give the same as the model we deploy ourselves? It’s the same. We won’t open-source a slightly worse model and then use a better model when we deploy ourselves—we won’t do that, it’s the same.

This also shows there’s actually no conflict. All of last year, on the C-end I basically open-sourced, and then the C-end service saw no conflict, really saw no conflict. So, this is the open source part.

Then next is the company’s long-term vision—I think our goal should be AGI. Everyone’s definition of AI isn’t necessarily the same, but that doesn’t stop us from treating AGI as our goal.

From a technical roadmap perspective, actually the AGI roadmap is relatively clear. With the current generation of AI technology, if you can describe a problem very clearly, giving it complete context and instructions, it already surpasses humans. But there’s a definition here, a premise: you give it complete context, you give it complete instructions. Then this definition is very hard to achieve. For example, in today’s meeting, we actually have a very long and strong context beforehand—everyone might have decades of context, and this is something AI doesn’t possess. So what AI can currently possess is: within a limited context, it can do better than humans.

But it still can’t replace humans. There’s still one thing missing here—continuous learning. Because humans can also learn continuously. If you hire an employee, he might spend two months getting familiar with the company environment, familiar with his work, and after two months he can get up to speed.

He can do many things. He can understand what you say—for example, if you say “call Xiao Wang over,” he knows who Xiao Wang is. But with AI, because of this context, it doesn’t have those prior two months of learning.

If you tell it “call Xiao Wang over,” you have to tell it who Xiao Wang is, what position, where he is, how to find him, what to note when finding him. You have to give AI all the context. So in this situation, AI can do it, but you can’t possibly give it all the context—it’s not realistic. So AI can’t replace your employees. But if AI has the ability to learn continuously, like your employees, coming to the company to learn for two months, then...

[00:51:02] ...but then it could replace everyone in the world. So we’re one step away—we’re missing learning-to-learn.

AI’s development, we can understand it as a staircase. Last year’s step was CoT, chain of thought. Because we discovered that through chain-of-thought methods, we can bring intelligence to a higher level. Through it thinking by itself, we can raise this ceiling, let this AI do more things.

Then we crossed another step. This year’s step is Agent, because we discovered that using the Agent approach, even more things can be done, its capability range is larger, its intelligence ceiling is higher.

Why is it a staircase? Because each subsequent step is based on the previous foundation. Agent uses CoT, and CoT also uses the previous step, which is the language model—so no step is wasted.

So, AI’s development, the trajectory of intelligence, is traceable. This year’s step is Agent, but this Agent step—it will also be walked to completion. That is, after it solves all the problems it possibly can, it still can’t replace your employees, but it’s already reached its capability ceiling.

Just like CoT—after CoT reached its ceiling, it already surpassed the most top-tier humans, in doing math olympiad problems and writing programs it already surpassed the most top-tier humans. But it still stops there, that technology couldn’t reach AGI.

So you see, AI’s intelligence trajectory is traceable. Then after Agent, we think the problem to be solved should be continuous learning—that is, how to let the model learn continuously, rather than you giving it strong training—it should be able to do relatively long-term continuous learning like a human.

This problem is the same thing as completing tasks and so on—they’re related, they solve the same problem. We now stand at Agent, and what we can see is the next bottleneck, which is continuous learning. The next problem to solve is how to learn continuously—this is visible, relatively clear.

That is, the obstacle stuck in front, you must cross over it, and there must definitely be a way to cross over, but it needs time. After continuous learning, we might arrive at a singularity. This singularity is: when this model can learn continuously, it can already do all the things humans can do.

It could then develop its own version, could research on its own, then develop its own next version, developing more advanced AI models. So it would reach a singularity, able to achieve its own iteration.

But this singularity—it’s not really a singularity, it’s also a gradual process. This process might also be a relatively long gradual change, not a mutation. But habitually, we all think it might be a singularity.

Because a long time ago, those prophets thought there would be a singularity here, but actually it’s not a singularity, it’s a continuous process. Then after this step is complete, I think that’s when embodied intelligence arrives.

This is our conjecture—this is what we think the timeline should be: first solve learning-to-learn, then reach that intelligence singularity, the singularity that can self-iterate, and then comes embodied intelligence. After embodied intelligence, it walks into the real world, can do housework for you, can care for you in old age.

We think this is a relatively ideal roadmap, but everyone’s views differ, there’s no right or wrong. Just that we feel this roadmap is the most relaxed.

This roadmap is because for each step, the new things you need to do are very few. This roadmap lets us avoid working overtime. But if the roadmap were reversed—say, it needs to first achieve embodied intelligence—then this is very tiring to do ourselves, a very bitter task. We don’t want this kind of roadmap—we want to do it a bit more relaxed.

If we first solve continuous learning, then solve that self-iteration singularity, then solve embodied intelligence, this path is very relaxed. Because later on, you can use earlier technology to help develop later technology. After the singularity, then doing embodied intelligence—that doesn’t need humans to do, doesn’t need us to do, the model itself can produce it. So this answers what our long-term goal is. As I told him, this is our long-term goal, which is what we call AGI.

Everyone, back to reality. Last year’s most important reality was that everyone wanted to make Chatbots, to grab C-end traffic. This year’s reality is that everyone wants to grab To B revenue, to participate in this, because if you don’t participate, you’re simply not at the table, right?

But we don’t think this is an important matter. Or rather, within our company, what we truly care about is actually those AGI roadmaps I just mentioned, and how to break through with this technology in the next step.

But one strange thing is, the thing you most want to get, you can’t get. Whereas the thing you don’t care much about, is actually fairly easy to get.

There’s a strategic advantage here—that is, in our hearts we’re thinking about AGI, what we do is AGI. So when we do those applications, do those C-end and B-end things, we simply don’t need to put too much thought into doing those things—actually very little effort suffices.

I think it’s that, standing at a technological high ground to do relatively lower-level technology, there’s this kind of dimensional-reduction strike. At least on last year’s C-end, we saw it was indeed this way. We didn’t put much effort into the C-end—at one point we even didn’t want to maintain those users, but the users just couldn’t be driven away.

Because they really couldn’t be driven away, so in the end they all remained. But slightly... Then this year’s B-end revenue, right now this growth looks relatively optimistic. I think this figure, compared to peers, should also be relatively good, I estimate. But we didn’t put much effort into doing this thing, we simply didn’t...

[01:02:45] ...do one thing—it was just done incidentally.

Doing the launch of internet intelligence, along the way, the AGI step is one I must walk, it’s a stair-step. On the road to AGI, I must pass this stair-step. So I provide these technologies to everyone through API—I didn’t do anything extra.

We’re still doing AI—this is a byproduct. I only need a few people to maintain this API, and there aren’t even customer service reps, no need for sales, nothing is needed—users come on their own.

Or rather, what’s considered the C-end users, the C-end and B-end, are all byproducts on our road to doing AGI, an intermediate output, with no conflict with my doing AGI. I’m not doing it for the sake of doing the C-end, or for the sake of doing the B-end—rather, our doing AGI is for the sake of doing AGI, and it happens to produce these things, so I take them for commercialization.

This is different from other companies. Other companies do this for the sake of doing this—for serving C-end users, or serving B-end users, they build this model. But for us, the original intention isn’t this—our original intention is still to pursue AGI.

I think to a certain degree this is a kind of dimensional-reduction strike. AGI is a bigger vision, and this vision can cohere more excellent people—it has stronger cohesion. So I have an advantage in organization, and then I use this advantage to... it’s a kind of dimensional-reduction strike.

But if you’re a commercial company, and your vision is to serve C-end users well, then it’s a different story. It has other advantages—in product, user service, traffic there will be advantages, but in technology there’s no advantage.

The currently favorable situation is that model technology is most important. You need to make the model well, others... Right, and this can explain our previous development trajectory. We truly chose AGI, and I simply didn’t think about wanting to have many users.

When we suddenly became popular last Spring Festival, that wasn’t in our script—we hadn’t thought about that thing at all. We just wanted to make this technology well. But at that time I discovered that actually our organization, compared to fully commercialized, product-goal-oriented organizations, has extra advantages in talent and organization.

This is also amazing. At that time everyone was fighting bloody battles over the C-end, and the result was led away by someone who didn’t compete. This also indeed shows the logic I mentioned earlier has merit—there indeed are talent and organizational advantages.

This talent advantage isn’t that my people are smarter than his—rather it’s how I organize this talent, how I motivate them, and then how they cooperate. This thing has an advantage. Because when you gather smart people together, it’s not that they naturally cooperate, naturally have great passion to charge toward a goal and complete it—so you need a vision.

Our earlier experience gave me the enlightenment that the AGI vision is very powerful.

OK, that’s the question. Then the next question is, what’s the importance of core interests?

I said earlier that we need to be very restrained in many aspects. But what are our core interests? Actually our core interest is only one point: our greatest core interest is maintaining the stability of the team. This is our greatest core interest, and can even be considered our only core interest.

As long as I can maintain the stability of the team, I will definitely succeed, definitely achieve AGI—it’s that simple. As long as everyone stays, we can continue, and I will definitely... basically there’s no risk. It’s just a matter of earlier or later; we’ll encounter setbacks; if we encounter setbacks and everyone stays, then I can continue again.

Money is definitely not a problem, resources aren’t a problem, the other elements are all easily obtainable. For us, there’s only one core interest, only one thing that can’t be conceded: we must maintain the stability of the team.

This is also a very big challenge we face, or rather, I think it’s the greatest risk. Of course, this risk has been substantially relieved with our recent round of financing. Because the options everyone received are relatively many, the amounts relatively large.

From a team stability perspective, as long as some of the most important employees, the oldest employees, can be stable, then others aren’t likely to leave. Even if others have fewer options, less income, they won’t leave. Because they’re not entirely here for money—everyone hopes to do this thing in an environment that can achieve AGI. So for talent, there’s still appeal. Historically, our talent turnover has been relatively low—compared to peers, our talent turnover is always relatively low. But this is still our greatest challenge, the only challenge, you could say.

Everything else is a matter of time—everything else at most causes us to be half a year late, a year late, but won’t mean we can’t produce it. We definitely won’t lack money, definitely won’t lack resources—actually none of these are lacking.

So many of the things we do now are to maintain the stability of the team. Aside from this point, everything else I think we can do without, everything can be restrained. We’ve always been very restrained, unwilling to become adversaries with any internet giant or small company.

I hope I can empower them, or hope I can assist everyone to do this thing, hope I can help everyone do this thing. This is also, as we said before, part of our commercial meaning. The premise is that everyone doesn’t...

[01:15:12] Under this premise, we’re very willing to assist and help anyone, even our competitors, including Alibaba, Zhipu, Moonshot, to do better. Because we don’t lose anything—we’re open source anyway, and open source also means not making boundaries too clear about how to do things—we hope you can reproduce it; if you can’t reproduce it, you speak up and I’ll tell you how to reproduce.

This is originally part of open source—it won’t be otherwise just because you’re a competitor. Of course, if you’re a partner, I’ll do more. But on the big interests there’s no conflict.

When dealing with the outside, our attitude is: we only do the AGI main line. This is what I just mentioned—GPT, CoT, Agent and so on—only the main line. The AI field is very broad, there are many things we feel aren’t on this main line—for example, 3D, video generation—I think they may not have much relation to the intelligence main line, so we won’t do them.

There are also some, for example world models—I think currently they don’t have much relation to the intelligence ceiling either, so we won’t do them. But if others do them, we’re happy to help. Whether we have time is one matter, but on interests there’s no conflict.

We also hope these AI technologies can be used in all kinds of production environments, can improve society’s production efficiency, can help all industries improve production efficiency. We’re very motivated to do this.

Whether I have time, whether I have manpower, or whether our colleagues themselves are interested—that’s another matter. But on interests there’s no conflict, and we hope to achieve this goal. And we believe this has no conflict with business whatsoever—the benefits I should obtain still aren’t reduced.

I think our having held this attitude before—actually we didn’t obtain any less because of it, we didn’t obtain any less because I open-sourced, nor because of our goodwill or my providing help to others. For example, last year’s C-end users—we still have relatively many C-end users now, and they’re still relatively solid.

Our B-end this year, I think, is also relatively optimistic. I didn’t have my commercial interests affected because of our goodwill—no effect at all, and it may even be a bonus. This looks counterintuitive, but it’s indeed this way. Or let’s think in reverse—if we violated it, it would also be inevitable—could I obtain more things? No.

There’s a question: how to understand that world models have no relation to raising the AI intelligence ceiling? I’m talking about this current stage—this is our judgment. We have seen it—this is our own AI roadmap, not the only roadmap.

From our understanding, from our judgment, the most important thing currently is to do AI training well. To do AI training well doesn’t require world models, doesn’t even require multimodality. Because if you narrow the AI training scope a bit—without multimodality, there’s just a portion of tasks you can’t do, but it doesn’t affect the algorithm’s validity. Multimodality ultimately still needs to be done. Right now what’s important is training, then the next step to solve is the continuous learning problem, and the step after that is it asking its own questions. But there’s no world model in this roadmap, no video generation.

When video generation first came out it was very hot—it seemed this must be done, as if if you didn’t do it you weren’t an AI company. So I found it strange—actually if you just think carefully, it has no relation to the intelligence roadmap.

In fact, you’ll also find that after video generation first appeared with Sora, everyone did it, big companies and small companies. But small companies later all cut it. It has no relation to the intelligence ceiling.

But commercially, it’s a good business—commercially it’s a good business. But this has no relation to intelligence. We won’t do something because it’s a good business—we’ll only do it because it’s something on the intelligence roadmap.

Video generation—because that’s relatively clear, I use it as an example. Then world models—the meaning of world models isn’t that clear, because many things can be said to be world models.

From our judgment, world models and intelligence aren’t the most important thing at this current stage. The most important is AI training, and how to solve continuous learning after AI training. This is our company’s judgment—of course each company’s judgment differs. What I just said—the most important issue for our company is personnel stability. From another dimension, what do we lack? What’s our gap with the US? Actually the gap is only one thing: resources.

We don’t have that many cards [GPUs]—the number of our cards is still relatively few. We currently have roughly 20,000 H-equivalent compute cards, and most of these just arrived, arrived in the last month or two, and there may be many more machines that haven’t arrived yet.

Our total compute was relatively little last year—this year we’re very aggressively expanding this compute. We currently have roughly 20,000 H-equivalent, and in the coming months we’ll have large batches of machines bought over, basically all NVIDIA.

How many cards do we need? Right now, definitely the more the better. Within the range we can afford, definitely the more cards the better—this is without question. So our current strategy is, within a reasonable price, buy as many cards as we can buy.

If, after this round of financing is spent, however many cards I can buy, I’ll buy. The pace of spending isn’t in the plan—rather, as long as the price is reasonable, however many there are, I’ll buy them all. So if I spend all the money within half a year, I think that’s a good thing. If I spend all the money within half a year, that would be too blissful, too ideal.

Actually, spending this much money is very difficult—you can’t buy that many cards, they’re very hard to buy, and prices are very high, and you can’t spend extremely high prices to buy them—you still have to ensure the price is reasonable.

**[01:26:41]**

If I spend all this money within half a year, that might be most ideal. Because turning money into NVIDIA cards is definitely better than putting it in the bank. Putting it in the bank, I can now maybe earn two points... but buying NVIDIA cards, ten months of social cost.

So definitely, however much you can buy, buy that much. First, after buying cards, I have lots of room afterward—through providing services or however, I can always have cash flow. If I have cash flow I can survive, so I don’t need to keep tons of money in my account.

So, we only worry about not being able to buy that many cards. If we could turn all the money into cards, then we’d unhesitatingly turn all the money into cards, and here we’re willing to pay a certain premium. That is, we’re willing to pay a certain premium to turn it into cards, because this is too worthwhile.

Even after paying the premium, actually it’s still hard to achieve this goal. So objectively speaking, if this year I could spend 20 billion [yuan], then it would count as our procurement department’s performance being super good.

Our gap with the US is mainly in resources, and the gap in people isn’t very big. In people there’s almost no gap, because it’s the same batch of people—maybe Chinese people. When Chinese people go abroad, some stay domestic, some stay abroad, some go abroad—it’s not that the smart people go abroad, no.

Actually it’s relatively random. The smartest people—maybe not that slightly more than half go abroad, but slightly less than half stay domestic. Domestic talent isn’t lacking, and our base is large—we have so many people replenishing every year.

Talent isn’t the bottleneck—resources are the biggest bottleneck. Resources first affect talent cultivation, because with little compute, our opportunities to run experiments are relatively few, so our talent overall has a gap with the US. The talent gap is essentially also because of the compute gap.

On the current largest models, we actually can’t afford to train. Even if we spent all 50 billion, we still couldn’t afford to train. Even if we could stack it up, we couldn’t afford to use it. The current largest model has about 800B active parameters; domestically, we’re still at the several-tens-of-B scale—the largest domestic model might be just several tens of B active, so that’s a difference of an order of magnitude.

If I wanted to train a model as large as [theirs], it should require 50,000 GB300s, or Huawei 950s, 200,000 cards. This is just training, not yet considering doing research. So the biggest gap between us and the US is in resources. Our current resources, and our resources within this year and the coming months, including soon-to-arrive large resources, are only enough for us to do more experiments at the 10B-active scale. Because at the several-tens-of-B active scale, there are still many experiments to do, still many things to figure out. We should be relatively far from being able to train an 800B model—still plenty of time, don’t have that many cards.

So the difference between us and the US, I believe, is a difference in resources. We might believe that all the differences we see—including the talent difference, the model capability difference, the application difference—can all be considered as due to the difference in compute resources.

Compute resources—on one hand, cards simply can’t be bought domestically; on the other hand, our capital investment is less than the US. We’re much less at the capital investment level, and talent-based salaries account for a very low proportion within this. You see them offering salaries of a hundred million dollars, but calculating it, talent salaries still account for very little—the bulk is still compute.

This problem is currently basically unsolvable, because Huawei’s output is also limited. Because if I want to train 800B, I need 200,000 of Huawei’s newest cards, and this is just training, not yet considering doing research.

So we currently simply won’t consider competing with the US at such a large scale. Right now we lose, but we should first do it well at the scale we can afford to train and afford to use—the several-tens-of-B active scale. Then when we have more resources in the next step, then bring it up to 150B, 156B, or 250B active scale.

So between us and the US there’s a gap here that currently looks very hard to close. If you insist on training such a large model, you can train it, but you can’t do sufficient research. That is, before training, you can’t do sufficient research.

There’s another question many people care about: where will the endgame gap of all large-model competition manifest? That is, where will the final large-model gap manifest? I think this gap might ultimately not be very large.

The final gap should be three aspects: one is cost, one is time, one is user experience. Aside from these, there might not be any gap.

Cost is relatively easy to understand—providing the same service, same quality of service, at what cost you can provide it. The same service—say, like BYD’s battery, under the same technical volume, can other companies provide it at this price? I think this is a relatively difficult matter, not that easy to achieve—definitely a barrier.

So cost is definitely a difference—I think cost is probably the number-one difference. Then the second is time, when you can achieve it. A few months earlier or a few months later—it’s different.

The third might be experience—user experience still has some differences. Among these users there will still be some user stickiness and user barriers, but it may not be essential. The essence is still, first, cost; second, time. Whether you produce it first or later, whether you make the same thing faster or slower.

For long-term commercialization paths and further enriching product lines, how to price? We feel that right now the most worthwhile thing to do, most worthwhile to spend effort on, is still doing AGI. Right now we need to push AGI forward again, that is, push the intelligence floor up, push it forward. This should, at the current stage, be more worthwhile—or a path with greater returns—than doing more product lines or considering more commercialization paths. I think for a period in the future, and for any period in the past, it’s probably been this way. That is, if we...

[01:39:46] We spent a lot of time thinking—what’s the use of enriching our products?

What was the commercialization discussed half a year ago? It must be that I need to do advertising, then I need to do e-commerce, I need to embed e-commerce in the product, then integrate with local services and such. It’s definitely useless, because change is too fast. This product—if you’re at the front, or at this current stage of ours, and spend a lot of time on commercialization considerations and paths, the product line’s lifecycle is very short.

I think it’s not yet the time for this. This is our judgment, or at least all the previous experiences support this judgment. In the past three years, at any point when you came to talk to me about commercialization paths and product lines, it was a waste of time, because you can’t predict, can’t foresee the future. What you can foresee is very little.

Because I’m watching the time—I don’t know if everyone wants a mid-session break? Want to eat something first? Or shall we continue talking? If no one objects, I’ll continue.

On doing world-leading AGI research, and commercializing at the appropriate time—I think we’ve always been doing commercialization, I believe we’ve always been doing commercialization, just not with commercialization as the goal. We take AGI as the goal, but we’ve always been doing commercialization—that’s why we have C-end users, why we have B-end revenue.

From historical experience, this strategy is successful. Then I think the point at which we completely turn toward commercialization should be very far off. So the biggest thing is still the extension of technology, then doing the next generation of technology, more so solving the current problems. These return ratios—personally I feel that, in the currently foreseeable future, at any point focusing on products is too early. So this is also part of our restraint. I hope these commercial opportunities can be done by others—we hope these commercial opportunities, how to use this AI, is something the entire society, all our partners do together, everyone shares this return together, rather than me wanting to swallow it all, which is impossible. And we don’t have that much energy, our organization doesn’t have that many people to do this thing.

For partners, actually this round of financing was carefully selected. As for the specific proposal of how to cooperate—but first, I think the interests are relatively aligned, that is, those most aligned with our interests, with the least hostility toward us, or those who most hope we can succeed. Not everyone hopes we can succeed, because we still harmed many others’ interests. Company major strategy, technical and business decision processes and decision mechanisms—our company is overall built on the foundation of consensus. I don’t decide all things by myself—rather, I need to seek consensus. My authority within the company and my influence within the company are built on the foundation of consensus.

For example, if I want to do something, I definitely first look at what everyone’s consensus is, whether everyone wants to do it. Then I might have some guidance or inclination, but the effect of this guidance is limited, the effect of guidance is very limited—it’s still built on a foundation of consensus. This decision mechanism is actually a consensus-seeking mechanism—it’s not that I can push something forward; it must be consensus for me to push it through, and then I’ll push it. Currently the main energy is basically all on the DeepSeek side. Let’s rest for five minutes. Let me write quickly. Everyone can unmute, give me some feedback. We’re fine on our end—otherwise let’s rest for five minutes first.

The final effect gap opening up between the various models should be a comprehensive one. Comparing model effects definitely has to be done at the same cost—that’s what’s meaningful. Because when you compare two cars, you also compare cars in the same price range.

A well-made model versus a poorly-made model—this difference shouldn’t be at some specific link; it should be overall.

Anthropic currently surpassing OpenAI—is this long-term? I think this isn’t long-term, this is definitely extreme. OpenAI and Google—in the future they’ll most likely still rise alternately, they should rise alternately. Actually right now Anthropic’s advantage in Code Agent isn’t that large—it’s not that it crushes OpenAI.

Our company maybe has half the people who usually feel OpenAI is better. Actually Anthropic has a first-mover advantage, but this first-mover advantage should disappear soon—it’s not an advantage it can hold long-term. All three companies are very impressive—among these three, efficiency is highest, the cost they spend, what they spend, the money they burn should be the least.

In the global AI division of labor, the role Chinese companies are very likely to play is still being the largest producer. By common reasoning, our production capacity is the largest, including chips—chips probably we have the largest production capacity, we have the most electricity, so our AI ultimately is very likely to be one of the three-body [players].

Chinese people will make this product the cheapest, and then in terms of effect—after all, foreign goods, now many goods where the Chinese-made and American versions don’t have much difference. In the future AI might be like this too, but Chinese-made AI might have a cheaper price. This cheapness might be systematically low, just like in other industries where China provides cheaper services.

When I want to do something, my habitual thinking is: what has the greatest return for me to do at this moment? If I feel doing products has the greatest return right now, I feel I’ll go do products; if I feel realizing AGI first has the greatest return right now, then I’ll go do AGI first. Obviously, I feel doing products right now doesn’t have the greatest return.

If you’re asking about the domestic-card adaptation problem—domestic cards actually have a historic opportunity right now. Because previously domestic-card adaptation had a difficulty called “poor ecosystem.” That is, after buying cards, you couldn’t use them—it didn’t have NVIDIA’s ecosystem. So NVIDIA’s moat is very strong. But this thing is changing. NVIDIA’s CUDA moat is rapidly disintegrating. The reasons for this rapid disintegration may be three aspects. One aspect is that now with AI, after having AI, building this ecosystem is much easier than before, because AI can write code.

[01:56:36] I can use AI to build this ecosystem, so I can build out an ecosystem exactly like NVIDIA’s. First, because there’s AI; second, there are some new technologies. For example, our company released a technology called TileLang, which is a high-level language.

Using this high-level language to write CUDA operators, we can quickly write out NVIDIA’s entire ecosystem all over again, and then combined with AI, this looks like it has no obstacles. But right now it’s not finished yet, not done yet, though this technology path looks like it has no obstacles. There’s another point—because CUDA, NVIDIA evolved from gaming cards, so in many places, the design of gaming cards and the settings of gaming cards are of one piece. CUDA is compatible with gaming cards. Previously because AI computation was a very small field, smaller than the gaming card market, this was reasonable.

But now the compute card market is already bigger than the gaming card market, so there’s no reason for these two to still need to be coupled. The current trend is that they’ll no longer be coupled going forward. So dedicated chips—whether Huawei or NVIDIA itself—will all be dedicated chips going forward, not the previous things.

Under this background, the role of NVIDIA’s original ecosystem is greatly reduced. Because for a dedicated chip, it has no relation to CUDA—it’s not bound to CUDA. Or rather, when this chip is designed, it already considered how to build this ecosystem.

It’s a bit complex, but anyway, domestic AI chip substitution now has a historic opportunity. We believe that within the next year, we’ll be able to see one thing verified: the domestic chip ecosystem has no problem whatsoever. Previously it was thought to have problems, thought to be unusable, not easy to use, but in the future, I think within a year, we’ll be able to reverse this perception, or reverse it with facts.

Domestic AI chips’ hardware and ecosystem both have no problem—the only problem is insufficient production capacity. On this point of domestic-card adaptation, there’s no obstacle—NVIDIA can’t block it. If it were in a normal business environment where I could buy NVIDIA cards, then domestic substitution would be relatively hard; but when NVIDIA cards can’t be bought, everyone is forced into it—everyone has to do domestic chips.

Under this background, domestic-card adaptation has no obstacle whatsoever. Domestic cards building an ecosystem as good as NVIDIA’s, even better than NVIDIA’s, I think has no obstacle, but still needs time.

We currently mainly cooperate with Huawei. Huawei does their own adaptation, but we ourselves will participate in this ecosystem, will deeply participate within Huawei. Huawei’s problem is still insufficient production capacity. Like Huawei gives us roughly 16,000 cards of capacity, internet giants maybe get a hundred-something thousand, we get ten-something thousand—I think this ratio is also relatively... but this is probably just how much capacity Huawei has.

So we also can’t count on training that bigger model on Huawei later, or training a model with several hundred B parameters active—there were some saying this year. But next year, the year after, there may be opportunities.

For Huawei-card adaptation, the main work we’re doing is making its high-level-language compiler well, making TileLang well. After making TileLang well, the problem might be solved readily. This matter is relatively complex to describe, but we’re doing it, and after it’s done it’ll be much clearer to explain. You can understand—when V3 was trained, it still used NVIDIA cards, but it already didn’t use NVIDIA’s ecosystem. V3 used NVIDIA cards but didn’t use NVIDIA’s ecosystem—rather, we first wrote a high-level compiler called TileLang, then based on TileLang’s ecosystem completed everything else, already almost not depending on NVIDIA’s ecosystem.

As long as I take this whole set of things, and redo this process on Huawei cards once, then it’s complete. I think this might be a historic mission—it can completely reverse everyone’s previous perception that domestic-card ecosystems are poor. Right now some favorable conditions of timing and location are basically all there—only time is lacking. I think within a year, there should be many people using it, or everyone understanding that this problem is essentially solved, and the rest is just the production capacity problem.

I’m relatively optimistic about domestic compute. I think on this point, NVIDIA is digging its own grave. Huawei’s supernode, Huawei’s 950 supernode, in performance and price can completely substitute for NVIDIA’s GB200, GB300. The price is definitely more expensive, but limitedly so. Fifty percent more expensive, a hundred percent more expensive—a hundred percent more doesn’t matter, two hundred percent more doesn’t matter.

For example, a hundred percent more expensive—I think it can already be considered a price-level substitute. On tasks it’s also a substitute—all the tasks GB300 can do, Huawei supernode can do, latency and everything the same. The only cost is that four Huawei cards equal one NVIDIA card, and simultaneously lagging by two years. Four-equal-one is understandable. Lagging two years means four Huawei 950s can equal one GB300. Lagging two years means lagging two years in time. Huawei 950 supernode is this year’s Q3 or Q4 product; NVIDIA GB200 is a product from two years ago’s Q3—two years apart.

NVIDIA this year’s Q3 might already have a new generation. So the gap between us and the US on chips—I believe there won’t be a gap on ecosystem going forward, but on chips it’s fourfold plus two years.

There’s another question: will we do vertical integration upstream? I hope not. So there’s a question: will we do vertical integration into upstream applications? We hope not to—I hope others do this thing. I don’t want to gobble up everything—I only want to eat one piece, only the piece I’m best at, or the piece we ourselves consider most core, then the piece directly related to the user...

[02:08:18] ...I think probably for many of our industry partners, they care more about similar... it should be others’ meaning, I shouldn’t interpret it.

Will we build large clusters ourselves in the future? I think building large clusters ourselves is definitely necessary—we’ve always been doing this ourselves, all our clusters are self-built. But whether to develop our own chips in the future, I think depends on how big the returns here are, depends on how big the returns are.

Tesla, cultivating... this matter. Suppose you operate a power plant—you don’t necessarily need to manufacture generators, right? The power-generation equipment can be made by others, and as long as its price is reasonable, why would you make it yourself?

So, I hope not to have to make chips. I hope I can buy chips at a reasonable price, so I don’t have to make chips. I think this is very likely the case—that is, recently NVIDIA chip profits aren’t necessarily... although...

We hope to only do one piece. I think AI is a very big thing, and it doesn’t need me... I only do one piece. If we focus, and I believe the business interest here is already big enough—that is, if the AI era will produce many trillion-level companies, I think we’re one of them.

We’ve always done one small piece of it—us being one of them is already no problem, and it doesn’t need me... I’m not confidently ambitious about doing everywhere. I think it’s most likely relatively hard to change, and if you really want to do other things, you’ll still have more ideas—this will... this is our attitude.

For example, at least in the To B and To C businesses, from what’s currently visible—those who really want to make a To C closed loop, we actually do it better; those who really want to make a To B closed loop, maybe don’t do it as well as us. The more you want... Then also, on To B I can say a few more words. The ceiling of the To B business should still be demand—under the background of the current generation of AGI/AI technology, To B demand should be limited. It’ll grow quickly, but it’s not an infinitely large thing—ultimately still constrained by demand, not compute.

Under current technical conditions, how big the revenue is ultimately still depends on demand. Because think about it, I can recover cost in ten months—I definitely would buy in bulk if there were somewhere to put it, and it’s because there isn’t that much to do. Right, demand should grow larger and larger—if with continuous technical breakthroughs, this demand will grow larger and larger.

Multimodal deployment—we’ve always been doing it. For products it’s very important; for C-end user products it’s very important. But for the intelligence ceiling, it’s a component, it’s not the main line itself.

But as a component, multimodality we’ll definitely do, and we’re doing it too. We’ll probably release related models—that is, our V4, V4’s subsequent versions will support native multimodality. But for us, multimodality—for intelligence, it’s a component—we don’t treat it as intelligence itself, its main line is the same as search.

Search is also a component, multimodality too—our understanding is it’s also a component. Scaling—we believe in scaling, definitely the larger the scale the better the effect, able to unlock more capabilities. What stops us from scaling is actually compute—it’s not that we don’t want to scale, it’s that we don’t have that much compute to do this scaling.

We haven’t touched where this ceiling is. We train models this large not because I feel a model this large is enough—rather it’s that I happen to have this many resources. I calculate based on my resources how large a model I can accept, can train—it’s calculated this way, not that this model is enough.

Currently, model returns are still very obvious. We haven’t yet had the chance to hit this scaling wall—it’s still very far from us. When Silicon Valley talks about scaling reaching its end, that’s for Silicon Valley; for Chinese people, we’re still very far from that—we simply haven’t scaled to that degree.

This scaling includes data scaling, model-scale scaling, and training cost. We’re relatively far from exploring this scaling ceiling—that is, don’t have that much compute. But we’ll also spare no effort to push this scaling ceiling.

Including after we finish financing, we’ll have more compute, and might train larger models. The shorter the time to buy, the better; if we could exhaust all energy in half a year, that would be best—actually can’t achieve it.

The next-generation model’s core capability—I think it must have continuous learning ability, only then can it be called the next-generation model. Before that, what we can do is cost, and making the effect better, making the speed faster. But for a major breakthrough, it should possess continuous learning.

Then there’s another question: why do we seem to especially care about this model’s computational efficiency? Because I’ve indeed found that not everyone cares much about this model’s efficiency. For a commercialized company, there’s no motivation to pursue model efficiency, because slightly higher model efficiency...

So it doesn’t have that much motivation to pursue very high model efficiency. So several startups—they themselves won’t... you don’t hear them say low cost is what they pursue, because that doesn’t conform to their interests. If low cost, what money do you still make? If low cost, you can’t collect much money.

On the contrary, this is part of our vision. Or rather, our vision, our colleagues care about this cost, because our colleagues are all ordinary people, they know using this costs money. They can have this empathy: others have to spend money to use it, so if others can have it cheaper, the degree others can accept will be higher.

So we have many colleagues who still hope we can make the cost lower. But if you speak from a business perspective, you actually wouldn’t think this way. From a business perspective...

**[02:21:31]**

From a cost or business perspective, this isn’t the first priority.

For either startups or big companies, this service cost simply isn’t high. But I think we hope it’s relatively light, I hope it’s something affordable, especially under the background of China’s compute shortage, that it’s affordable, that it can be used on domestic cards.

I think low cost is first of all a result. Our models have indeed always been moving in a lower-cost direction on model architecture—this relates to our vision. We still have many algorithmic methods, and cost can still go lower.

Another reason cost goes lower is: the lower the cost, the more I can train larger models, the more I can afford larger models. On the same compute, when compute is limited, if my computational efficiency is higher, I can afford larger models.

For big companies, they don’t necessarily think this way. For big companies, resources can be added—they can solve it by adding resources. But we’ll prioritize cost efficiency. The value of data-type models—data is a relatively broad range, data should be almost equal to half the model. Why would I feel that if I want that, or suppose I assume AI can occupy twenty percent of GDP, and then if I want to occupy five percent within it—absolutely won’t work. Because I’ll definitely be defeated by another person—if another person says I only want to occupy one percent, then I’ll definitely be defeated by him.

If my goal is that I want to occupy five percent of AI, of all-humanity GDP—theoretically the math still works out. Look at OpenAI—their math seems to work out, theoretically no problem. But they have a problem: they’ll be defeated by another person willing to occupy only one percent. Because another person says, I do it this well, but I only need to take one percent of global GDP—then he’d defeat him. At this point if yet another person emerges saying I only need zero-point-one percent, then he’d defeat the previous person.

Macroscopically, no matter which of these few percent it’s taken from, there’s no difference between them—they’re all the same. Those who take more will be defeated by those who take less. You don’t even actually need to take more—if your vision is to take more, you’ll be defeated by those whose vision is to take less.

Actually no one has gotten money yet—it’s just a vision. If your vision is to take more, you’ve already lost first, you’ll face greater difficulties. This is just how the world is.

OpenAI from the start felt they could truly monopolize this world, but actually they’ll encounter many many challengers. They’ll encounter challenges, so they won’t have it that easy. Challenges the US has encountered—in the future they may also encounter China’s challenge, because Chinese people are willing to take less and can provide this service.

In China, there will also be people willing to take a bit less. But ultimately there’s a balance here, because if you take too little, the company’s business logic doesn’t hold and it can’t survive. So if you take too little, you can’t survive; if you take too much, you’ll be defeated by those who take less.

So for us, we’re not about taking the most profit, or about return-maximizing pricing—rather it’s earning just a reasonable return. This is one explanation.

I believe in this thing—I’m not looking for reasons for this matter, because there’s no need to find reasons. I do it this way to begin with, and I definitely have reasons for doing it this way. These reasons may not be very conventional, but I think the company isn’t conventional to begin with.

Our company’s management actually has two lines: one line is top-down, one is bottom-up. Bottom-up is—each person does whatever they want to do themselves, no one manages them, no KPIs.

Top-down is the formal—we collectively want to do something, needing everyone in the company to cooperate. For example, if we want to release V4, then there has to be division of labor, each person does one part. That’s top-down, and this top-down we call formal.

Generally we hope this formal doesn’t take up more than half of employees’ total time—that is, formal shouldn’t exceed half. They still have half their time that’s unassigned, and they do whatever they want. This is a research scope, letting them explore on their own—according to what they feel is important, they explore, with no prerequisite requirements.

As long as the company can support it—company compute can support them doing it, or they don’t need compute, they need to do very little—then they simply don’t need to coordinate. So this is our current organizational method.

Some people think we’re top-down, some think we’re bottom-up—I think both are right. My one standard is that formal preferably shouldn’t exceed half.

We generally also don’t work overtime much. Overtime has two reasons. The first is that doing research needs a relatively relaxed environment. If you push very tightly, you can’t do research. Because it’s precisely that you need to have this interest yourself, you need to think about these problems in your spare time, so it has to be in a relatively relaxed environment for there to be a chance to explore. This is a need arising from research culture.

The second is that we’re very focused. Being very focused means the things we need to do are few. Then I don’t have that many things to do, so I don’t need to work overtime.

This is of one piece with the earlier restraint. Because I’m restrained, so many times I just don’t do it. Then the things I need to do become fewer, and the work each person gets assigned becomes less. You see many of our products are incomplete—we haven’t gone to fill them in.

This is also a culture of ours.

OK, because there are many questions, I’ve mostly scanned through them. Everyone, if you have any more questions, please ask.

“Please, investors, feel free to unmute and exchange. Let me remind you slightly—Mr. Liang talked about a lot of relatively sensitive information, please absolutely don’t spread some numbers or situations externally, including card quantities and such. Also absolutely don’t screen-record and share externally. Thank you all very much—if you have questions, please feel free to unmute and exchange.”

[02:34:59] “Brother Yang, could you share more about the timeline for when continuous learning will bring a breakthrough? Including if continuous learning is realized, what other architectural and algorithmic innovations are needed, and other key elements?”

Few people, research is needed. Right now the whole world is researching this problem. Or rather, for investors, what investors see most now is AGENT; but for us research people, what we now see more of is learning, and how to solve this learning problem.

Actually learning may not be a single technology—it’s a problem. How to solve this problem may have many technologies—not one technology, it’s not one thing, it’ll be many things. Or rather, AGI is composed of many things, needing models, and then needing many other things.

Actually it’s also an engineering and algorithm problem. This problem is relatively specialized, but there are many methods, and many research efforts.

“Mr. Liang, thank you. Thank you very much for this opportunity today. First, I really want to respond to and express gratitude—what you said at the very beginning moved me greatly and gave us much inspiration. You mentioned that this team carries the greatest goodwill, hoping to, within this industry, push forward society and the development of human intelligence, making a little contribution.”

“And within this, carrying this sense of mission and vision. I think this is very close to the corporate culture of the company we serve, which is ‘cultivate oneself to benefit others.’ I deeply understand why you would lead the team to do open source.”

“I would vividly feel it’s like us building a banyan tree, a little-birds’-paradise ecosystem—benefiting all things without competing, but this way it’ll be accepted by everyone, all things coexisting, and ultimately it’ll be everywhere. So we’ll use this investment to likewise express our recognition, support, and respect for this mission and vision.”

“At the same time, we also hope to contribute some strength in the future of this industry, in some areas we’re good at. On this I also want to continue consulting and discussing with you. For example, in future ecosystem co-building, now after open-sourcing, how many partners, talents, and teams in this industry can relatively well reproduce our currently open-sourced models and results?”

“In the future, when taking the next step to further develop this ecosystem, in which areas do you feel more high-quality talent is needed to connect to our models and reproduce them? Or is it that GPU compute is now relatively scarce? Will everyone in the future use a kind of model-matrix approach?”

“For example, we make the large base models better and better, and partners and teams in all industries make some vertical-industry models or application models within the model matrix. How is this developing currently? In the future, step by step, two years, three years, what kind of ecosystem do you feel it’ll grow into?”

“This is the first thing I want to consult you about. Second, you just now shared with many partners many observations about AI hardware. Like AI global big companies, each might make hundred-billion-dollar-level investments, and China currently seems to have some shortcomings in hardware and compute.”

“How long do you feel this can be solved, to support our AI development, so hardware and compute shortcomings don’t hold AGI back? Do you feel this is a mission Chinese people must accomplish in the future—we can definitely make it, it’s just a matter of time and capital investment?”

“But at the same time, it may be two-sided. On one hand, the progress of models will improve model intelligence, causing single tasks or certain intelligence units’ consumption of hardware and compute to gradually decrease, no longer needing such large compute operations, because model progress will make it compute cleverly. I don’t know if I understand correctly.”

“On the other hand, hardware technical progress will make compute’s efficiency more powerful. Will this be a path where the two sides move toward each other? Right now, at this point in time, using the current 960, or H200, to make hundred-billion-dollar-level compute investments—you just mentioned you amortize it over three years, so what’s its actual lifecycle? Do you feel technical iteration is four or five years?”

“Or put bluntly, will it be that a compute center now builds a ten-thousand-card cluster with current cards, and maybe three years later it’s actually relatively not-so-advanced compute? Will there be a situation where the current stage is under construction, not enough, and three years later becomes relatively not-so-high-quality compute with redundancy? I don’t know if there’d be such a phenomenon. These two questions I consult you on, thank you.”

Thank you. The first question is the ecosystem question.

We currently feel that maybe the problem every company faces is insufficient talent. But I think this talent shortage will be phased. In the early period of every industry’s development, talent is insufficient.

Including previously making websites—when websites were first being made, there were very few people making websites, talent was very scarce. Later the internet needed server-side work, and talent was also very scarce. But this kind of talent shortage is very quickly resolved—just two or three years, because large amounts of people are cultivated.

The AI talent shortage is also phased, and we’ve already seen it substantially alleviated. Because AI people truly aren’t lacking—every company will quickly cultivate people, cultivating people is fast. So the AI industry overall—whether ecosystem, model companies, or whatever—talent isn’t lacking.

Talent scarcity is definitely a short-term phenomenon. Historically there’s never been a long-term shortage of a certain type of person. I still remember over ten years ago it was said pilots were very scarce—pilots’ cultivation cycle is very long, but it was also quickly resolved.

So everyone doesn’t need to worry about the talent-shortage problem. Also, there are currently a bit too many companies domestically doing models—still too many. The US maybe has just three; China has too many things doing base models.

Ultimately there definitely won’t need to be that many people doing base models—it’ll definitely converge. So resources are also relatively scattered, and to some degree relatively wasteful. It first...

[02:44:00] Every company has to do the same thing, but the US has just three companies doing it, resources concentrated only in these three. China’s resources are very scattered—each company gets even fewer resources. I think this will definitely converge—definitely—but it needs a process, and ultimately it will definitely converge.

Don’t need that many companies, because right now maybe everyone feels the profit margin of doing this thing is very high, so they definitely want to do it themselves. But when they discover this thing might not have such high profit, they might stop doing it. Recently there definitely isn’t such high profit—I don’t believe there’s such high profit, because it doesn’t conform to objective law.

This means we’re at a stage: if there’s a very high profit margin, this definitely doesn’t conform to objective law. We should have a reasonable profit. So this is a current state of the industry—I think it will definitely converge.

That is, everyone doing the large-model part—among them, don’t say some company monopolizes, saying “I want to take away all the profit,” this definitely won’t work. If each company only takes a reasonable profit, then actually there’s no need for that many people to do large models. China ultimately having three or four companies competing—competition is already very sufficient, prices absolutely already enough for a price war.

Large models—maybe not to say two big companies, two small companies—that’s probably already quite enough.

As for the ecosystem, I don’t have too many thoughts. We hope to support more people, but we don’t have that much energy. We have this willingness, and there won’t be interest conflict, but whether we actually do it is another matter.

But at least there’s no interest conflict here—we hope for win-win cooperation. First, I absolutely don’t think large-model companies can take away most of the profit—this is impossible, because with so many large-model companies, the gaps aren’t that large now.

The gap is only two things: one is time, one is cost. So it won’t reach any company having windfall profits—I think it won’t reach windfall profits. Those who control cost well earn a bit more, those who control cost poorly earn a bit less—that’s simply all.

Is that an answer? Was the first question answered?

“Everyone can... you believe that in the future there will definitely be many people who can... in the future actually... everyone’s data-application cycle iterations...”

“Can you hear now? Thank you. Right, thank you for your answer, and I very deeply understand and respect your ecosystem-strategy positioning in this industry. For example, on the data side, publicly available data now—I believe model companies all already have channels to obtain it, this method should be no problem, just a matter of time and cost.”

“Then subsequently, for example, when truly reaching AGI, one setting or ideal state everyone imagines is that the model can self-iterate, self-learn—that is, train itself. So for this part, currently on data, do you feel simulated data can be put to use, or is real data’s quality highest?”

“If it still needs to come from real data, will that limit AI’s intelligence to still being within humanity’s past... this level, because it depends on the real data humans truly once had? Or can this ceiling be broken through, letting the model’s capabilities surpass all humanity’s past real [data] through simulated data, synthetic data, created data, and so on?”

I think it can be surpassed. I think there are two points of surpassing—for example, Go, AlphaGo played a move humans had never seen. That is, it’s definitely surpassing humans within a certain range.

But it may also have a ceiling, may also have limitations. But this limitation we currently can’t see. We believe, so broadly speaking, it can surpass based on knowledge humans already have, that we can already articulate.

“Then subsequently, is this by real data or simulated data? Will it work?”

There can’t help but be many methods.

“Okay, thank you. I’ve also taken up your time. I’d also like to continue consulting on the earlier AI Infra question.”

The second question was what? A bit...

“Alright, let me quickly repeat it briefly. I want to consult—for AI Infra, in the future I believe everyone is investing at the hundred-billion-dollar level in compute. So on this side, maybe we believe Chinese people in the future are mission-bound to accomplish it on hardware, that one day there might be high-efficiency compute, but maybe in the practical process, currently it’s still a constraint.”

“So in the future, will it be two directions moving toward each other? One aspect is that after model capabilities iterate, they actually go from ‘brute compute’ to ‘clever compute,’ so per-unit-model or per-task compute requirements and consumption gradually decrease marginally. On the other aspect, hardware like cards’ capability improving will make iteration speed faster and faster, single-card efficiency improve.”

“So what kind of phenomenon is this in the process? Will it be that building a ten-thousand-card cluster now, buying H200 or 960, but after two or three years, it becomes relatively not-so-high-quality compute, relatively becomes an outdated component?”

NVIDIA cards you can basically depreciate over five years. Huawei cards at most depreciate over three years. Huawei 950—this year using it is pretty good, next year using it I think is still okay, later using it I think might really be too power-hungry.

Huawei cards’ lifecycle will definitely be a bit shorter, because it’s already two years behind NVIDIA to begin with. But I think the gap isn’t that large.

If you could buy however many B200s right now, I think it’d all be worthwhile. Suppose for Tencent, if Alibaba could buy them, then look at the quantity; if they can be bought at a reasonable price, it’s definitely worthwhile. But when it’s not a matter of calculating cost—I believe they can’t be bought. Understood. Our compute lagging is a fact. This fact is dissolved through three aspects. The first aspect is we accept model lag—we can only use smaller models than they can, and how much smaller is a training question. We have to accept a certain model lag, and smaller model sizes.

This lag has a benefit—lagging means you have more time, you have certain technologies, and this way you can use clever methods...

[02:53:59] Thank you.

So the gap between us and the US might be lagging the US by 12 months, lagging the US maybe 12 to 18 months, or 6 to 12 months. Anyway, simply put, lagging the US by two years, then using only one-twentieth of the US’s compute to accomplish this thing.

This narrative is lagging one to two years, but using only one-twentieth of its compute. So in the future we want to rewrite this narrative—that is, we use one-nth of its compute, but shorten this time even more, shorten it to 6 months, 3 months—I think this is a goal.

And we can even surpass them in certain aspects. But under the situation where total compute still has an order-of-magnitude gap, comprehensive surpassing is unrealistic; but in certain key, trade-off-selected places, us surpassing in some areas might be possible.

Understood, thank you Mr. Liang. Full of confidence, let’s work hard together. I’ll leave time for other partners, thank you.

“Thanks for the earlier sharing. I have two quick technical questions here. In the earlier discussion of technical roadmap, you mentioned the core problem to solve at this stage is continuous learning, which is also a current research hotspot abroad, called Recursive Improvement. I’d like to ask—currently, what’s the biggest technical difficulty? From your perspective, when can this be solved? This is the first question.”

“The second question—you also mentioned first solving continuous learning, then doing intelligence. I also want to understand—what’s the technical root behind you framing it this way? Does it mean that after solving the continuous learning problem, DeepSeek will subsequently also do general intelligence? I consult you on these two questions.”

The technical question is actually a bit... [hard] to explain—its difficulty lies in that we currently still haven’t found a method that really works. The whole world still hasn’t found a good method—everyone is groping. So right now it’s still at the groping stage—that is, we don’t know who can grope to the next method to solve this problem.

Right now it’s still at the exploration stage. We have many ideas, many ideas that currently look promising, but none have been made to work yet. Right, that’s the first one.

The second is—right now internally we relatively value this narrative: training our next-version model, we hope it can help our own development. It can improve DeepSeek’s efficiency—our model first and foremost improves DeepSeek’s own work efficiency, letting it provide more help when we develop the next-version model.

Or put simply, the model we make—the first goal isn’t for everyone to find it useful, but for us ourselves to find it useful. First it’s useful for ourselves. After it’s useful for ourselves, when I develop the next-version model I’ll be faster.

Many people internally think this way: first it should be useful for ourselves, first it’s for us ourselves to use. Then this is the fastest method to realize AGI. When it’s useful for ourselves, that means others might find it useful too, but first we have to ensure we ourselves find it useful.

This narrative is a bit odd, but indeed many people think this way. Rather than users finding it useful, I hope it helps ourselves more, so we can realize AGI, will be much faster. So the logic of this narrative is: to help us realize AGI.

But first it’s to help us realize it. Realizing AGI needs this help. Right now it’s very certain—we indeed need AI to help us realize AGI. Although it doesn’t yet work autonomously, it’s still just paired with humans, but it’s already very useful.

The second question—just now it was mentioned, first solve the continuous learning problem, then enter general intelligence. Is this an expectation of yours for what follows? I want to understand what the technical root behind this is. Why must continuous learning be solved first, then enter general intelligence? Your understanding of this subsequent part.

Because solving the continuous learning problem can greatly accelerate our R&D progress. If I first solve the continuous learning problem, then the general intelligence problem is no big deal. I’d have AI’s assistance, and if AI can learn continuously, its capabilities should be very strong.

The current Agent’s capabilities are limited because it can’t learn continuously, it can’t effectively learn continuously. If continuous learning could be done first, then AI’s capabilities are very strong—it can greatly improve the efficiency of our own research.

If continuous learning is made first, general intelligence might become very easy—using it to do it becomes very easy. So I say this is a result we relatively hope to see—we’re relatively effort-saving, we’re relaxed. Otherwise, if you were to manually do general intelligence now, it’s relatively tiring, relatively bitter, a data-intensive, manpower-intensive matter, and cost-effectiveness isn’t high either. Thanks for sharing.

A small question—the online questions, take a look at the chat group. “How much longer until AGI? Can domestic hardware catch up in this time?”

In the Zoom meeting’s chat window.

Okay, I saw this.

Huawei 950—right now Huawei gives us 16,000 cards, this should be publicly stateable. It should be an order of magnitude less than internet giants. Huawei can only give us this much, because this price isn’t cheap either.

Internet giants’ pursuit is bigger—for internet giants, they need it more. For us, we can buy some non-compliant cards. So our purpose in buying Huawei 950 is still hoping to help Huawei build this ecosystem well.

16,000 Huawei 950 cards only equal 4,000 B-series cards. So it’s not a very large quantity, the significance isn’t very big. Not enough to train a next-generation model—it’s only enough to train our current generation model, not enough for the next generation. But it can let Huawei do this well first—that’s about Huawei 950.

Then, how much longer until AGI? Can domestic hardware catch up in this time?

I think on this AI endeavor, on this AI matter, domestically in one or two years we should be able to reach roughly the same as abroad, or maybe this year we can achieve substituting for foreign models. On this AI matter, with the current approach, current paradigm, it’s not a very hard thing, so this year we should be able to achieve it. But it’s not yet AGI.

[03:05:48] At least I feel it needs to be able to learn continuously.

Can domestic hardware catch up in this time? I think domestic hardware might need a few years. Domestically we first need to solve the ecosystem problem, because ecosystem is a confidence problem. Solve the ecosystem problem, then solve the production capacity problem—I think it should be able to be gradually solved.

I don’t quite believe that five years from now, we’ll still be stuck on the production capacity problem. Right now we’re definitely stuck on the production capacity problem—this year, next year, the year after, I think we might still be stuck on the production capacity problem, but five years later, I think maybe not necessarily—I’m still relatively optimistic.

Then the second question is thinking about future organizational structure, planning the scale of personnel. First, our previous organizational structure was very decentralized, because there was no organizational structure. But after our people expand further, these definitely need some changes made.

Right now I can only say there will need to be many changes made here, but right now it’s hard to express it all at once. Ultimately there should still be different departments—some departments we need to build a relatively rigorous hierarchical structure for. Other departments might still maintain a relatively loose, relatively flat structure.

As personnel increase, we’ll make this adjustment. It should be that we have to make this adjustment right away, because I’m already making this adjustment. If this adjustment weren’t made, many things couldn’t be pushed forward. There indeed are many departments that should have organizational structure.

Then everyone still has the question—CV is prior to which version, right? I think right now, we released this GCV4 version online—it’s still relatively rough, and in terms of capabilities there’s still a lot that needs time.

Generally, for me, a relatively comfortable release cadence is roughly releasing one version every two or three months. The last release was maybe late April, so the next release might be late June—roughly like this. If there are no surprises, each version should still be better than the last.

At the 50B-active scale, I feel ultimately we won’t have too much difference from the current wave of open source. In inference speed and effect, I feel there might not be too much difference.

But compared to their large-size model—their one unreleased model—the gap should still be relatively large. That gap should be that our active parameters can’t achieve it—that definitely has to be a larger-size model, maybe say 150B.

So 150B—I think based on our current training progress, this year, optimistically by year-end we can start training, at least next year April... the gap is still relatively large. Right, this is the gap with OCE.

“Hello, I actually have a question—you often mention this AGI, that its realization process is a gradual process, rather than there being a mutation. So can I understand this as a process with no critical point?”

It has no critical point, but it’s nonlinear. We currently relatively believe a narrative that AI can accelerate AI research, AI can accelerate AI research. That is, it’s not linear, because you can use AI to accelerate your own research, so it might be nonlinear later on.

Understood. So currently your—I understand the earlier conclusion might be that continuing to scale language models is enough, is sufficient, to reach this state.

I can only say language model scaling—I currently don’t see a ceiling. Our current intelligence level, or the intelligence level reaching the US, neither has seen a ceiling.

Understood. Because I’m very curious—you said, for the US, this 800B-active model, they can train it but they also can’t afford to use it, so they can only train it out—it’s also hard to bring out for everyone to use, because it’s indeed a bit expensive.

What I was actually curious about earlier is—humans actually having language ability is only a matter of the last hundred thousand years, and before that they maybe evolved for 3.7 billion years. But on this matter of training AI, maybe it’s possible—you say that order can be reversed—but ultimately it might still enter into so-called, not necessarily called world models, still entering into physical models or the embodied part, right? That is, after this ceiling.

Right, I think embodiment definitely still has to be entered—ultimately embodiment. So for our company, naturally, the endpoint might all be embodiment. Because for a normal person, their needs aren’t computers, right? Because normal people—their eating, drinking, entertainment, clothing, food, housing, transportation—they don’t need computers.

What they need is—so they still need embodied intelligence to solve concrete human-labor needs. If the goal is to relieve human-labor needs, then embodiment I think is unavoidable.

Understood. So speaking in stages, suppose reaching something like—not necessarily a critical point—that is, can self-evolve, can self-evolve relatively well, or approach that... the rising to that time point, I’m very curious what state of AI you’d hope lands first is...

Might be different from now. We hope it can—that is, directly, suppose there’s no embodiment, then our definition of AGI, or what we hope AGI can do—it can help me iterate the next-version model, help me iterate the next-version model.

Then after there’s embodiment, what we hope it does is likewise—let it iterate the next-version embodiment, let it make the next-version robot.

I’m still very curious about one point—because I read previous DeepSeek interviews and such, it should be that in some important direction choices and research choices, taste and intuition are very important, rather than simple engineering optimization. So if AI can self-evolve later, are these things—taste, taste, intuition—still important, or what will be important?

AI currently doesn’t lack taste and intuition—what it lacks is continuous learning ability. AI’s taste and intuition are no problem. If you have it write an article, its taste and intuition—I think there’s no problem.

There are still several questions earlier—let me look at those questions. I saw a few questions discussed on screen, but I can’t see them now on my end.

“Mr. Liang, I left a question on my screen, let me read it to you again. Actually I’d like to consult—continuous learning, as you also mentioned, many researchers also mention it’s a still-unsolved research problem. Then the coding agent, especially catching up to MILES, that is...”

[03:17:47] “...Scaling to this Office level, MIS level is a relatively definite goal. For a still-unsolved research goal and a relatively definite scaling goal, how do you think research resources should be allocated, especially the talent resources of researchers, to achieve the best balance and effect?”

Model Office is a relatively definite goal, but model MIS—I think it’s not yet certain to say it’s definite, can only say it’s a goal. Model Office, I think, should be relatively definite.

CoT doesn’t take up resources. Doing this kind of free-form research doesn’t consume cards, needs only very few cards—what it needs is ideas; it also doesn’t consume talent resources, because it doesn’t need someone constantly doing it. It’s not a project—rather it needs many people all thinking about this problem.

So there’s no need to allocate any resources to it, because it doesn’t need resources. To train models, to make models, release models, do model-efficiency experiments—that needs resources. As I just said, the resource consumption on people and cards is both low.

So we call this “drawing lottery tickets.” The threshold is very low—anyone can go draw, but who can draw out something—this maybe I also don’t know whether it depends on talent or what.

So there’s no need for us to allocate resources here. It’s just that where we differ from other companies is that we’ll spend time discussing this problem, will think about this problem, and then treat it as an important matter.

In the company, it’s an important problem, a problem we’ll spend time thinking about, but it doesn’t need much resource to do.

Then below I’ll look at one more question: large models’ hallucination problem relatively affects user experience. The hallucination problem also has a method that can solve it, but this is a long proposition. The hallucination problem can be considered something solvable through better post-training—it’s a solvable, improvable problem.

It’s just that everyone hasn’t put much effort into it. Or rather, for me, hallucination is a problem, but we probably attribute it to a product problem. We’ll solve it, but it’s not a key problem.

Earlier there’s also a data-labeling problem. Our data annotation—this relates to our capital investment. With our capital-investment structure, we can’t support the cost of that much high-quality data annotation, because the cost is very high.

US data-annotation cost and China data-annotation cost don’t have much difference. Labeling data in China doesn’t have a cost advantage—especially on labeling high-end data there’s no cost advantage—which makes it very hard for us to invest in labeling data like the US. This path in China is very hard, because labeling data is simply too expensive—whether we outsource labeling or label ourselves, both are very tough.

So right now it’s basically walking on two legs. It’s not that we completely can’t label—rather it’s that labeling data has some low-cost, some high-cost. We label the low-cost ones first.

So you can also consider—right now half the people in our company are labeling data. Half of the core researchers, the most important people, half are labeling data. We just concentrate on labeling data. Solving AI at this current stage relies on labeling data. You just need to look—it’s about data.

“Mr. Liang, thank you. You spoke especially—we feel it very much, very good. Thank you. I’ll consult a few questions.”

“First question—you just said Chinese models are definitely stronger than US models in efficiency. You also mentioned some other aspects where in the future we might also be stronger than the US. Which aspects do you feel we’ll be stronger than the US in intelligence or other aspects?”

I think in many experience aspects, it’s possible we can do better than the US. That is, not to speak of our own experience—we feel our own product experience is pretty good—I think in user experience we won’t necessarily be worse than the US.

In products, in product capability we won’t necessarily be worse than the US. Cost should also be lower than the US, so it’s possible China will still have competitiveness.

Other aspects—if you ask whether there are structural advantages, I think maybe there aren’t. But in cost and product, these two aspects, I think there indeed are certain structural advantages.

Cost is easy to understand—it’s because they don’t need to do it, so they don’t develop this capability. They definitely don’t value this thing as much as we do. We can treat it as a very important matter, but for them, this is unimportant.

Products too—natively many companies’ product capabilities are still okay. So I think these two aspects might have structural advantages.

Okay. Second question—I consult you—you just mentioned post-training, our investment relatively high cost, like Anthropic and OpenAI both invest huge amounts. After this round of financing, do you feel we’ll increase investment in post-training?

The gap is mainly on high-quality data annotation, then mainly in AI research. We’ll definitely increase investment, but like high-quality data annotation, typically that’s not capital investment.

The bottleneck of high-quality data annotation, I think, is time—that is, needing time. Because for OpenAI, for abroad, for Anthropic, they’re all earlier, and have more capital, more cards too.

Under this situation, domestically we can consider it as having just started in the last half year, so in time, I think more time is needed. This has little relation to capital investment, because even without more capital investment, the original capital is enough for it to expand at the fastest speed.

But this speed has a ceiling—the bottleneck isn’t that I can immediately have more people, it’s not stuck on money, nor stuck on cards. But it’s indeed in a rapid-expansion process.

So we feel within a year, the high-quality data problem done relatively well—I think domestically it should be expectable. I think the outlook might not be that cold, but it indeed needs time.

Thank you. Then the third question—we’ve seen Anthropic use their own model to make their own products, launching many vertical finance, legal...

[03:29:17] “...even in the future wanting to move toward the medical direction. Do you feel that at some stage we’ll also consider this kind of vertical application?”

I haven’t yet thought clearly about what our domestic business model will be in the future, or what the smoothest path is—we haven’t reached that stage. Domestic and foreign situations aren’t necessarily the same—what it’ll be like domestically at that time, I think is relatively hard to judge right now.

Domestically, looking at the current situation, I think the most reasonable approach should be to go all-out on general Agent, and other Agents’ priority should be lower, including finance, doctor and such Agents. Do Coding first, because Coding Agent can accomplish a lot, and there are many vertical Agents. At the current stage, what we feel is most important should still be Coding Agent.

Explained very clearly, thank you. There’s one more question we’d like to consult you on. Actually we do DeepSeek, and we also admire you greatly—you’ve always done DeepSeek in a very pure scientific-research manner.

But now this industry has indeed also walked onto the capital markets, walked onto the capitalization path. You’re also a very responsible person—whether to partners or investors, you’re all very responsible. So, between the direction of pure scientific research, purely doing AGI, and the capital markets, the balance—you’ll definitely still go to the capital markets in the future, definitely still have public-shares channels—how do you feel about balancing this going forward? How do you consider this?

I currently feel it should be achievable—that is, both. Suppose this year I can have several hundred million dollars of B-end revenue, plus we have C-end users, then this itself already has a certain commercial foundation. With next year’s B-end revenue, if this demand can grow further, the company isn’t far from net profit—it might already be net profit.

It might already not be a pure cash-burning stage, so I feel the actions we can take, can operate afterward should still be relatively large. Or rather, in the worst case, selling API might already be able to support a public company. That is, if there’s no new technical progress later, if our technology is frozen here, then ultimately we’d go all-out selling API, do these services well—I think that’s enough too.

So I think there’s still confidence—it’s indeed not that hard. Because it’s indeed in a high-leverage place, and in a very fast-developing field, it might just not be that hard. I can only say we hope to have bigger dreams, but we also have a floor of performance we can bring out.

Thank you, well spoken. Questions are done, thank you.

“Thank you for sharing. Some questions mentioned earlier, I’d like to ask you again. Because I think DeepSeek’s biggest difference from other companies is that our organization is different from others, or our organizational form is different.”

“But the organizational form relates to our goal, and maybe also needs to consider the organization’s own boundaries and efficiency. I don’t know—from a macro perspective, does our organizational form have a good learning object? Historically maybe Bell Labs, or what kind of form might be relatively ideal? Or do we ourselves feel there isn’t an ideal one, and it more needs us to gradually explore?”

Maybe in a new era, we can only rely on ourselves to explore. Because our organizational form is definitely different from the several US companies, right? The three companies are also different from each other, but they at least explore from the angle of a commercial company.

I mainly want to ask again about this organization question.

First, we have no imitation object. Every step is us proceeding from the actual situation, seeking truth from facts, making decisions according to the actual situation, finding out how we should do it. So it’s a product of the times, or a reflection of the actual situation—it’s not a result of imitation.

That is, in this situation, the optimal solution I feel might just be this way, or I myself believe the path we chose ourselves is this way. Every step we’ve definitely thought through, definitely chosen—anyway the choice result is chosen this way. It’s not that because we saw someone choose this way, we chose this way—rather it’s because we analyzed the pros and cons and chose this way.

So in the future it might be the same—we haven’t imitated anyone. I think we’re still different from Bell Labs, because it explicitly didn’t need commercialization, because it was a... But we explicitly need commercialization. We ultimately still need to survive—we’re a company after all, the government won’t give me a single cent.

So we can have a very grand mission, but at bottom we’re a company, we need to consider how to survive. So B-end is definitely important for us, because maybe in the future we’ll need it to survive. It’s just that right now it’s unimportant, because right now it’s a cost line.

So I think this is still different—still different from Bell Labs. We’re essentially still a company. Historically there have also been many companies that have pursuits beyond profit, but you can’t say that because it has pursuits beyond profit, it’s not a company.

Many companies are great precisely because they have a pursuit beyond profit. That pursuit ultimately not only didn’t affect its commercialization, but rather let it commercialize better. We’re essentially still a company—it’s just that in considering which money to earn, when to earn money, how much to earn, what to earn by, we have trade-offs.

“Mr. Liang, I have two small questions to quickly consult you on. One is—you actually mentioned MILOS earlier, that it might not be a very definite goal, but you’ll definitely work toward this direction. Then you also mentioned active parameters might for example...”

[03:41:02] “...the next generation might be around 150 to 250B. In this situation, do you feel 150 to 250B benchmarks against O4.7, or might benchmark against something else? This is the first question.”

“Second question—you also mentioned earlier that we now use some compilation languages other than CUDA on the whole inference side. I understand originally we were based on NVIDIA’s ecosystem, maybe using PTX and such more. Now using more of things like the TileLang you just mentioned—will it substantially reduce some of our efficiency on inference, or reduce efficiency short-term?”

“I don’t know how you’d view the efficiency loss brought by this change in compilation language, or whether long-term it’s actually a supplementable state, improving efficiency?”

It improves efficiency, substantially improves efficiency.

“OK, so there’s no negative impact at all?”

Right, it substantially improves efficiency. So this is an opportunity. It’s equivalent to—previously you couldn’t leave CUDA’s ecosystem, now we can abandon its ecosystem and use a simpler method, just use TileLang. It’s a high-level language, writing those programs is also very fast, the code you need to write is very small now, I can rewrite it all over again.

Understood. So both of these are actually one—like the relatively big opportunity brought by AI you mentioned, not a shortcoming that maybe needs to be short-term made up for.

Right, it’s a big opportunity of technological development. It’s not AI’s opportunity—because we also have a project, we’re using AI to write TileLang.

Is that faster?

Right now all TileLang is human-written, but it’s already much faster than originally writing CUDA.

Understood. So even for the hardware-underlying execution efficiency, you feel there’s also no impact?

Losing 1% to 2%, I think is acceptable.

OK, understood, got it. Alright, thank you, very clear.

Does anyone else have any questions?

No questions, let’s wrap up here today.

Okay, thank you.

Thank you all very much for your time, thank you Mr. Liang, thank you.

Bye-bye.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @liang wenfeng 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/deepseek-s-liang-wen…] indexed:0 read:97min 2026-07-23 ·