cd /news/artificial-intelligence/20vc-are-openai-and-anthropic-overva… · home topics artificial-intelligence article
[ARTICLE · art-66661] src=vuci.ai ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

20VC: Are OpenAI and Anthropic Overvalued? The Open-Source AI Reality | How Token Costs Will Fall 10x And Usage Will Explode 100x | The Future Is Not One AGI; It's Millions of Specialised Models with…

Fireworks AI hit $1 billion in annual recurring revenue with 200 employees by betting on millions of specialized AI models rather than a single AGI, and founder and CEO Lin Qiao predicts token costs will fall 10x while usage explodes 100x in three years. The company raised $1.5 billion at a $17 billion valuation and processes over 40 trillion tokens daily, with the vast majority from customized models. Qiao argues that open-weight models are structurally more viable than proprietary frontier models, warning that AI companies face 'scaling to bankruptcy' if infrastructure costs outpace revenue.

read18 min views5 publishedJul 20, 2026
20VC: Are OpenAI and Anthropic Overvalued? The Open-Source AI Reality | How Token Costs Will Fall 10x And Usage Will Explode 100x | The Future Is Not One AGI; It's Millions of Specialised Models with…
Image: Vuci (auto-discovered)

Fireworks raised $1.5B at $17B valuation Fireworks AI raised $1.5 billion at a $17 billion valuation, a remarkable outcome for a 200-person company.

Fireworks AI hit $1B ARR with 200 people by betting that millions of specialised AI models will beat one AGI — and Lin Qiao says token costs will fall 10x while usage explodes 100x in three years.

The Twenty Minute VC (20VC): Venture Capital | Startup Funding | The Pitch

Fireworks AI hit $1B ARR with 200 people by betting that millions of specialised AI models will beat one AGI — and Lin Qiao says token costs will fall 10x while usage explodes 100x in three years.

TL;DR

Lin Qiao, founder and CEO of Fireworks AI, makes the case that the future of AI isn't one AGI ruling everything — it's millions of specialised models, one per application [1] — Lin Qiao "The AGI believers assume one model will solve everything. Lin Qiao thinks that's both technically wrong and philosophically depressing. The…" 08:35 . Fireworks has hit $1B in ARR with just 200 people and processes over 40 trillion tokens a day, the vast majority from customised rather than off-the-shelf models [2] — Harry Stebbings "$1B ARR, 200 people: Fireworks AI reached $1 billion in annual recurring revenue with only 200 employees, scaling in roughly 4 years." 00:55 . Lin predicts a 10x reduction in token costs over three years will unlock 100x more usage [3] — Lin Qiao "10x cost drop drives 100x usage: Lin Qiao argued that a 10x cost reduction in AI tokens will trigger a 100x explosion in usage as AI become…" 44:50 , and argues every company will eventually own its own intelligence stack the same way every company owns its own software stack. The single most useful takeaway: don't wait for AGI to solve your problem — own your model now or risk losing competitive differentiation.

Lin Qiao, Co-Founder and CEO of Fireworks AI, discusses why the company bet on inference over training, the open-source AI revolution, enterprise trust in Chinese models, the multi-model future, token cost trajectories, and whether Fireworks will eventually need to build its own data centres.

The episode opens with a punchy teaser from Lin Qiao before Harry Stebbings delivers a rare investor endorsement: a $10 million check written after just a 15-minute meeting. Harry breaks down the five reasons — a world-class team, a fast-growing inference market, triple-digit ARR growth to $1 billion in four years, the ability to hire stars like ex-Salesforce President George Hu, and the sheer upside potential of a company he thinks could reach $500 billion. The framing sets the episode's bullish tone before three sponsor integrations (JPMorgan, Navan, Base44) round out the opening block.

Harry asks the valuation question directly: if 90% of enterprise workflows can now be handled by open-weight models at a fraction of the cost, are Anthropic and OpenAI dramatically overvalued? Lin agrees the market is beginning to realise this, then introduces a more uncomfortable concept — 'scaling to bankruptcy.' In the SaaS era, finding product-market fit was the hard problem; scaling was cheap. In the AI era, these are decoupled. Companies with genuine customer demand and willingness to pay can destroy themselves by growing because the AI infrastructure COGS spiral faster than revenue. This creates a structural pull toward open-weight models that companies can control, customise, and optimise.

Harry — declaring a conflict of interest as a Lagoora investor — uses the Harvey-Lagoora dynamic as a test case. A year ago, not building your own model seemed fine because frontier models were improving so fast. Now the picture is reversed. Lin reframes the question: in legal, where error tolerance is near-zero and workflows are deeply specialised, the competitive advantage isn't whether you own a model but whether you've encoded your proprietary knowledge into the orchestration layer and the fine-tuned models that power it. He points to Cursor as evidence — they were the first coding AI company to tune their own models, and almost all coding companies have followed.

Harry asks whether the relentless pace of model launches is sustainable. Lin separates two layers: base general IQ advancement, which comes in occasional large step functions (like the introduction of chain-of-thought reasoning in early 2024), and specialisation, which builds rapidly on top of each step function. As base model quality improves, the number of specialised branches that can be grafted onto the trunk multiplies. If anything, Lin expects the specialisation wave to accelerate faster than general intelligence progress — a view with direct implications for Fireworks' market opportunity.

One of the most technically rich chapters in the episode. Lin explains that early AI tech adopters are hackers who want control, and Cursor — flush with frontier lab researchers — is the prime example. The problem they solved together is profound: hyperscalers run RL training on 100,000 interconnected chips. Cursor and Fireworks had no such cluster. Their solution was to decouple the trainer (which updates model weights) from the RL rollout (which deploys the model into a synthetic environment to collect rewards), distributing the system across five or six global data centre regions and syncing fresh model weights efficiently enough that the reward signal stays numerically sound. This distributed design enabled Cursor's recent model launches without a hyperscaler-level budget.

Harry opens with the customer concentration question — can Fireworks survive a Cursor churn given a potential SpaceX acquisition? Lin pivots to the broader portfolio: 2024 was the year of coding, 2025 is the year of co-work, and co-work customers span legal, finance, healthcare, consumer-facing recommendation systems, and beyond. He then drops a striking operational disclosure: Fireworks processes more than 40 trillion tokens per day, the majority from customised models. When Harry asks where that number goes in a year, Lin offers a range of 20x to 100x — and argues that at those growth rates, the idea that AI CapEx is in a bubble is simply wrong. The world is bottlenecked by physical supply chains, not by lack of demand.

This is the most quotable chapter in the episode. Lin explains that current token prices are artificially elevated by supply chain constraints, not by any fundamental cost floor. As competition increases and infrastructure scales, prices will fall sharply. He quantifies three levers: solving tasks requires fewer tokens as models become more precise; Fireworks' platform optimises inference unit economics for each customised workload; and underlying GPU and memory infrastructure will improve structurally over two to three years. The combined effect: a 10x cost reduction over three years that unlocks a 100x usage explosion. He then dives into zero KLD — Fireworks' commitment to bit-exact numerical equivalence from training to inference — as the quality guarantee that justifies its premium over commoditised inference providers.

Harry observes that China's lack of regulatory friction means data centres go up in weeks versus years in the US, and asks whether that translates into a structural AI advantage. Lin acknowledges the construction velocity differential but argues the US has comparable specialisation — the bottleneck is global supply chain constraints (electricians, transistors, materials) rather than policy alone. The conversation then broadens to sovereign models: the brief TikTok ban illustrated the existential vulnerability of depending on foreign-controlled infrastructure. Lin's analogy is stark — if your country's AI is like electricity, another country should never control the switch. The principle applies equally to individual enterprises: owning your intelligence is not optional.

The quickfire round delivers some of the episode's most memorable moments. Lin admits he was wrong to fear growing too fast — aggressive AI tool adoption and a refined hiring filter for extreme ownership changed his calculus. His Jensen Huang insight crystallises into a philosophy: replying to emails in one minute isn't ego, it's the only way to maintain the information freshness needed for precise, fast leadership judgement. His biggest regret is underinvesting in marketing early, framing it not as fluff but as the discipline of educating customers on the right direction. He closes with his sharpest three-year prediction: owning your own intelligence stack will become mandatory for every company, just as owning your own software stack is today. Sponsor reads for JPMorgan, Navan, and Base44 close the episode.

Chapter 1 · 00:07

The episode opens with a punchy teaser from Lin Qiao before Harry Stebbings delivers a rare investor endorsement: a $10 million check written after just a 15-minute meeting. Harry breaks down the five reasons — a world-class team, a fast-growing inference market, triple-digit ARR growth to $1 billion in four years, the ability to hire stars like ex-Salesforce President George Hu, and the sheer upside potential of a company he thinks could reach $500 billion. The framing sets the episode's bullish tone before three sponsor integrations (JPMorgan, Navan, Base44) round out the opening block.

Fireworks AI raised $1.5 billion at a $17 billion valuation, a remarkable outcome for a 200-person company.

Fireworks AI hit $1 billion in ARR with just 200 employees by betting on specialised inference when everyone else was chasing training. The company processes over 40 trillion tokens a day, mostly from customised rather than off-the-shelf models.

Fireworks AI reached $1 billion in annual recurring revenue with only 200 employees, scaling in roughly 4 years.

Navan claims the industry average for booking a business trip is 45 minutes, versus 7 minutes on their platform.

Lin Qiao founded Fireworks AI at 48 years old, after 7 years at Meta and earlier stints at LinkedIn and in academia.

Most of the world's valuable data sits locked inside enterprise applications, never touching a general model's training set. Fireworks AI was built on the conviction that activating this private data through specialised models is the real frontier of AI.

The AGI believers assume one model will solve everything. Lin Qiao thinks that's both technically wrong and philosophically depressing. The future is millions of specialised models — one per application, per use case, per company.

After Jensen Huang told Lin Qiao that every company must be special to justify its existence, Lin realised the implication was profound: all of a company's product design, data, and user relationships encode irreplaceable private intelligence that no external model can learn.

Chapter 2 · 13:00

Harry asks the valuation question directly: if 90% of enterprise workflows can now be handled by open-weight models at a fraction of the cost, are Anthropic and OpenAI dramatically overvalued? Lin agrees the market is beginning to realise this, then introduces a more uncomfortable concept — 'scaling to bankruptcy.' In the SaaS era, finding product-market fit was the hard problem; scaling was cheap. In the AI era, these are decoupled. Companies with genuine customer demand and willingness to pay can destroy themselves by growing because the AI infrastructure COGS spiral faster than revenue. This creates a structural pull toward open-weight models that companies can control, customise, and optimise.

When Fireworks was founded, open models were in their infancy. Betting on them was a huge gamble. The PyTorch roots gave the team conviction in open ecosystems, and the payoff came as open models crossed quality thresholds that now rival closed models for the vast majority of enterprise use cases.

Lin Qiao suggests that as open models handle 90% of enterprise use cases at a fraction of the cost, the market is beginning to re-examine whether frontier model companies are priced for a world that will actually materialise. The power-line metaphor applies: important, yes — irreplaceable, no.

In the SaaS era, finding product-market fit was the hard part — scaling was cheap. In the AI era, companies with genuine demand can still destroy themselves by growing, because AI infrastructure costs don't scale gracefully. Lin Qiao calls it 'scaling to bankruptcy.'

Chapter 3 · 19:00

Harry — declaring a conflict of interest as a Lagoora investor — uses the Harvey-Lagoora dynamic as a test case. A year ago, not building your own model seemed fine because frontier models were improving so fast. Now the picture is reversed. Lin reframes the question: in legal, where error tolerance is near-zero and workflows are deeply specialised, the competitive advantage isn't whether you own a model but whether you've encoded your proprietary knowledge into the orchestration layer and the fine-tuned models that power it. He points to Cursor as evidence — they were the first coding AI company to tune their own models, and almost all coding companies have followed.

The top six open-weight models globally are currently Chinese-built. Lin Qiao argues that once a model is open, enterprises can wrap their own guardrails around it, but the deeper point is that every model — Chinese or American — encodes its creator's judgment and taste, which always needs tuning.

Lin Qiao believes the future will feature millions of specialised AI models, one per application or use case, rather than a single dominant AGI.

Chapter 5 · 28:00

One of the most technically rich chapters in the episode. Lin explains that early AI tech adopters are hackers who want control, and Cursor — flush with frontier lab researchers — is the prime example. The problem they solved together is profound: hyperscalers run RL training on 100,000 interconnected chips. Cursor and Fireworks had no such cluster. Their solution was to decouple the trainer (which updates model weights) from the RL rollout (which deploys the model into a synthetic environment to collect rewards), distributing the system across five or six global data centre regions and syncing fresh model weights efficiently enough that the reward signal stays numerically sound. This distributed design enabled Cursor's recent model launches without a hyperscaler-level budget.

Fireworks CTO Dima embedded at Cursor for months to build a distributed reinforcement learning infrastructure that decouples the trainer from RL rollout across six global data centre regions. This let a capital-constrained startup run training jobs that previously required 100,000 interconnected chips at a hyperscaler.

Lin Qiao sees 2024 as the year of coding AI and 2025 as the year of co-work. Co-work is dramatically more diverse than coding — spanning legal, finance, healthcare, customer support, and consumer-facing AI — and Fireworks is already landing customers across all of those verticals.

Fireworks processes more than 40 trillion tokens a day today. Lin Qiao projects that number could be 20x to 100x higher by end of next year. At those volumes, worries about a CapEx bubble look completely backwards.

Fireworks processes more than 40 trillion tokens per day, the majority from customised rather than off-the-shelf models.

Chapter 6 · 37:00

Harry opens with the customer concentration question — can Fireworks survive a Cursor churn given a potential SpaceX acquisition? Lin pivots to the broader portfolio: 2024 was the year of coding, 2025 is the year of co-work, and co-work customers span legal, finance, healthcare, consumer-facing recommendation systems, and beyond. He then drops a striking operational disclosure: Fireworks processes more than 40 trillion tokens per day, the majority from customised models. When Harry asks where that number goes in a year, Lin offers a range of 20x to 100x — and argues that at those growth rates, the idea that AI CapEx is in a bubble is simply wrong. The world is bottlenecked by physical supply chains, not by lack of demand.

Lin Qiao projects Fireworks' token throughput could grow 20x to 100x by end of next year as AI adoption accelerates.

Cursor, an early Fireworks customer when it was a single-digit-million company, grew approximately 1,000x over two years.

Marc Benioff reportedly spends about 3.8% of Salesforce developer salaries on Anthropic's Claude Code product.

Chapter 7 · 43:00

This is the most quotable chapter in the episode. Lin explains that current token prices are artificially elevated by supply chain constraints, not by any fundamental cost floor. As competition increases and infrastructure scales, prices will fall sharply. He quantifies three levers: solving tasks requires fewer tokens as models become more precise; Fireworks' platform optimises inference unit economics for each customised workload; and underlying GPU and memory infrastructure will improve structurally over two to three years. The combined effect: a 10x cost reduction over three years that unlocks a 100x usage explosion. He then dives into zero KLD — Fireworks' commitment to bit-exact numerical equivalence from training to inference — as the quality guarantee that justifies its premium over commoditised inference providers.

Token costs are artificially high because of supply chain constraints. Once competition and infrastructure catch up, Lin Qiao expects a 10x cost reduction in three years. Cheaper tokens will unlock use cases that are currently economically impossible, driving a 100x surge in usage.

Lin Qiao predicts token costs will fall 10x over the next three years due to supply chain improvements and infrastructure efficiency.

Lin Qiao argued that a 10x cost reduction in AI tokens will trigger a 100x explosion in usage as AI becomes a utility.

Fireworks achieves zero KL-divergence between training and inference — meaning model weights transfer with bit-exact numerical equivalence. This matters enormously at scale: even tiny quality drops at the train-inference boundary waste customers' entire training investment.

Fireworks AI achieves zero KL-divergence between training and inference, meaning model weights transfer with full numerical equivalence and no quality loss.

Chapter 8 · 49:00

Harry observes that China's lack of regulatory friction means data centres go up in weeks versus years in the US, and asks whether that translates into a structural AI advantage. Lin acknowledges the construction velocity differential but argues the US has comparable specialisation — the bottleneck is global supply chain constraints (electricians, transistors, materials) rather than policy alone. The conversation then broadens to sovereign models: the brief TikTok ban illustrated the existential vulnerability of depending on foreign-controlled infrastructure. Lin's analogy is stark — if your country's AI is like electricity, another country should never control the switch. The principle applies equally to individual enterprises: owning your intelligence is not optional.

Traditional GPU hardware depreciated over six years. Now multiple new SKUs launch annually, and the newest models always prefer the newest chips. This collapse in depreciation cycles fundamentally changes whether companies should own or rent data centre capacity.

Traditional hardware depreciation cycles of 6 years are being upended as multiple GPU SKUs launch within a single year, making older hardware obsolete faster.

Watching TikTok get briefly banned in the US crystallised a broader truth: if your critical intelligence infrastructure sits on another country's AI, it can be cut off overnight. Lin Qiao argues every nation — and every company — needs sovereign AI the same way they need sovereign electricity grids.

Meta has been building its own custom AI chips (MTIA) since at least 2018, well before the generative AI era.

Chapter 9 · 1:00:15 The quickfire round delivers some of the episode's most memorable moments. Lin admits he was wrong to fear growing too fast — aggressive AI tool adoption and a refined hiring filter for extreme ownership changed his calculus. His Jensen Huang insight crystallises into a philosophy: replying to emails in one minute isn't ego, it's the only way to maintain the information freshness needed for precise, fast leadership judgement. His biggest regret is underinvesting in marketing early, framing it not as fluff but as the discipline of educating customers on the right direction. He closes with his sharpest three-year prediction: owning your own intelligence stack will become mandatory for every company, just as owning your own software stack is today. Sponsor reads for JPMorgan, Navan, and Base44 close the episode.

Everyone talks about HBM or energy as the AI bottleneck. Lin Qiao's answer is different: the industry still has no great system designed for 10-trillion-parameter models. Closing that gap requires co-design from model to serving platform to chip system — a level of integration that doesn't yet exist.

Fireworks AI expects to at least double from $800M–$1B ARR by the end of 2026, targeting roughly $2B ARR.

Lin Qiao met George Hu, the former President of Salesforce, a year before hiring him — and turned him down because Fireworks had only 50 people. The relationship evolved through board-level advice before Hu formally joined as Fireworks crossed into hypergrowth. The lesson: build the relationship early, even if the timing isn't right yet.

Fireworks doesn't hire for competence first. They hire for extreme ownership — people who claim end-to-end problems without being asked, see them through to delivery, and treat every outcome as their own responsibility. That mindset, Lin Qiao says, compounds faster than any skill.

Jensen Huang replies to emails within one minute. Lin Qiao spent years marvelling at the sheer volume before understanding the wisdom: in a fast-moving company, information loss between layers is guaranteed. The only antidote is a leader who stays close to the ground and makes precise, fast judgements.

Just as every company builds its own software stack because its problems are unique, every company will eventually own its own AI intelligence stack. Renting general intelligence from a handful of providers will be seen as a competitive liability — not a convenience.

No indexed bits in this chapter.

Sign in to keep viewing

Create a free account to keep exploring this episode's insights, snapshots and quotes.

We scan show notes for social handles, websites and apps. Nothing matched on this episode.

We use essential and analytics cookies to run Vuci. To understand how the site is used:

Privacy Policy. Install Vuci on your phone

Add it to your home screen for a faster, app-like experience.

Install Vuci on your phone

Tap the Share button, then “Add to Home Screen”.

A new version is available

Reload to get the latest Vuci.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @fireworks ai 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/20vc-are-openai-and-…] indexed:0 read:18min 2026-07-20 ·