{"slug": "gradually-then-suddenly", "title": "Gradually, then Suddenly", "summary": "Meta unveiled Muse Glimmer, Alibaba released Qwen3.8-27B, and startup Ornith dropped Ornith-1.5-35B, a mixture-of-experts model that runs efficiently on consumer hardware, enabling agents to install and run locally. Mark Pesce, writing in The Watershed, reported that Ornith-1.5-35B runs about fifteen times faster than Qwen on his six-year-old PC, and he now runs it on all three household PCs, predicting an explosion of AI agents on existing hardware.", "body_md": "# Gradually, then Suddenly\n\nIt's been quite a week for the Home Watershed.\n\nLast week Meta unveiled [Muse Glimmer](https://research.meta.ai/blog/introducing-muse-glimmer-open-agentic-model?ref=thewatershed.markpesce.com), which gave us a taste of what might be coming.\n\nOver the weekend Alibaba dropped [Qwen3.8-27B](https://thewatershed.markpesce.com/52-70-byoai/), which scores 52 on the Artificial Analysis Intelligence Index. The only complaint anyone has is that, as a 'dense' model, it really demands a dedicated, high-end GPU - or you accept that it inferences slowly.\n\nTo get around that you need a different kind of model with a different architecture, known as a 'mixture of experts', or MoE. MoE models inference far more efficiently, and don't demand the same level of hardware.\n\nThat's just what we got yesterday.\n\nOrnith - a startup obscure enough that few people working in AI had heard of it - dropped [Ornith-1.5-35B](https://ornith.ai/ornith_1_5.html?ref=thewatershed.markpesce.com), an MoE model. According to Ornith's own benchmarks (taken with an appropriate grain of salt), possibly a quite good one.\n\nI put it to the test immediately - by firing up an agent and asking *it* to install Ornith on my six-year-old PC. Half an hour later, everything was up and running, and I hadn't touched a config file. That's how things are on this side of the home watershed: a new model came out yesterday, and an agent did the installing. Once it was running, I could see it inferencing about fifteen times faster than Qwen had managed on the same machine. That's a massive difference.\n\nBut is Ornith 'smart enough'?\n\nThat's an easy question to answer in theory and a hard one to answer in practice. The best rule of thumb: you learn whether a model can do the job *by giving the model the job*, then watching how it performs.\n\nWhich I did, across a range of tasks, including one that demands deeply considered critical analysis. That last is my own personal benchmarking tool - a unique problem, and one that every model answers uniquely. Ornith wasn't quite as stellar as Qwen. But it was \"good enough\" to leave me seriously impressed.\n\nImpressed enough that every PC in the house - there are three, including one that's now a decade old - is running Ornith. **Because they can.** Each has a 'good enough' agent running with just enough speed and capacity to be useful across countless applications, available around the clock for whatever projects I have in mind.\n\nIn the course of one week I've gone from no \"good enough\" models running at home to four - likely five, once I get around to setting up the MacBook Air.\n\nThe installed base isn't the good computer you own. It's *all of them* - and this week, one by one, mine came to life.\n\nThat gives a sense of what's coming: now that models are both \"good enough\" and \"fast enough\" to run on a broad range of *already installed* consumer hardware, everything is in place for an explosion in the number of agents.\n\nIt's just a software distribution problem now - and my six-year-old PC already showed you how that problem ends. When an agent can install the next agent, distribution stops being a download and becomes a sentence: *set this up for me*.\n\nAll three of these models are open weight and cost nothing. Nothing stands between some clever person's killer app and the hundreds of millions of machines that could be running it by the weekend after it appears.\n\nWe saw a glimpse early this year with OpenClaw. But OpenClaw needed a \"big\" model, running in a data centre, to be truly useful. Now that model - or near enough - runs on the home computer.\n\nHemingway's bankrupt explained how it happened: two ways. Gradually, and then suddenly. The gradual part took seventy years, and finished at Dartmouth's anniversary last week. This was the first week of suddenly.\n\nWe've seen the lightning. Now we wait for the thunderclap.", "url": "https://wpnews.pro/news/gradually-then-suddenly", "canonical_source": "https://thewatershed.markpesce.com/gradually-then-suddenly/", "published_at": "2026-08-20 22:00:37+00:00", "updated_at": "2026-08-20 22:12:42.308272+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-agents", "ai-products"], "entities": ["Meta", "Alibaba", "Ornith", "Qwen3.8-27B", "Ornith-1.5-35B", "Muse Glimmer", "Mark Pesce", "The Watershed"], "alternates": {"html": "https://wpnews.pro/news/gradually-then-suddenly", "markdown": "https://wpnews.pro/news/gradually-then-suddenly.md", "text": "https://wpnews.pro/news/gradually-then-suddenly.txt", "jsonld": "https://wpnews.pro/news/gradually-then-suddenly.jsonld"}}