I, like many people, have witnessed and felt the whiplash between the seemingly-orgastic highs of the Astra launch and the experience of using the model.
Whether intentional or otherwise, a lot of airtime was given in the brief pre-release hypecycle to Astra's ability to generate things that looked cool. Lots of games, renders, and canvas examples cluttered my various timelines.
This makes sense: 3D stuff is cool; most people don't really have a sense of the toolchain involved in things like this, and how much of the effort is on the part of the LLM versus the structures it builds upon is somewhat obfuscated; platforms like xtheeverythingapp.com tend to reward videos and animations.
I try (unsuccessfully) to remain sober about all of this. But the release of the model, coupled with things like the deeper information around the Hugging Face incident and Collusion.wiki, did ignite a bit of a spark in the back of my head — thinking, as always, however briefly, about the chance that things are really different this time. Did we solve software?
Then I started using the model.
And much like my quasi-review of Fable 5.1, I found that the model is good! It's hard for me to really compare and contrast it against Fable 5.1 because I didn't do formal comparisons or anything like that; I just tried throwing a couple of problems I was working on against it, and it handled them all well and the things it couldn't handle were generally due to lack of context and agency, not lack of horsepower.
It's very smart and capable, but limited by the tools and harnesses and contexts it has access to (words I deploy in the broader sense, not the industry-specific ones.). All of this mirrors almost exactly my sense of wonder followed by disillusion when I started really kicking around Fable for the first time. Though I do maintain that Fable pre-DOD was a slightly different beast, and felt more akin to a real change in how I programmed.
Astra is very good at a lot of things; for me, though, it is increasingly extremely competent at tasks that have already been solved by faster, cheaper models. And at least personally, as someone who runs a non-trivial business where the risk of false positives vastly outweighs the attractiveness of sheer velocity, I struggle to find the right balance here between exploring the boundaries of what is possible and remembering that right now my job is no longer that of a technologist but of an industrial designer.
I run a small business where we have absolute license to use models at will, so long as we do so judiciously and securely. I talk about them for no economic or personal reason Disclaimer: I am, as of this writing, an investor in Anthropic through an acquisition they made.; I want to understand (and thereby master) them out of a conviction that they'll improve my business.
I bring that up because I found myself vigorously nodding along to this essay from Armin Ronacher, who writes (amongst other things, which I agree with):
I'm more and more convinced that all of AI engineering is Neijuan (内卷, meaning curl inwards). In China it describes a system that demands ever more effort and competition without improving output. The way in which it sometimes shows up in the West is the 996 nonsense. The English term for Neijuan is "Involution" from the book Agricultural Involution. Agricultural involution describes the intensification of farming that raises productivity per square meter while leaving productivity per head unchanged.
One way of thinking about LLMs (or any technology) is dichotomous:
- There is the raw power a given model can generate,
- and then there's the design and manufacturing of interfaces that can use that model in useful ways.
The biggest step function in software engineering over the past few years was not a specific model but Claude Code. Astra and Fable have improved my life less than full support for subagents and worktrees did.
I get that it's unfair of me to pretend those two things are discrete! I don't have enough intelligence to understand how raw power acts as a proxy for various other implications outside my industry. But I do — at least for now — have enough experience to understand that right now it's a poor proxy for operational leverage in mine.