I am curious what actually happens between two generations of AI models.
For example, how do you go from Sonnet to Opus? Is Opus trained from scratch, built on Sonnet, or mostly the same model with more compute and training? And how do models like Astra suddenly make a big jump in some capabilities? What is stopping Anthropic, Mistral, or others from doing the same thing? Is the main difference just more compute and money, or are there training methods, data, architecture, and research breakthroughs that competitors may not know about?
I can't think of a better place to ask this. I am guessing there are people here who actually work on these models and know what goes on behind the scenes.
Comments URL: [https://news.ycombinator.com/item?id=49662978](https://news.ycombinator.com/item?id=49662978)
Points: 1