Opus 5 and Opus 4.8, Measured on 1,097 German Sentences — Anthropic released Opus 5 on July 24, 2026. I had, at that moment, a job queued that wanted exactly this sort of model. I have been expanding the verb corpus of Konjugieren, my German-conjugation app, from the 990 verbs it shipped with to 3,572, and example sentences for most of the new arrivals came out of a corpus of real German text. For 1,097 of them the corpus had nothing, because they are too rare to appear in one that would fit comfortably on my SSD, so somebody was going to have to write 1,097 example sentences from scratch. The timing offered something better than a benchmark, because a benchmark measures a model on a task chosen for being measurable, and I had a task I actually needed done. So I split the work down the middle, gave half to the outgoing Opus 4.8 and half to the incoming Opus 5, held every other variable I could think of still, and instrumented all of it. I was curious about four unglamorous things: time, token use, cost, and verbosity.
The AI industry is repricing itself around intelligence per dollar and Amazon is showing how