{"slug": "glm-5-3-flash-ox-alpha-the-difference-they-didn-t-tell-you", "title": "GLM-5.3-Flash / Ox-Alpha The Difference They Didn't Tell You!", "summary": "Z.ai's stealth preview model, branded 'ox-alpha' on OpenRouter.AI and OpenCode.AI, has launched as GLM-5.3-Flash, an efficient open-weights model priced at $0.07 per million input tokens and $0.25 per million output tokens. A developer testing the model found it generates original jokes rather than memorized ones, calling this a 'killer feature' that sets it apart from other models. The model has been praised for handling long-running agentic tasks at low cost, with one user reporting a 12-hour coding session costing just $0.47.", "body_md": "If you looked the other way, you may have missed the stealth preview model that was free on OpenRouter.AI and OpenCode.AI, which was branded \"ox-alpha\", that was making a bit of a splash. Everyone had figured out that it was a new Z.ai multi-modal, but no one could figure out how they were giving away 5,835,184,092,873 prompt tokens and 94,402,126,602 completion tokens in a single day!\n\nThe answer when it launched as GLM-5.3-Flash is that its an very efficient Flash open-weights model that you can get hosted by a US hardware by San Francisco firm [OpenCode.AI](https://opencode.ai/data/unknown/ox-alpha) for an input cost of $0.07 / 1M and output cost of $0.25 / 1M.\n\nYet they have hidden one dark truth that I can now reveal. The ultimate proof that this model is different. I have hooked it into my personal fork of the rust based apache2 codex, which I prefer, and I asked it to tell me a joke, and it bombed:\n\nIn the era of all the models memorising a snake game and \"tell me a joke\" it made one up on the spot and bombed 💖💚💛❤️\n\nI have promoted all the models \"tell me a joke\" on all the harnesses, and you get the same set of responses. They have memorised that. They do not make up stuff. They stick with some very common jokes. I have given up asking for \"tell me another\" as they only have a limited set.\n\nYet there we are GLM-5.3-Flash, as a flash model is not trying to go from memory, it has a crack at it. And that makes this model special. That is NOT a bug, that is a killer feature!\n\nWe are talking about a model that Theo.gg puts up there with Claude Opus and OpenAI Sol as being that good for getting on with huge, long-running agentic tasks in a 45 minute break down [Ox Alpha is INSANE](https://youtu.be/Xdxp3lbQKyQ)\n\nIt ran all night doing great work, and it cost me less than a coffee at my local Gail's Bakery. I had Kimi K3 cross-check the work, and it was 10x the cost to just review the git diff!\n\nLast night I had OpenCode.AI coding agent up and running working across all my current projects and repos going flat out. On one massive repo dust off it is still running after twelve hours, having spent $0.02 max on each bug fix todo, and a grant total for $0.47 total.\n\nLove GenAI or hate it, this new mode is no joke. Boom boom!", "url": "https://wpnews.pro/news/glm-5-3-flash-ox-alpha-the-difference-they-didn-t-tell-you", "canonical_source": "https://dev.to/simbo1905/glm-53-flash-ox-alpha-the-difference-they-didnt-tell-you-1clo", "published_at": "2026-08-28 10:49:48+00:00", "updated_at": "2026-08-28 11:19:31.268372+00:00", "lang": "en", "topics": ["large-language-models", "generative-ai", "ai-products", "ai-tools"], "entities": ["Z.ai", "GLM-5.3-Flash", "OpenRouter.AI", "OpenCode.AI", "Theo.gg", "Kimi K3", "Claude Opus", "OpenAI Sol"], "alternates": {"html": "https://wpnews.pro/news/glm-5-3-flash-ox-alpha-the-difference-they-didn-t-tell-you", "markdown": "https://wpnews.pro/news/glm-5-3-flash-ox-alpha-the-difference-they-didn-t-tell-you.md", "text": "https://wpnews.pro/news/glm-5-3-flash-ox-alpha-the-difference-they-didn-t-tell-you.txt", "jsonld": "https://wpnews.pro/news/glm-5-3-flash-ox-alpha-the-difference-they-didn-t-tell-you.jsonld"}}