{"slug": "ai-agents-weekly-glm-5-3-flash-hy4-preview-qwen3-8-flash-claude-s-built-in-bench", "title": "🤖 AI Agents Weekly: GLM-5.3-Flash, Hy4 Preview, Qwen3.8-Flash, Claude's Built-In Browser, Terminal-Bench-Science, Jalapeño, Skild S1, and More", "summary": "Z.ai released GLM-5.3-Flash, a natively multimodal 320B-A18B model with a 1M-token context window, under the MIT license, priced at $0.15 per 1M input tokens and $0.50 per 1M output tokens. The model scores 84.3 on Terminal Bench 2.1, 63.4 on DeepSWE v1.1, 48.8 on AutomationBench v1.0.6, 55.3 on HLE with tools, and 1773 on GDPVal-AA v2, beating GLM-5.2 on every benchmark. Z.ai served all anonymous traffic on Chinese AI chips, reporting 3x better end-to-end serving performance than its earlier baseline.", "body_md": "In today’s issue:\n\nGLM-5.3-Flash ships under MIT\n\nTencent opens Hy4 preview weights\n\nQwen previews the Qwen4 architecture\n\nClaude gets its own browser\n\nTerminal-Bench-Science scores agents on science\n\nOpenAI reports first Jalapeño results\n\nSkild S1 learns from one video\n\nHeadlong keeps agents always thinking\n\nX launches Chat Agents\n\nMCP publishes its next roadmap\n\nAI4AI-Bench tests recursive self-improvement\n\nAgents close 81.7% of the speedrun gap\n\nRepo-wide migrations survive 5.4% of runs\n\nAnd all the top AI dev news, papers, and tools.\n\n## Top Stories\n\n### GLM-5.3-Flash Ships Under MIT\n\nZ.ai released GLM-5.3-Flash, a natively multimodal 320B-A18B model with a 1M-token context window, published under the MIT license. It was previously previewed as Ox Alpha.\n\n**Agentic benchmarks:** 84.3 on Terminal Bench 2.1, 63.4 on DeepSWE v1.1, 48.8 on AutomationBench v1.0.6, 55.3 on HLE with tools, and 1773 on GDPVal-AA v2, ahead of GLM-5.2 on every one.**Coding performance:** On Z.ai Code Bench v1.0, run through Claude Code, GLM-5.3-Flash beats GLM-5.2 at every effort level and at max effort comes within half a point of Claude Opus 4.8 at 29.0 against 29.5.**Priced to run in a loop:**$0.15 per 1M input tokens, $0.50 per 1M output, and $0.03 for cached input, which makes long agent trajectories cheap to iterate on.**Hybrid attention carries the efficiency:** Linear attention captures local dependencies while sparse attention retrieves global context through a lightweight indexer, cutting attention compute 3.0x and KV cache 4.4x against GLM-5.3. Against GLM-4.5 it nearly halves both activated parameters (18B against 32B) and layers (45 against 92).**Served on Chinese silicon:** Z.ai ran the model anonymously as ox-alpha on OpenCode and OpenRouter before release and served all of that traffic on Chinese AI chips, reporting 3x better end-to-end serving performance than its own earlier baseline on the same hardware.", "url": "https://wpnews.pro/news/ai-agents-weekly-glm-5-3-flash-hy4-preview-qwen3-8-flash-claude-s-built-in-bench", "canonical_source": "https://nlp.elvissaravia.com/p/ai-agents-weekly-glm-53-flash-hy4", "published_at": "2026-08-29 15:33:13+00:00", "updated_at": "2026-08-29 15:49:40.211215+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-products", "ai-infrastructure"], "entities": ["Z.ai", "GLM-5.3-Flash", "GLM-5.2", "Claude Opus 4.8", "OpenCode", "OpenRouter"], "alternates": {"html": "https://wpnews.pro/news/ai-agents-weekly-glm-5-3-flash-hy4-preview-qwen3-8-flash-claude-s-built-in-bench", "markdown": "https://wpnews.pro/news/ai-agents-weekly-glm-5-3-flash-hy4-preview-qwen3-8-flash-claude-s-built-in-bench.md", "text": "https://wpnews.pro/news/ai-agents-weekly-glm-5-3-flash-hy4-preview-qwen3-8-flash-claude-s-built-in-bench.txt", "jsonld": "https://wpnews.pro/news/ai-agents-weekly-glm-5-3-flash-hy4-preview-qwen3-8-flash-claude-s-built-in-bench.jsonld"}}