Post-training to build better Roblox UIs
Lemonade post-trained Hy3, a 295B model, to generate Roblox UI in its agentic Roblox Studio harness, targeting parity with Grok 4.5 at lower cost for its hundreds of thousands of users. The team built…
Lemonade post-trained Hy3, a 295B model, to generate Roblox UI in its agentic Roblox Studio harness, targeting parity with Grok 4.5 at lower cost for its hundreds of thousands of users. The team built…
A developer used GPT 5.6 Luna via Simon Willison's `llm` CLI tool to label git commits as "maintenance" or "new development", reporting that the model matched their own manual labels across the entire…
AI Gateway's model leaderboard for the period from June 21, 2026 to September 18, 2026 shows Jev leading in reach at 15.6% of teams and in preference at 13.3% of teams using it as their primary model …
A new benchmark shows that the Qwen3.8 27B model outperforms GPT 5.6 Luna on the DABstep SQL benchmark while costing under 50 cents in electricity, over 17x less, according to a blog post by Alex Mona…
1endpoint, a new API gateway, offers access to multiple AI models through a single endpoint with usage-based pricing, starting at $0.0420 per 1M input tokens for GLM 5.2. The service includes prompt c…
Between October 2018 and July 2026, AI models progressed from simple systems like BERT to massive agents that solve complex math and write software, with the ability to resolve real coding issues impr…
OpenAI's price cut for GPT 5.6 Luna by 80% has made AI-powered analytics answers cost less than half a penny each, with Luna achieving 99.8% accuracy on an agentic SQL benchmark at a 5x lower price th…
Pathway, a startup founded by Zuzanna Stamirowska, unveiled a 150-million parameter reasoning model, BDH-CQ, claiming it achieves comparable performance to leading frontier models at a fraction of the…
GitHub reported degraded availability for GPT 5.6 Luna, an AI product, and is offering email and SMS notifications for updates on the incident.…
Grok 4.5's cache price per 1 million tokens has dropped from $0.50 to $0.30, making it the cheapest model for agentic tasks according to internal benchmarks, even cheaper than GPT 5.6 Luna. The price …
Google's new Gemini 3.6 Flash model outperforms its predecessors Gemini 3.5 Flash and Gemini 3.1 Pro on every published benchmark, scoring 58.7% on SWE-Bench Pro versus 55.1% and 54.2%, and 49% on Dee…
OpenAI released GPT 5.6 in three models—Sol, Terra, and Luna—now available on Vercel's AI Gateway. The models offer improved agentic capabilities in coding, biology, and cybersecurity with greater tok…
SpaceXAI launched Grok 4.5, its smartest model yet, as a coding and agentic-work tool priced at $2 per million input tokens. The model outperforms Anthropic's Opus 4.8 on key benchmarks but trails com…