Evaluating Hy3 on Hack The Box Challenges
Tencent's Hy3 scored 34.8% on the HTB-Challenger Benchmark, the third-lowest among all tested models, and got stuck on 9 of 16 Hack The Box challenges, generating a median of 101,509 output tokens per…
Tencent's Hy3 scored 34.8% on the HTB-Challenger Benchmark, the third-lowest among all tested models, and got stuck on 9 of 16 Hack The Box challenges, generating a median of 101,509 output tokens per…
OpenAI's GPT-5.6 Terra and Sol models initially rejected security-related prompts from a researcher, but joining the Trusted Access for Cyber program resolved the issue, allowing testing of the full G…
SpaceXAI's Grok 4.6 topped the Newsroom Reliability v0.2 benchmark with a mean score of 0.78, outperforming OpenAI's GPT-5.6 Luna Pro (0.77) and GPT-5.6 Sol Pro (0.76), Meta's Muse Spark 1.2 (0.73), A…
OpenAI's GPT-5.6 Luna Pro solved 11 of 16 Hack The Box challenges, scoring 55.0% on the HTB-Challenger Benchmark, compared with GPT-5.6 Luna which solved only one Medium and no Hard challenges. The mo…
OpenAI's GPT-5.6 lineup introduces three tiers—Sol, Terra, and Luna—each with a Pro variant that allocates extra compute for deeper reasoning. Sol targets complex agentic tasks, Terra balances capabil…