Solving Hack The Box Challenges with GPT-5.6
OpenAI's GPT-5.6 Terra and Sol models initially rejected security-related prompts from a researcher, but joining the Trusted Access for Cyber program resolved the issue, allowing testing of the full G…
OpenAI's GPT-5.6 Terra and Sol models initially rejected security-related prompts from a researcher, but joining the Trusted Access for Cyber program resolved the issue, allowing testing of the full G…
XAI's Grok 4.6 achieved an 80% benchmark score on the HTB-Challenger Benchmark, solving 14 of 16 Hack The Box challenges with zero false positives, surpassing all previously tested models in both spee…
DeepSeek V4 Pro solved 8 of 16 Hack The Box challenges in the HTB-Challenger Benchmark, scoring 36.2% with no false positives, but its median cost per challenge was $0.35, comparable to GPT-5.6 Luna d…
DeepSeek V4 Flash 0731 solved only 2 of 16 Hack The Box challenges, scoring 8.2% on the HTB-Challenger Benchmark, and reported incorrect flags in 9 challenges, a false-positive rate unmatched by any o…
Qwen3.8 Max, released by Alibaba Cloud, solved 10 of 16 Hack The Box challenges with a benchmark score of 54.8%, but its median cost per challenge was $2.00, making it the most expensive model tested …
OpenAI's GPT-5.6 Luna Pro solved 11 of 16 Hack The Box challenges, scoring 55.0% on the HTB-Challenger Benchmark, compared with GPT-5.6 Luna which solved only one Medium and no Hard challenges. The mo…
Hack The Box's new HTB-Challenger benchmark evaluates LLM models on offensive security challenges, addressing the high cost and comparability issues of existing methods. The benchmark, created by an a…