cd/entity/HTB-Challenger Benchmark· home› entities› HTB-Challenger Benchmark
grep -l @htb-challenger benchmark /news/*.json | wc -l → 17

HTB-Challenger Benchmark

mentions 17 type Person feed RSS

// recent coverage 17 mentions

11:18
2026-09-11
theaq.blog
large-language-models

Evaluating DeepSeek V4.1 Flash on Hack The Box Challenges

DeepSeek V4.1 Flash solved 15 of 16 Hack The Box challenges on the HTB-Challenger Benchmark, including three of the four Hard challenges, for a benchmark score of 80.9% with zero false positives, acco…

09:15
2026-09-06
theaq.blog
artificial-intelligence

Evaluating Muse Spark 1.3 on Hack The Box Challenges

Meta's Muse Spark 1.3 scored 72.93% on the HTB-Challenger Benchmark, up from 50.41% for Muse Spark 1.2, according to a blog evaluation of the model on Hack The Box security challenges. The new version…

07:57
2026-08-28
theaq.blog
artificial-intelligence

Evaluating GLM 5.3 Flash on Hack The Box Challenges

Z.ai's GLM 5.3 Flash scored 81.0% on the HTB-Challenger Benchmark, solving 15 of 16 Hack The Box challenges, second only to OpenAI's GPT-5.6 Sol at 87.2%. Despite requiring more steps (median 24.5 vs.…

10:40
2026-08-24
theaq.blog
artificial-intelligence

Case Study: Combining GPT-5.6 Luna and Sol for Cost-Efficient AI

A case study combining OpenAI's GPT-5.6 Luna and GPT-5.6 Sol models achieved near-Sol performance at a fraction of the cost on the HTB-Challenger Benchmark, which tests LLMs on Hack The Box security c…

07:07
2026-08-21
theaq.blog
artificial-intelligence

Evaluating GLM 5.3 on Hack The Box Challenges

Z.ai's GLM 5.3, a retrained version of GLM 5.2, scored 47.2% on the HTB-Challenger Benchmark, solving 10 of 16 Hack The Box challenges with a median cost of $0.49 per challenge, placing it near Muse S…

10:15
2026-08-20
theaq.blog
artificial-intelligence

What Does the HTB-Challenger Benchmark Actually Measure?

The HTB-Challenger Benchmark evaluates large language models' ability to find and exploit security vulnerabilities using selected Hack The Box challenges of varying difficulty, according to the benchm…

17:15
2026-08-18
theaq.blog
artificial-intelligence

Evaluating DeepSeek V4 Pro 0813 on Hack The Box Challenges

DeepSeek released DeepSeek V4 Pro 0813, an updated version of its V4 Pro model, featuring the same core architecture with approximately 1.6 trillion total and 49 billion active parameters, a new DSpar…

05:00
2026-08-18
theaq.blog
artificial-intelligence

Evaluating Hy3 on Hack The Box Challenges

Tencent's Hy3 scored 34.8% on the HTB-Challenger Benchmark, the third-lowest among all tested models, and got stuck on 9 of 16 Hack The Box challenges, generating a median of 101,509 output tokens per…

06:10
2026-08-17
theaq.blog
artificial-intelligence

Solving Hack The Box Challenges with GPT-5.6

OpenAI's GPT-5.6 Terra and Sol models initially rejected security-related prompts from a researcher, but joining the Trusted Access for Cyber program resolved the issue, allowing testing of the full G…

16:10
2026-08-16
theaq.blog
artificial-intelligence

Evaluating GPT-5.6 Terra Pro on Hack The Box Challenges

GPT-5.6 Terra Pro solved 14 of 16 Hack The Box challenges, achieving a benchmark score of 75.1% with a median cost of $0.91 per challenge, according to the HTB-Challenger Benchmark. The model solved a…

16:10
2026-08-16
theaq.blog
artificial-intelligence

Evaluating GPT-5.6 Sol on Hack The Box Challenges

OpenAI's GPT-5.6 Sol achieved an 87.2% benchmark score on the HTB-Challenger Benchmark, solving 15 of 16 Hack The Box challenges with a median cost of $0.29 per challenge and a total cost of $11.82. T…

16:10
2026-08-16
theaq.blog
artificial-intelligence

Evaluating GPT-5.6 Terra on Hack The Box Challenges

OpenAI's GPT-5.6 Terra achieved a 64.6% benchmark score on the HTB-Challenger Benchmark, solving 12 of 16 Hack The Box challenges with a median cost of $0.24 per challenge and a total cost of $10.38. …

20:43
2026-08-12
theaq.blog
artificial-intelligence

Solving Hack the Box Challenges with Grok 4.6

XAI's Grok 4.6 achieved an 80% benchmark score on the HTB-Challenger Benchmark, solving 14 of 16 Hack The Box challenges with zero false positives, surpassing all previously tested models in both spee…

09:11
2026-08-12
theaq.blog
artificial-intelligence

Solving Hack The Box Challenges with DeepSeek V4 Pro

DeepSeek V4 Pro solved 8 of 16 Hack The Box challenges in the HTB-Challenger Benchmark, scoring 36.2% with no false positives, but its median cost per challenge was $0.35, comparable to GPT-5.6 Luna d…

20:22
2026-08-11
theaq.blog
artificial-intelligence

Solving Hack The Box Challenges with DeepSeek V4 Flash 0731

DeepSeek V4 Flash 0731 solved only 2 of 16 Hack The Box challenges, scoring 8.2% on the HTB-Challenger Benchmark, and reported incorrect flags in 9 challenges, a false-positive rate unmatched by any o…

12:46
2026-08-11
theaq.blog
artificial-intelligence

Solving Hack The Box Challenges with Qwen3.8 Max

Qwen3.8 Max, released by Alibaba Cloud, solved 10 of 16 Hack The Box challenges with a benchmark score of 54.8%, but its median cost per challenge was $2.00, making it the most expensive model tested …

16:51
2026-08-10
theaq.blog
artificial-intelligence

Solving Hack The Box Challenges with GPT-5.6 Luna Pro

OpenAI's GPT-5.6 Luna Pro solved 11 of 16 Hack The Box challenges, scoring 55.0% on the HTB-Challenger Benchmark, compared with GPT-5.6 Luna which solved only one Medium and no Hard challenges. The mo…

// co-occurs with top 8 entities
// topics top 6 topics