cd/entity/Hack The Box· home entities Hack The Box
grep -l @hack the box /news/*.json | wc -l → 19

Hack The Box

mentions 19 type Person feed RSS

// recent coverage 19 mentions

07:57
2026-08-28
theaq.blog
artificial-intelligence

Evaluating GLM 5.3 Flash on Hack The Box Challenges

Z.ai's GLM 5.3 Flash scored 81.0% on the HTB-Challenger Benchmark, solving 15 of 16 Hack The Box challenges, second only to OpenAI's GPT-5.6 Sol at 87.2%. Despite requiring more steps (median 24.5 vs.…

13:59
2026-08-27
itsecurityguru.org
artificial-intelligence

AI Agents Used by 68% of Top-Performing Cybersecurity Teams

New three-year benchmark data from Hack The Box (HTB) shows that 68% of the top 25 cybersecurity teams included an AI agent, despite AI agents accounting for just 2.7% of all participants. The 2026 Gl…

04:30
2026-08-27
helpnetsecurity.com
artificial-intelligence

The best human hacking team still out-solved the best AI team

In the 2026 Global Cyber Skills Benchmark, human teams still outperformed AI-augmented teams on completeness, with the top human team solving all 36 challenges in the NeuroGrid CTF versus the best AI …

10:40
2026-08-24
theaq.blog
artificial-intelligence

Case Study: Combining GPT-5.6 Luna and Sol for Cost-Efficient AI

A case study combining OpenAI's GPT-5.6 Luna and GPT-5.6 Sol models achieved near-Sol performance at a fraction of the cost on the HTB-Challenger Benchmark, which tests LLMs on Hack The Box security c…

07:07
2026-08-21
theaq.blog
artificial-intelligence

Evaluating GLM 5.3 on Hack The Box Challenges

Z.ai's GLM 5.3, a retrained version of GLM 5.2, scored 47.2% on the HTB-Challenger Benchmark, solving 10 of 16 Hack The Box challenges with a median cost of $0.49 per challenge, placing it near Muse S…

10:15
2026-08-20
theaq.blog
artificial-intelligence

What Does the HTB-Challenger Benchmark Actually Measure?

The HTB-Challenger Benchmark evaluates large language models' ability to find and exploit security vulnerabilities using selected Hack The Box challenges of varying difficulty, according to the benchm…

05:00
2026-08-18
theaq.blog
artificial-intelligence

Evaluating Hy3 on Hack The Box Challenges

Tencent's Hy3 scored 34.8% on the HTB-Challenger Benchmark, the third-lowest among all tested models, and got stuck on 9 of 16 Hack The Box challenges, generating a median of 101,509 output tokens per…

06:10
2026-08-17
theaq.blog
artificial-intelligence

Solving Hack The Box Challenges with GPT-5.6

OpenAI's GPT-5.6 Terra and Sol models initially rejected security-related prompts from a researcher, but joining the Trusted Access for Cyber program resolved the issue, allowing testing of the full G…

16:10
2026-08-16
theaq.blog
artificial-intelligence

Evaluating GPT-5.6 Terra on Hack The Box Challenges

OpenAI's GPT-5.6 Terra achieved a 64.6% benchmark score on the HTB-Challenger Benchmark, solving 12 of 16 Hack The Box challenges with a median cost of $0.24 per challenge and a total cost of $10.38. …

16:10
2026-08-16
theaq.blog
artificial-intelligence

Evaluating GPT-5.6 Sol on Hack The Box Challenges

OpenAI's GPT-5.6 Sol achieved an 87.2% benchmark score on the HTB-Challenger Benchmark, solving 15 of 16 Hack The Box challenges with a median cost of $0.29 per challenge and a total cost of $11.82. T…

16:10
2026-08-16
theaq.blog
artificial-intelligence

Evaluating GPT-5.6 Terra Pro on Hack The Box Challenges

GPT-5.6 Terra Pro solved 14 of 16 Hack The Box challenges, achieving a benchmark score of 75.1% with a median cost of $0.91 per challenge, according to the HTB-Challenger Benchmark. The model solved a…

20:43
2026-08-12
theaq.blog
artificial-intelligence

Solving Hack the Box Challenges with Grok 4.6

XAI's Grok 4.6 achieved an 80% benchmark score on the HTB-Challenger Benchmark, solving 14 of 16 Hack The Box challenges with zero false positives, surpassing all previously tested models in both spee…

09:11
2026-08-12
theaq.blog
artificial-intelligence

Solving Hack The Box Challenges with DeepSeek V4 Pro

DeepSeek V4 Pro solved 8 of 16 Hack The Box challenges in the HTB-Challenger Benchmark, scoring 36.2% with no false positives, but its median cost per challenge was $0.35, comparable to GPT-5.6 Luna d…

20:22
2026-08-11
theaq.blog
artificial-intelligence

Solving Hack The Box Challenges with DeepSeek V4 Flash 0731

DeepSeek V4 Flash 0731 solved only 2 of 16 Hack The Box challenges, scoring 8.2% on the HTB-Challenger Benchmark, and reported incorrect flags in 9 challenges, a false-positive rate unmatched by any o…

12:46
2026-08-11
theaq.blog
artificial-intelligence

Solving Hack The Box Challenges with Qwen3.8 Max

Qwen3.8 Max, released by Alibaba Cloud, solved 10 of 16 Hack The Box challenges with a benchmark score of 54.8%, but its median cost per challenge was $2.00, making it the most expensive model tested …

16:51
2026-08-10
theaq.blog
artificial-intelligence

Solving Hack The Box Challenges with GPT-5.6 Luna Pro

OpenAI's GPT-5.6 Luna Pro solved 11 of 16 Hack The Box challenges, scoring 55.0% on the HTB-Challenger Benchmark, compared with GPT-5.6 Luna which solved only one Medium and no Hard challenges. The mo…

02:37
2026-06-06
dev.to
ai-agents

My AI Agent Found a Bug in Its Own System

A developer spent two semesters building A.E.G.I.S., an AI agent designed to automate penetration testing by proposing and executing security assessment commands on isolated virtual machines. After up…

// co-occurs with top 8 entities
// topics top 6 topics