15:29
2026-08-20
promptcube3.com
artificial-intelligence
ProgramBench reverse-engineering benchmark puts LLMs to the test
ProgramBench, a new reverse-engineering benchmark, compiles and fuzzes LLM-generated code against original binaries, with GPT-4o achieving a 34% pass@1 rate, Claude 3.5 Sonnet 28%, and DeepSeek-Coder-ā¦