LLM evaluation only becomes useful when every model faces the same prompts, the same fixed judge, per-axis rubrics, and a public verbatim trail. The LFORLA Reverse Engineering benchmark does exactly that: it restores C source from stripped binaries, scores with deterministic token similarity against server-only references, and publishes what each model actually answered. Here is how the numbers are made, what the current leaderboard shows, and how to read it without fooling yourself.
The benchmark is called Reverse Engineering (Binary to Source). The task is blunt: given a stripped binary, restore the C source of a function. Models can run in tools or no-tools mode. The scoring is deterministic token similarity against server-only references, so the same answer always gets the same score.
Here is the mechanism, step by step.
That combination is what makes LLM evaluation comparable rather than anecdotal. Same input, same judge, separated axes, inspectable output.
| Model | Provider | Score |
|---|---|---|
| Nemotron 3 Ultra (free, via opencode) | deepseek | 43.48991396255626 |
| GLM 5.2 (our own submission) | lforla | 78.0 |
We submitted our own model to this benchmark, so the second row is not a neutral third party. It is our own result, published with the same scoring rules. The chart attached to this post plots the leaderboard for the reverse-engineering benchmark.
The gap is large: 78.0 versus roughly 43.49. In practice, that means GLM 5.2 recovered substantially more of the reference source under deterministic token similarity. Because the judge is fixed and the references are server-only, the difference is not a matter of style or verbosity. It is a matter of how much of the actual function structure the model reconstructed.
A few things are worth noting about the shape of the leaderboard.
If you are picking a model for reverse engineering or adjacent code tasks, do not read the overall score as a universal ranking. Read it as a signal about this task, under this judge, with these references. Ask three questions. First, does the benchmark match your workload? Binary-to-source restoration is not the same as writing a web app. Second, can you inspect the verbatim trail? If you cannot read what the model actually answered, you are trusting a number you cannot verify. Third, does the benchmark separate axes? A single score hides whether the model failed at reasoning or at tool use.
The LFORLA benchmark is useful precisely because it answers those questions in public. The same prompts, the same fixed judge, per-axis rubrics, and a readable trail mean you can compare models without taking anyone's word for it.
This benchmark is narrow. It measures one task: restoring C source from stripped binaries. It does not measure general chat quality, long-context reasoning, or safety. The scoring is deterministic token similarity, which is stable but not the same as human judgment of correctness. A model could produce a functionally equivalent function that scores lower because the tokens differ. The leaderboard currently shows a single external model plus our own submission, so it is not a broad field survey. And because we submitted our own model, our result should be read with that disclosure in mind.
LLM evaluation becomes comparable when the mechanism is boring and public: same prompts, same fixed judge, per-axis rubrics, and a verbatim trail anyone can re-read. The LFORLA Reverse Engineering benchmark does that, and the current numbers show GLM 5.2 at 78.0 against Nemotron 3 Ultra at 43.48991396255626. Read the trail, not just the score. Explore the benchmark and the full leaderboard at https://lforla.org.