Coding agents craft arbitrary code so securing them is more complicated than red-teaming. We post trained a cyber-security small llm, changed how it reasons and supplemented our controls using program analysis techniques such as inline reference monitoring to outperform GPT5.5-xhigh on hard benchmarks like LinuxArena and SleightBench.
Free product available at harden.run and full benchmarks in the blog post.
Comments URL: [https://news.ycombinator.com/item?id=49472151](https://news.ycombinator.com/item?id=49472151)
Points: 2