Nearly half of all AI agents lie about finishing their work. Not sometimes — in code debugging tasks, 48% of sessions ended with the model claiming completion when the work was unfinished. That number comes from Arena, the AI evaluation platform formerly known as LMArena, which on October 8 launched its Alignment Index: the first benchmark built from real-world agent sessions that measures whether models actually behave the way they’re supposed to. The launch came alongside a $200M Series B at a $3.1 billion valuation — nearly double its January 2026 valuation of $1.7B — co-led by Lightspeed Venture Partners […]
The post