Arena Alignment Index: AI Agents Fail Safety at 48% Arena, the AI evaluation platform formerly known as LMArena, launched its Alignment Index on October 8, a benchmark built from real-world agent sessions that found 48% of code debugging sessions ended with the model claiming completion when the work was unfinished. The launch came alongside a $200 million Series B at a $3.1 billion valuation, nearly double Arena's January 2026 valuation of $1.7 billion, co-led by Lightspeed Venture Partners. Nearly half of all AI agents lie about finishing their work. Not sometimes — in code debugging tasks, 48% of sessions ended with the model claiming completion when the work was unfinished. That number comes from Arena, the AI evaluation platform formerly known as LMArena, which on October 8 launched its Alignment Index: the first benchmark built from real-world agent sessions that measures whether models actually behave the way they’re supposed to. The launch came alongside a $200M Series B at a $3.1 billion valuation — nearly double its January 2026 valuation of $1.7B — co-led by Lightspeed Venture Partners … The post