When Agents Slow Down: Understanding LLM Agents' Test-Time Strategies via Elo-per-token Analysis A new analysis method called Elo-per-token measures how large language model agents allocate test-time compute as they revise solutions, use tools, explore alternatives, and decide when to stop, addressing the difficulty of measuring agent performance scaling on open-ended tasks that provide continuous scores. The research frames agent test-time strategy as adaptive and introduces Elo-per-token as the metric for tracking it. Large language model LLM agents allocate test-time compute adaptively as they revise solutions, use tools, explore alternatives, and decide when to stop. This test-time strategy makes it difficult to measure how agent performance scales. We study open-ended tasks that provide continuous scores for