Stop lying to yourself with mocks and use a staging cluster for agent evals
Monday.com AI Engineering Director Dor Cohen said in a September 15 session that agent evaluations are unreliable when run against mocked APIs and databases, citing an agent tasked with retrieving 600…