RENDER: Controlling Reader-Facing Evidence in LLM Memory Evaluation
A new benchmark called RENDER, introduced in an arXiv paper (arXiv:2608.23568v1), shows that the format of reader-facing memory artifacts can swing LLM memory evaluation scores by up to 72.6 points. On 500 LongMemEval qu…