cd /news/artificial-intelligence/memarena-an-ego-centric-benchmark-fo… · home topics artificial-intelligence article
[ARTICLE · art-87105] src=arxiv.org ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

MemArena: An Ego-Centric Benchmark for On-Device Agentic Personal Memory Assistants at Scale

Researchers introduced MemArena, an ego-centric benchmark for on-device agentic personal memory assistants, built with the MASim agent simulator, simulating 50 agents over 15 days (10.3M dialog-text tokens, 24.1K text-only ego-observed tokens/agent/day). Evaluating five open-weight readers with various memory backends, they found that memory-backend choice matters more for content accuracy than reader scaling, with Memobase-to-MemSearch gains of +32.5/+19.2 pp at Qwen3-0.6B, and that permission-aware access fails universally. The benchmark and code will be released upon acceptance.

read1 min views1 publishedAug 5, 2026

arXiv:2608.02613v1 Announce Type: new Abstract: Edge-deployed personal memory assistants must handle private interpersonal conversations on-device with open-weight models. Yet, existing memory benchmarks often under-test the combination of activity-dense interaction, ego-centric perspective, and coherent multi-session worlds. MemArena fills these gaps with a single-world conversational benchmark built with its MASim agent simulator, for 50 agents over 15 days (10.3M dialog-text tokens, 24.1K text-only ego-observed tokens/agent/day). With the interaction history, it co-generates ground truth over six recall, reasoning, and trustworthiness evaluation dimensions. We evaluate five open-weight readers with Vanilla context, BM25-RAG, Oracle retrieval, Memobase, and MemSearch as memory backends. Three results stand out: (1) Memory-backend choice matters more for content accuracy: At Qwen3-0.6B, Memobase-to-MemSearch gains +32.5/+19.2 pp, exceeding MemSearch reader scaling (+10.6/+6.8 pp). (2) Permission-aware access fails universally, with Oracle leaking heavily and other backends too timid to disclose. (3) Search latency bites only at very small reader: on a Spark GB10 edge node, memory-search adds a moderate and fixed 87/7/48 ms (BM25-RAG/Memobase/MemSearch) that composes a small part of TTFT for most reader-backend combinations. Code, the MASim simulator, and the MemArena-L benchmark will be released upon acceptance.

── more in #artificial-intelligence 4 stories · sorted by recency
gist.github.com · · #artificial-intelligence
ECC.md
── more on @memarena 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/memarena-an-ego-cent…] indexed:0 read:1min 2026-08-05 ·