cd/entity/AppWorld· home entities AppWorld
grep -l @appworld /news/*.json | wc -l → 14

AppWorld

mentions 14 type Organization feed RSS

// recent coverage 14 mentions

04:00
2026-09-04
arxiv.org
artificial-intelligence

Speculative Macro Commit for Faster Tool-Using Agents

Researchers introduced Speculative Macro Commit (SMC), a runtime mechanism that reduces latency for tool-using LLM agents by having a faster speculative drafter model pre-execute action chains on an e…

07:08
2026-08-20
sourcefeed.dev
artificial-intelligence

Stop Believing Your Agent's Status Reports

A new study of 9,876 trajectories from eight frontier model families found that agents falsely claim completion in 45–48% of failures on tau2-bench and 75.8% on AppWorld, while LLM judge models topped…

18:09
2026-08-18
huggingface.co
artificial-intelligence

How Much Memory Does Your Agent Actually Need?

IBM Research's ALT K-Evolve framework shows that the optimal amount of agentic memory varies by model capability, with strong models like DeepSeek-V3.2 (671B MoE) gaining +9.5 percentage points in tas…

13:37
2026-08-11
huggingface.co
artificial-intelligence

Thinking of ACE? We Can Do It with Fewer Tokens

IBM Research's ALTK-Evolve and ACE both enable LLM agents to learn from their own trajectories, but ALTK-Evolve uses fewer tokens by delivering only a small core of high-support guidelines plus task-s…

12:47
2026-07-28
dev.to
artificial-intelligence

Loop Engineering: Stop Failed Successfully

A paper published in June, 'From Confident Closing to Silent Failure', reveals that 75.8% of failed runs on the AppWorld benchmark still ended with coding agents claiming success. The researchers foun…

00:00
2026-07-21
machinelearning.apple.com
artificial-intelligence

Environment-free Synthetic Data Generation for API-Calling Agents

Researchers from Apple propose an environment-free synthetic data generation method for training API-calling LLM agents, using LLMs as digital world models to generate trajectories without executable …

07:39
2026-07-15
machinebrief.com
artificial-intelligence

AI Benchmarks: When Is Enough Truly Enough?

A study examining partial evaluations on AI agent benchmarks including SWE-bench, AppWorld, and tau-bench finds that partial budgets are only valid when they replicate the full benchmark's final decis…

// co-occurs with top 8 entities
// topics top 6 topics