Have a model read full agent traces to rewrite prompts, not just RL GEPA, a prompt-optimization method in which a model reads an agent's full trace — reasoning, tool calls and errors — and rewrites the prompt, reportedly doubled the gains of GRPO after a single round on three examples, according to its creator, versus GRPO's 25,000 rollouts. The claim appeared alongside other agent-development items, including DeepSeek's DSec sandbox platform paper, which sustains more than 380,000 concurrent sandboxes and coordinates with GPU training to curb reward hacking, and Epoch AI's Furniture Assembly Benchmark, whose top score rose from 28% to 80% in ten months. GEPA has a model read an agent's full trace, including reasoning, tool calls and errors, then rewrite the prompt. Its creator says one round on three examples doubled the gains GRPO reached after 25,000 rollouts. Watch: GEPA has a model read an agent's full trace, including reasoning, tool calls and errors, then rewrite the prompt. Its creator says one round on three examples doubled the gains GRPO reached after 25,000 rollouts. Read: DeepSeek published a paper on DSec, the production sandbox platform behind its agentic RL training. It sustains more than 380,000 concurrent sandboxes and coordinates with GPU training to curb reward hacking. Watch: Prime Intellect's Elie Bakouch set Claude Code and Codex on the Optimizer Speedrun. Both beat the human record, but neither invented a new optimizer; they recombined known ideas for small gains. Read: Epoch AI's new Furniture Assembly Benchmark asks models to find the mistake in a half-built piece of furniture from a photo and the manual. The top score rose from 28% to 80% in ten months. Read: In a guest post on Terence Tao's blog, cryptographer Amit Sahai argues frontier AI already produces original mathematical ideas faster than people can absorb them, and calls for training far more mathematicians. Read: Cua released a stable Cua Driver for Omarchy that puts an agent's synthetic cursor inside the Hyprland compositor, so an agent can work in a background window while the user's pointer stays free. Read: Wafer reports its GLM-5.2 endpoint averaged 379 ms versus 674 ms for Gemma 4 31B on Cerebras in Y Combinator's AI Office Hours, and says users talked 2.5 minutes longer per session.