How a Ticket System Became My Agentic AI Lab
A developer built Lutions, a self-hosted ticket system, as a testing ground for AI coding agents in mature codebases. The project explores how agents can work effectively within established convention…
A developer built Lutions, a self-hosted ticket system, as a testing ground for AI coding agents in mature codebases. The project explores how agents can work effectively within established convention…
A new arXiv preprint (2609.04518v1) finds that in multi-harness reinforcement learning for coding agents, the evaluation harness is the dominant variable, moving the mean solve rate from 2.14% to 9.27…
Arize AI and Atlan report that enterprises running over 100 million AI evaluations monthly still face agent failures because model choice alone cannot ensure reliability, with context and harness engi…
Researchers propose LivePlan, a system that monitors and corrects programming agents in real time, improving issue resolution rates by up to 15.2% (average 9.9%) over vanilla SWE-agent across SWE-benc…
An engineer at Anthropic distills hard-won lessons from production AI agents including Claude Code, OpenHands, and SWE-agent into a practical field guide for building enterprise-ready agents. The guid…
Modem, a startup founded in early 2025, has built a 680,000-line TypeScript codebase with 99.9% of code generated by LLMs, and shares findings on how coding agents like Claude Code and Codex navigate …
GitHub's Copilot code review performance declined after swapping in shared CLI tools (grep, glob, view) due to poorly adapted instructions, but rewriting those instructions for pull request review wor…
Researchers at iNLP-Lab have developed PACT (Protocolized Action-state Communication and Transmission), a new inter-agent communication strategy for large language model-based multi-agent systems that…
A new open-source coding agent, mini-SWE-agent, achieves up to 74% on the SWE-bench verified benchmark using just 100 lines of Python code. Developed by the Princeton and Stanford team behind SWE-benc…
Princeton University and Stanford researchers released SWE-agent, an open-source AI agent that autonomously fixes GitHub issues and has earned 19,310 GitHub stars since its NeurIPS 2024 debut. The pro…