Ornith-1.0-35B
Ornith AI released Ornith-1.0-35B, a 35B-parameter mixture-of-experts model with 3B active parameters per token, on June 25, 2026, under an MIT license on HuggingFace, featuring a 262K context window.…
Ornith AI released Ornith-1.0-35B, a 35B-parameter mixture-of-experts model with 3B active parameters per token, on June 25, 2026, under an MIT license on HuggingFace, featuring a 262K context window.…
A developer proposes a 100-lens framework for evaluating context-aware AI coding agents, arguing that passing tests alone is insufficient to measure whether an agent understood the current state of a …
Ornith-1.5, a family of open-weight foundation models built on Qwen3.5 and Gemma4, trains through a self-improvement loop that jointly generates tasks, constructs scaffolds, and runs reinforcement lea…
AI coding agents have made building software cheap, but AdaL's developers found that cheap implementation increases the risk of building the wrong thing. After building a benchmark first, they discove…
SWE-bench remains the most widely used open-source benchmark for AI coding agents, with 2,294 tasks from 12 Python repositories, but newer benchmarks like Terminal-Bench, SWE-Bench Pro, and Senior SWE…
NVIDIA has unveiled NOOA (NVIDIA Object-Oriented Agents), an open-source framework that models AI agents as single Python classes, unifying capabilities, state, prompts, and memory. The framework's de…
A new framework from Collinear AI advises teams to choose AI models per workload rather than once company-wide, using a build-vs-buy test across capability, cost, compliance, and control. The guide re…
An AI agent session can consume 50,000–200,000 tokens, compared to a few hundred to a couple thousand for a chat turn, with multi-agent jobs exceeding one million tokens, according to illustrative ran…
Qwen has released Qwen3.8-27B, a 27-billion-parameter multimodal language model with integrated vision capabilities, built on the Qwen3.5 architecture. The model supports native image and video unders…
An engineer argues that passing test suites is insufficient to prove an AI coding agent made the correct engineering decision, highlighting the need for context-adaptation benchmarks. The developer pr…
A new arXiv preprint (2608.08239v1) finds that replay-based evaluation of LLM routers in multi-step agents scores the wrong world, with branching rollouts showing that model swaps alter 61-94% of post…
Researchers introduced Mendel Gödel Machine (MGM), a self-improving coding agent that uses comparative evolution to rewrite its own source code, achieving faster and better convergence than single-tra…
A developer argues that AI-assisted coding is stunting junior engineers' debugging skills, citing an Anthropic study showing AI users scored 50% on comprehension quizzes versus 67% for hand-coders, an…
Tencent's open-source AI agent memory system, TencentDB Agent Memory, reached #1 on GitHub Trending this week, adding nearly 2,000 stars in a single day to reach 14,600 total. The tool provides persis…
Bespoke Labs is hiring a contract researcher to design and evaluate reinforcement learning environments and benchmarks for long-horizon agentic tasks, which require hours, days, or weeks of coherent m…
Asif Waliuddin, an engineer, rolled back his fleet of chief-of-staff agents to an older model version after finding the newer model performed worse in real sessions, despite better benchmark scores. C…
A new study from arXiv (2607.28641v1) introduces the Agentic Formalism Trap and the Evaluative Dissonance Index (D_E), showing that LLM-as-a-Judge systems conflate structural proceduralism with semant…
A developer has open-sourced a framework-agnostic testing methodology for AI agents, including a 61-source benchmark map and 58 universal test blocks across seven tiers, with full coverage of the OWAS…
Anthropic's Claude Sonnet 5 and Claude Opus 5, released in mid-2026, offer distinct strengths and pricing trade-offs. Sonnet 5 excels at high-volume coding and content tasks with a 72.7% SWE-bench Ver…
A developer's context-selection tool, cognitive-cache, initially appeared no better than grep in a small benchmark, but a larger SWE-bench evaluation showed it significantly outperformed grep and TF-I…