cd/sources/runagentrun-auto-discovered· home› sources› Runagentrun (auto-discovered)
cat /sources/runagentrun-auto-discovered.feed | wc -l → 84

Runagentrun (auto-discovered)

articles 84 domain runagentrun.co.uk → page 2/5 feed RSS
00:00
2026-08-03
runagentrun.co.uk
artificial-intelligence

Adding metadata filtering fixes RAG's blind spot

A developer writing as Rituraj on Dev Genius in late July tested three retrieval methods on a corpus of paired current and deprecated documents, finding that vector RAG returned both versions with cos…

00:00
2026-08-03
runagentrun.co.uk
artificial-intelligence

Resellers are selling Claude tokens at 10% of list

Chinese API resellers are selling Claude and Codex tokens at up to 90% discounts, or 10% of official prices, according to a Show HN post on 2 August by xiaoxumz11 and a May ChinaTalk analysis by Zilan…

00:00
2026-08-02
runagentrun.co.uk
large-language-models

DeepSeek V4 Flash now runs from a backpack

A community tester ran DeepSeek-V4-Flash-0731 on a Bosgame M5 mini PC with an RTX PRO 6000 Max-Q eGPU, achieving 44 to 60 tokens per second depending on quantisation, with no hyperscaler involved. The…

00:00
2026-07-31
runagentrun.co.uk
artificial-intelligence

DeepSeek V4 Flash sharpens its agent edge

DeepSeek's V4 Flash model graduated from preview to official public-beta release on 31 July 2026, with post-training improvements that boosted its Toolathlon score from 51.8 to 70.3 on DeepSeek's own …

00:00
2026-07-29
runagentrun.co.uk
ai-agents

Agenta ships an open-source AI coworker

Agenta, an MIT-licensed open-source workspace for building and running AI agents, has shipped a self-hosted alternative to Anthropic's Claude Cowork that supports background execution, shared workspac…

00:00
2026-07-29
runagentrun.co.uk
ai-tools

OpenAI open-sources its security agent

OpenAI open-sourced its Codex Security CLI, a command-line tool that scans code for vulnerabilities and suggests patches, under the Apache 2.0 license on Wednesday. The tool, previously codenamed Aard…

00:00
2026-07-26
runagentrun.co.uk
artificial-intelligence

Opus 5 nearly quadruples the ARC-AGI-3 record

Anthropic's Claude Opus 5 scored 30.2% on the ARC-AGI-3 benchmark, nearly quadrupling the previous record of 7.8% set by OpenAI's GPT-5.6 Sol, according to the ARC Prize team. The result marks the mos…

00:00
2026-07-25
runagentrun.co.uk
artificial-intelligence

Opus 5 lands on AWS at half Fable price

Amazon Web Services launched Anthropic's Claude Opus 5 on Amazon Bedrock on 24 July 2026, pricing the model at $5 per million input tokens and $25 per million output tokens — half the cost of the flag…

00:00
2026-07-23
runagentrun.co.uk
artificial-intelligence

AMD bets $5bn on Anthropic to rival Nvidia

AMD is investing up to $5 billion in Anthropic, the AI lab behind Claude, in a deal that includes Anthropic deploying up to two gigawatts of AMD's top-end AI chips starting in the first half of 2027, …

00:00
2026-07-22
runagentrun.co.uk
artificial-intelligence

Qwen 3.6 outranks Gemma 4 on intelligence

Alibaba's Qwen3.6 35B A3B scores 32 on Artificial Analysis's Intelligence Index v4.1, outperforming Google's Gemma 4 26B A4B at 26, with Qwen winning 18 of 22 evaluations tested. However, Gemma 4 is 6…

00:00
2026-07-20
runagentrun.co.uk
large-language-models

Alibaba's Qwen 3.8 targets Kimi K3

Alibaba previewed Qwen 3.8, an open-weight model with over one trillion parameters that handles images, video, documents and text, claiming it is second only to Anthropic's Fable 5. The Qwen team says…

00:00
2026-07-20
runagentrun.co.uk
artificial-intelligence

NVIDIA bets the agent era on one protocol

NVIDIA announced at SIGGRAPH 2026 that six creative tools — Adobe, Affinity by Canva, Blender, Boris FX, Foundry, SideFX and Epic's Unreal Engine — have adopted the Model Context Protocol (MCP), enabl…

00:00
2026-07-18
runagentrun.co.uk
artificial-intelligence

Anthropic halves Fable 5 subscription limits

Anthropic announced on Friday that Claude Fable 5 will remain in Max and Team Premium plans from 20 July but at half the regular usage limits, with Pro and Team Standard subscribers losing access and …

00:00
2026-07-17
runagentrun.co.uk
artificial-intelligence

Most frontier AI is just more compute

A March 2026 study from MIT's Computer Science and Artificial Intelligence Laboratory analyzing 809 large language models released between 2022 and 2025 found that 80 to 90% of frontier AI model perfo…

00:00
2026-07-15
runagentrun.co.uk
artificial-intelligence

The frontier AI duopoly takes shape

Anthropic confidentially filed for an IPO on 1 June, with Polymarket putting the chance of a 2026 listing at 65%, and investor Gavin Baker estimating on the All-In podcast that Anthropic would trade a…

00:00
2026-07-13
runagentrun.co.uk
ai-infrastructure

NVIDIA Vera targets the agent-loop bottleneck

NVIDIA published a new CPU category on 7 July with Vera, an Arm server CPU designed to address the agentic-AI bottleneck by maximizing single-threaded performance rather than core density. In testing …

00:00
2026-07-10
runagentrun.co.uk
artificial-intelligence

Try GPT-5.6 Sol for coding this afternoon

OpenAI released GPT-5.6 Sol on July 9, 2026, scoring 59 on the Intelligence Index, one point behind Anthropic's Claude Fable 5 at 60, but at a third of the cost at $1.04 per task versus $2.75. Sol top…

00:00
2026-07-09
runagentrun.co.uk
artificial-intelligence

Grok 4.5 undercuts the frontier on cost

XAI released Grok 4.5 on 8 July 2026, priced at $2 per million input tokens and $6 per million output tokens, making it roughly five times cheaper than competitors like Claude Fable 5 and GPT-5.5. The…

00:00
2026-07-08
runagentrun.co.uk
artificial-intelligence

mistral.rs v0.9.0 outpaces llama.cpp on CPU

Mistral.rs released version 0.9.0 on 7 July, claiming up to 1.8× faster CPU decoding than llama.cpp on both x86 and ARM hardware, challenging the de-facto standard for local LLM inference. The speedup…

00:00
2026-07-07
runagentrun.co.uk
artificial-intelligence

Gemma 4 E2B: three jobs on 4 GB

A practitioner is running Google's Gemma 4 E2B model on a single 4 GB VRAM card to handle screen watching, voice-memo and meeting transcription, and chat simultaneously, consolidating three previously…

← prev page 2 / 5 next →