cd/sources/runagentrun-auto-discovered· home sources Runagentrun (auto-discovered)
cat /sources/runagentrun-auto-discovered.feed | wc -l → 68

Runagentrun (auto-discovered)

articles 68 domain runagentrun.co.uk → page 2/4 feed RSS
00:00
2026-07-10
runagentrun.co.uk
artificial-intelligence

Try GPT-5.6 Sol for coding this afternoon

OpenAI released GPT-5.6 Sol on July 9, 2026, scoring 59 on the Intelligence Index, one point behind Anthropic's Claude Fable 5 at 60, but at a third of the cost at $1.04 per task versus $2.75. Sol top…

00:00
2026-07-09
runagentrun.co.uk
artificial-intelligence

Grok 4.5 undercuts the frontier on cost

XAI released Grok 4.5 on 8 July 2026, priced at $2 per million input tokens and $6 per million output tokens, making it roughly five times cheaper than competitors like Claude Fable 5 and GPT-5.5. The…

00:00
2026-07-08
runagentrun.co.uk
artificial-intelligence

mistral.rs v0.9.0 outpaces llama.cpp on CPU

Mistral.rs released version 0.9.0 on 7 July, claiming up to 1.8× faster CPU decoding than llama.cpp on both x86 and ARM hardware, challenging the de-facto standard for local LLM inference. The speedup…

00:00
2026-07-07
runagentrun.co.uk
artificial-intelligence

Gemma 4 E2B: three jobs on 4 GB

A practitioner is running Google's Gemma 4 E2B model on a single 4 GB VRAM card to handle screen watching, voice-memo and meeting transcription, and chat simultaneously, consolidating three previously…

00:00
2026-07-05
runagentrun.co.uk
large-language-models

A new workbench for running local AI models

Solo developer released Kivarro, an open-source local inference workbench for running AI models on personal hardware, built on Rust and Tauri and targeting GGUF models. The creator posted it to r/Loca…

00:00
2026-07-04
runagentrun.co.uk
artificial-intelligence

Frontier AI lost a finance test

Bridgewater's AIA Labs and Thinking Machines Lab fine-tuned a Qwen3-235B model on the hedge fund's internal investor judgement, achieving 84.7% accuracy on triage tasks versus 78.2% for the best front…

00:00
2026-07-03
runagentrun.co.uk
large-language-models

A Gemma 4 fine-tune targets marketing copy

A community-built fine-tune of Google's Gemma 4 31B model has beaten the base model by 290 Elo points on the EqBench3 benchmark for marketing copy, according to a Reddit post. The fine-tune leverages …

00:00
2026-07-01
runagentrun.co.uk
ai-agents

NVIDIA Turns BioNeMo Into Agent Tools

NVIDIA turned its BioNeMo life-sciences stack into an agent toolkit, allowing AI agents to call accelerated models and libraries for scientific tasks. Anthropic's Claude Science workbench is the first…

00:00
2026-06-30
runagentrun.co.uk
large-language-models

Anthropic ships Claude on Azure Blackwell racks

Anthropic's Claude models became generally available on Microsoft Foundry, running on NVIDIA GB300 Blackwell Ultra GPUs, as part of a strategic partnership involving a $30 billion Azure compute commit…

00:00
2026-06-30
runagentrun.co.uk
large-language-models

Sonnet 5 closes in on Opus 4.8

Anthropic released Claude Sonnet 5 on 30 June 2026, claiming it nearly matches Opus 4.8 in agentic performance at a lower price. The model is the default on Free and Pro plans, with introductory prici…

00:00
2026-06-28
runagentrun.co.uk
large-language-models

OpenAI ships GPT-5.6 Sol under restricted US access

OpenAI launched GPT-5.6 Sol, Terra, and Luna on June 26, with Sol outperforming Anthropic's Claude Mythos on agentic coding benchmarks. The US government restricted access to trusted partners under a …

00:00
2026-06-28
runagentrun.co.uk
large-language-models

The LLM tier that actually fits your work

Two new 2026 comparisons from DeepInfra and GMI Cloud conclude that the gap between open and closed LLMs has narrowed to 5-10% on overall capability, with no clean leaderboard existing. Closed models …

00:00
2026-06-27
runagentrun.co.uk
ai-agents

Wrap, don't rebuild: AWS's agentic overlay pattern

AWS published a technical how-to on June 25, 2026, detailing a pattern for adding agent-to-agent (A2A) capabilities to existing REST services without rewriting core code. The proposed overlay approach…

00:00
2026-06-26
runagentrun.co.uk
artificial-intelligence

DeepSeek Flash breaks the agent cost curve

Retriever, a browser-agent startup, cut the cost of automated web workflows by over 100x by swapping its planning model from a frontier API to DeepSeek V4 Flash, an openly licensed Chinese model. A mu…

00:00
2026-06-25
runagentrun.co.uk
ai-agents

Anthropic puts a permanent Claude in Slack

Anthropic launched Claude Tag, an always-on AI agent that joins Slack channels as a persistent team member, replacing its older Slack connector. The agent builds long-term memory of projects and decis…

00:00
2026-06-25
runagentrun.co.uk
large-language-models

Gemma 4 outpaces Qwen 3.6 on code review

Google's Gemma 4 31B outperforms Alibaba's Qwen 3.6 27B on agentic code review tasks, finishing faster due to superior Multi-Token Prediction (MTP) design, according to benchmarks and field reports. W…

00:00
2026-06-24
runagentrun.co.uk
artificial-intelligence

Ai2 ships Tmax-27B terminal agent

Ai2 released Tmax-27B on 23 June 2026, an open-weight terminal-agent model built on Qwen3.6-27B that scores 43% on Terminal Bench 2.0 and 69% on TB Lite. The dense 27B model outperforms the sparse 397…

00:00
2026-06-23
runagentrun.co.uk
ai-tools

Sage Router: one endpoint, every model

Earl Co released Sage Router, an open-source self-hosted gateway that exposes a single endpoint for AI agents to route requests to multiple model providers with automatic failover. The tool targets sm…

00:00
2026-06-22
runagentrun.co.uk
artificial-intelligence

Google makes Interactions the default Gemini API

Google promoted its Interactions API to general availability on June 22, 2026, making it the default interface for Gemini models and agents. The new API introduces typed steps, managed agents with Lin…

00:00
2026-06-22
runagentrun.co.uk
ai-agents

A business assistant for under £50 a month

A new open-source AI assistant stack combining the Hermes agent and MiniMax-M3 model on Nous Portal costs under £50 per month and can automate tasks like market briefings, inbox triage, lead research,…

← prev page 2 / 4 next →