cd/entity/GLM-5· home entities GLM-5
grep -l @glm-5 /news/*.json | wc -l → 28

GLM-5

mentions 28 type Organization page 1/2 feed RSS

// recent coverage 28 mentions

18:09
2026-08-18
huggingface.co
artificial-intelligence

How Much Memory Does Your Agent Actually Need?

IBM Research's ALT K-Evolve framework shows that the optimal amount of agentic memory varies by model capability, with strong models like DeepSeek-V3.2 (671B MoE) gaining +9.5 percentage points in tas…

15:12
2026-08-17
usewire.io
artificial-intelligence

Context window blindness: agents can't see their limits

A June 2026 study by LightSpeed and Tencent found that four frontier models—Claude Sonnet 4.5, DeepSeek-V4-Pro, GLM-5, and Gemini-3-Flash—misjudged their total context size with median relative error …

23:50
2026-08-15
promptcube3.com
artificial-intelligence

Vernor Vinge predicted the Singularity by 2030 and he was

Vernor Vinge, a computer scientist and science fiction author, predicted in his 1993 essay that the Singularity—a point where machines surpass human intelligence and accelerate progress beyond human c…

04:51
2026-08-10
newsletter.semianalysis.com
artificial-intelligence

Ultra-High Interactivity on NVIDIA GPUs? - TileRT InferenceX

TileRT's persistent engine on NVIDIA GPUs achieves up to 500 tokens/s/user on the InferenceX GLM5 FP8 744B benchmark on a single B200 decode server, approximately 3× faster than GB300 NVL72 running tr…

04:00
2026-08-03
arxiv.org
artificial-intelligence

Chain-of-Models: Cross-Model Auditing for Bias-Robust LLM Judges

A new arXiv preprint (2607.28636v1) introduces Chain-of-Models (CoM), an automated audit pipeline in which a second model inspects a first model's reasoning trace to reduce cognitive biases in LLM jud…

04:13
2026-07-15
dev.to
large-language-models

I Was Shocked I'm Overpaying for AI by 40x as a Bootcamp Grad

A bootcamp graduate discovered they were overpaying for AI by up to 40x after analyzing API costs. By switching from GPT-4o to models like DeepSeek V4 Flash via Global API, the developer reduced a $50…

19:00
2026-07-13
dev.to
large-language-models

Let Me Show You Which AI Model Actually Writes the Best Code

A developer benchmarked 10 large language models on five coding tasks, finding that DeepSeek V4 Flash offers the best value-to-quality ratio at $0.25 per million output tokens, while Qwen3-Coder-30B e…

00:14
2026-07-12
dev.to
large-language-models

Migrating Off OpenAI: A Backend Engineer's Notes From Production

A backend engineer migrated three production services from OpenAI to DeepSeek V4 Flash via a Global API endpoint, reducing monthly costs from $500 to approximately $12.50—a 40× price difference—while …

21:26
2026-07-10
machinebrief.com
artificial-intelligence

Optimizing GLM-5: Why Bigger Isn't Always Better

OpenClaw's GLM-5 inference optimization study found that adjusting parameters like chunked prefill size and request concurrency improved throughput and reduced latency, cutting serving costs by 10.4% …

11:58
2026-07-01
dev.to
large-language-models

I Cut My AI Bill 97.5% in One Afternoon — And You Can Too

A developer cut their monthly AI bill from $487.92 to $12.50 by switching from OpenAI's GPT-4o to DeepSeek V4 Flash via the Global API, achieving a 97.5% cost reduction. The migration required changin…

16:36
2026-06-30
dev.to
large-language-models

How I Found the Best AI Coding Model Without Going Broke

A bootcamp graduate tested ten AI coding models on five tasks, scoring them on code quality, readability, and explanation clarity. The experiment found that DeepSeek V4 Flash and DeepSeek Coder offer …

00:02
2026-06-30
dev.to
large-language-models

From $500 to $12.50: My Real Migration Off OpenAI in 2026

A developer migrated from OpenAI's GPT-4o to DeepSeek V4 Flash via a global API provider, reducing monthly costs from $487 to $12.50 with only two lines of code changed. The switch required only alter…

09:37
2026-06-27
dev.to
large-language-models

Cutting OpenAI Costs From Scratch: What Nobody Tells You

A B2B SaaS startup cut its LLM inference costs by 97% by switching from GPT-4o to cheaper alternatives like DeepSeek V4 Flash, reducing a $14,200 monthly OpenAI bill to an estimated $355. The develope…

10:43
2026-06-26
dev.to
large-language-models

I Wish I Knew About This OpenAI Swap Sooner — Full Breakdown

An engineer at a company using OpenAI's GPT-4o for LLM inference discovered they were overpaying by up to 40x compared to alternatives like DeepSeek V4 Flash served through Global API. After benchmark…

19:45
2026-06-22
arxiv.org
large-language-models

Self-Harness: Harnesses That Improve Themselves

Researchers introduced Self-Harness, a new paradigm enabling LLM-based agents to iteratively improve their own operating harnesses without human intervention. In tests on Terminal-Bench-2.0, Self-Harn…

page 1 / 2 next →
// co-occurs with top 8 entities
// topics top 6 topics