cd/entity/GLM 5.3 Flash· home› entities› GLM 5.3 Flash
grep -l @glm 5.3 flash /news/*.json | wc -l → 35

GLM 5.3 Flash

mentions 35 type Person page 1/2 feed RSS

// recent coverage 35 mentions

00:00
2026-10-08
mindstudio.ai
large-language-models

Step 5 Preview Pricing: API Costs and October 15 Weights Release

Stepfun has set Step 5 Preview API pricing at $1 per million uncached input tokens, $5 per million cached input tokens, and $2.70 per million output tokens, with reasoning tokens billed at the output …

00:00
2026-10-07
mindstudio.ai
large-language-models

Mistral Large 4 (Le Chat): Hands-On Testing of the New Flagship

Mistral AI released Mistral Large 4, its new flagship mixture-of-experts model with roughly a trillion total parameters and about 50 billion active at inference, in public preview via its API and Le C…

07:00
2026-10-03
dotnetperls.com
large-language-models

Problems with Local LLMs

A developer's comparison found that generating complicated Rust functions on a local LLM took 30 to 60 minutes of 100% CPU and 100% GPU usage at an estimated electricity cost of $0.02 to $0.04, while …

15:29
2026-10-02
wagtail.org
ai-tools

One month coding with GLM 5.3 Flash

A one-month experiment to run all of September on the open model GLM 5.3 Flash ended with only 50% of the 2 billion tokens generated going to the target model, according to the developer's AgentsView …

01:55
2026-10-01
news.ycombinator.com
ai-agents

Ask HN: My agent went haywire, does this output mean anything?

A Hacker News user reported that a freshly provisioned Hermes agent running on a headless Debian server with Z.ai's GLM 5.3 Flash model entered a 1m36s "thinking" loop of ostensibly random words and c…

06:57
2026-09-30
twitter.com
large-language-models

GLM 5.3 flash on AMD GPUs 670 tok/s

GLM 5.3 Flash is now live on RunInfra running on AMD GPUs, delivering 670 tokens per second on the Vercel AI Gateway at $0.11 per 1M input tokens, $0.45 per 1M output tokens and $0.03 per 1M cached to…

11:44
2026-09-29
x.com
large-language-models

Qwen3.8-Flash-Next Is on TensorFold with Speed Boosts

TensorFold 0.3.6.2 delivered decode speeds over 62 tokens per second on a single stream and 119 tokens per second across five concurrent streams running Qwen3.8-Flash-Next on a single Nvidia DGX Spark…

07:00
2026-09-27
dotnetperls.com
large-language-models

Experimented with Claude Opus 5.5

A developer reported that Claude Opus 5.5 generated correct Rust code from a spec for $0.09 per final generation, improving the code by switching lookup table values from u32 to u8 for slightly faster…

07:00
2026-09-26
dotnetperls.com
large-language-models

Costs of Online LLM Usage

Spec-driven development on OpenRouter cut the cost of evaluating a Markdown prompt of up to 180 lines and generating a function to between 1 and 7 cents, according to the author's experiments, far bel…

07:00
2026-09-24
dotnetperls.com
large-language-models

Used GLM 5.3 Flash for Code

A developer used OpenRouter to run GLM 5.3 Flash on two Rust functions generated from written specs, paying about $0.03 total for the work. The developer reported the model coded the functions well, w…

08:00
2026-09-23
spectrocloud.com
ai-infrastructure

41 billion tokens later: dogfooding local inference routing

Spectro Cloud reported that 85 of its engineers processed 41 billion tokens in a one-month pilot of its PaletteAI Inference Launchpad, with 40 billion tokens handled locally on a single server with ei…

09:00
2026-09-21
labqoat.com
computer-vision

I'm afraid of spiders. So I made AI look at 2k of them

Gemini 3.8 Flash identified spider species correctly on 997 of 2,000 photos, a 49.85% accuracy rate that was the highest among nine AI models tested in a benchmark built from research-grade iNaturalis…

17:36
2026-09-20
github.com
ai-tools

Directional steering is a runtime activation edit for DS4

Ds4 now supports directional steering, a runtime activation edit that applies a normalized f32 direction per transformer layer during inference, with steering files shaped 43x4096 for DeepSeek V4 Flas…

19:50
2026-09-17
kagifeedback.org
large-language-models

GLM 5.3 Flash unusably slow

Users of GLM 5.3 Flash reported the model has become unusably slow, with response times of at least 3 minutes for simple questions, according to complaints posted on a discussion thread. Commenter shu…

00:00
2026-09-15
mindstudio.ai
large-language-models

Run GLM 5.3 Flash Locally: GSQ and RCO Quantization Explained

An independent research group in Austria has developed two quantization techniques, GSQ and RCO, that compress Z.ai's 320-billion-parameter GLM 5.3 Flash vision-language model from roughly 320GB at fu…

page 1 / 2 next →
// co-occurs with top 8 entities
// topics top 6 topics