cd/entity/GLM-4.7-Flash· home› entities› GLM-4.7-Flash
grep -l @glm-4.7-flash /news/*.json | wc -l → 12

GLM-4.7-Flash

mentions 12 type Organization feed RSS

// recent coverage 12 mentions

04:00
2026-09-28
machinebrief.com
artificial-intelligence

RAZOR: Pruning Replaceable Experts in LLMs

RAZOR, a training-free expert pruning method for mixture-of-experts (MoE) large language models, achieved the highest nine-task macro average among evaluated pruning methods in all eight settings test…

17:59
2026-09-21
twitter.com
ai-tools

Halo: Post-train LLMs 3x faster than TRL and Megatron

White Circle launched Halo, a post-training framework for open-source models that delivers up to 2.8x the throughput of stock TRL with lower peak memory while keeping models in native HuggingFace form…

21:08
2026-09-16
uprouter.online
ai-products

Editorial re-verification: GLM Coding (China)

Zhipu AI's GLM platform lists GLM-5.3 at 8 yuan per million input tokens, 2 yuan per million cached tokens and 28 yuan per million output tokens with a 1M-token context on its official bigmodel.cn pri…

20:00
2026-09-15
akitaonrails.com
large-language-models

Novo LLM Benchmark v4: retestando 39 LLMs (Parte 2)

AkitaOnRails published Part 2 of its LLM Benchmark v4, retesting 39 large language models on a seven-sprint Rails app suite seeded with 14 real CVE-based sabotages, after the author rejected the entir…

00:00
2026-09-04
llmstatus.ai
large-language-models

GLM-4.7-Flash deprecated, retires 2026-09-10

Z.ai (Zhipu) has deprecated its GLM-4.7-Flash model, which will retire on September 10, 2026, six days from now. The model, released on January 19, 2026, supports a 200k context window and 131k max ou…

01:01
2026-07-26
promptcube3.com
large-language-models

HotPin: Running 120B MoE on 24GB RAM

A new llama.cpp patch called HotPin enables running a 120-billion-parameter mixture-of-experts (MoE) model on just 24GB of RAM, achieving up to 67% memory savings and a 45% speedup over standard swapp…

20:14
2026-06-30
dev.to
large-language-models

GLM Is the New Hotness, So Let's Test It On the Homelab

GLM, the model family from Z.ai (formerly Zhipu AI), is gaining attention for its open weights and strong benchmark numbers. A developer tested GLM-4.7-Flash, a 30B-A3B MoE model, on a single-GPU home…

12:31
2026-05-27
github.com
large-language-models

I ran GLM-5.1 on a 16GB RAM machine

A team of engineers successfully ran the 754-billion parameter GLM-5.1 large language model on a consumer PC with only 16GB of RAM and a Ryzen 5 5600G CPU, achieving zero crashes or out-of-memory erro…

// co-occurs with top 8 entities
// topics top 6 topics