cd/entity/llama-server· home› entities› llama-server
grep -l @llama-server /news/*.json | wc -l → 23

llama-server

mentions 23 type Organization page 2/2 feed RSS

// recent coverage 23 mentions

01:00
2026-05-20
dev.to
large-language-models

Unload All llama.cpp Router Models Without Restarting

While llama.cpp's router mode allows loading and unloading individual models via HTTP API calls to `/models/unload`, there is no built-in "unload all" endpoint. The recommended approach for unloading …

05:15
2026-05-04
gist.github.com
large-language-models

MTP benchmark

This article presents benchmark results comparing the performance of a Qwen3.6 model running in standard mode versus with Multi-Token Prediction (MTP) enabled. The MTP configuration with a draft of 3 …

← prev page 2 / 2
// co-occurs with top 8 entities
// topics top 6 topics