cd/entity/llama-server· home entities llama-server
grep -l @llama-server /news/*.json | wc -l → 13

llama-server

mentions 13 type Organization feed RSS

// recent coverage 13 mentions

00:00
2026-08-18
mindstudio.ai
artificial-intelligence

Qwen3.8-27B Ridge Quant: Run the Full Model on 12GB VRAM

Empero-ai released Qwen3.8-27B-Ridge-3.7bpw.gguf, a quantized version of the 27-billion-parameter Qwen3.8-27B hybrid reasoning and vision model that shrinks from roughly 50GB to 11.7GB, enabling it to…

07:00
2026-07-26
dotnetperls.com
developer-tools

Using ui-mcp-proxy in Llama-cpp

A developer created a simple MCP server for local LLMs to call Rust-coded tools, but encountered CORS errors. The llama-cpp feature ui-mcp-proxy, passed as an argument to llama-server, sets up a proxy…

19:40
2026-07-09
gist.github.com
artificial-intelligence

Tess localbench

Tess-4-27B, a reasoning-native agentic fine-tune on Qwen3.6-27B, achieved 81% (122/150) on the local benchlocal full suite, ranking #1. The model scored 94% on core instruction/tools/structure tasks b…

00:00
2026-07-07
runagentrun.co.uk
artificial-intelligence

Gemma 4 E2B: three jobs on 4 GB

A practitioner is running Google's Gemma 4 E2B model on a single 4 GB VRAM card to handle screen watching, voice-memo and meeting transcription, and chat simultaneously, consolidating three previously…

09:36
2026-06-18
dev.to
developer-tools

llama-bench skipped FA on capable GPUs — b9437 corrects it

Build b9437 of llama.cpp fixes two default-value bugs in llama-bench that caused flash attention to be skipped on capable GPUs and GPU-layer count to use a legacy sentinel. The flash attention flag no…

07:23
2026-05-29
github.com
ai-tools

Econd-Brain-MCP

Second-brain-mcp is a self-maintaining personal knowledge database that uses MCP, DuckDB, and biological memory models to automatically link, compress, and index saved papers, notes, and figures. The …

01:00
2026-05-20
dev.to
large-language-models

Unload All llama.cpp Router Models Without Restarting

While llama.cpp's router mode allows loading and unloading individual models via HTTP API calls to `/models/unload`, there is no built-in "unload all" endpoint. The recommended approach for unloading …

05:15
2026-05-04
gist.github.com
large-language-models

MTP benchmark

This article presents benchmark results comparing the performance of a Qwen3.6 model running in standard mode versus with Multi-Token Prediction (MTP) enabled. The MTP configuration with a draft of 3 …

// co-occurs with top 8 entities
// topics top 6 topics