cd/entity/llama-server· home› entities› llama-server
grep -l @llama-server /news/*.json | wc -l → 23

llama-server

mentions 23 type Organization page 1/2 feed RSS

// recent coverage 23 mentions

19:09
2026-09-28
returneditor.ai
ai-tools

Why we removed Ollama from Return (and what replaced it)

Return, a document analysis tool for lawyers, replaced Ollama with a bundled llama-server from llama.cpp as a sidecar process after Ollama 0.12 introduced cloud models in September 2025 that proxy req…

15:03
2026-09-27
dev.to
ai-tools

Mixer — write in text, get 3D. Editor based on local LLMs

Developer ArtemPodloboshnikov released Mixer, an open-source 3D editor that generates GLB/GLTF models from text prompts using local LLMs via a built-in llama-server. The MIT-licensed Tauri v2 applicat…

03:06
2026-09-22
gist.github.com
large-language-models

Qwen 3.8 27B on RTX 5090 at 90-120tps

A developer published a llama-server configuration that runs a Qwen 3.8 27B NVFP4 model with MTP speculative decoding on an RTX 5090, reporting throughput of 90-120 tokens per second. The setup uses a…

20:15
2026-09-17
dev.to
large-language-models

I let a local 27B LLM audit and fix my Splunk + Sysmon stack

A security analyst studying for CySA+ demonstrated that a local 27B-parameter LLM running entirely on their own GPU could audit and remediate a Splunk and Sysmon home SOC stack without any data leavin…

13:12
2026-09-16
forum.level1techs.com
large-language-models

VBR k/v cache is actually usable (buun-llama)

A user running Qwen3.8-Flash-Next at IQ3 quantization on a system with 64GB of RAM, an AMD 9950X3D, a 9070XT, and a ZFS stripe of three mid-range NVMe drives reported fitting 255K tokens of context us…

19:02
2026-09-15
forum.level1techs.com
ai-infrastructure

Homelab critique

A homelab user identified as gessha detailed a home setup that uses two Nvidia RTX 3090 GPUs in an i5-8600K workstation to host Qwen 3.8 27B and Qwen 3-VL 4B models via llama-server, alongside three m…

10:14
2026-09-14
dev.to
ai-tools

llama.cpp vs Ollama in 2026: Which Runtime Should You Run?

A technical comparison examines the tradeoffs between Ollama and llama.cpp for local LLM inference, framing the choice as one between a managed model service and a toolkit operated directly. The guide…

00:00
2026-08-18
mindstudio.ai
artificial-intelligence

Qwen3.8-27B Ridge Quant: Run the Full Model on 12GB VRAM

Empero-ai released Qwen3.8-27B-Ridge-3.7bpw.gguf, a quantized version of the 27-billion-parameter Qwen3.8-27B hybrid reasoning and vision model that shrinks from roughly 50GB to 11.7GB, enabling it to…

07:00
2026-07-26
dotnetperls.com
developer-tools

Using ui-mcp-proxy in Llama-cpp

A developer created a simple MCP server for local LLMs to call Rust-coded tools, but encountered CORS errors. The llama-cpp feature ui-mcp-proxy, passed as an argument to llama-server, sets up a proxy…

19:40
2026-07-09
gist.github.com
artificial-intelligence

Tess localbench

Tess-4-27B, a reasoning-native agentic fine-tune on Qwen3.6-27B, achieved 81% (122/150) on the local benchlocal full suite, ranking #1. The model scored 94% on core instruction/tools/structure tasks b…

00:00
2026-07-07
runagentrun.co.uk
artificial-intelligence

Gemma 4 E2B: three jobs on 4 GB

A practitioner is running Google's Gemma 4 E2B model on a single 4 GB VRAM card to handle screen watching, voice-memo and meeting transcription, and chat simultaneously, consolidating three previously…

09:36
2026-06-18
dev.to
developer-tools

llama-bench skipped FA on capable GPUs — b9437 corrects it

Build b9437 of llama.cpp fixes two default-value bugs in llama-bench that caused flash attention to be skipped on capable GPUs and GPU-layer count to use a legacy sentinel. The flash attention flag no…

07:23
2026-05-29
github.com
ai-tools

Econd-Brain-MCP

Second-brain-mcp is a self-maintaining personal knowledge database that uses MCP, DuckDB, and biological memory models to automatically link, compress, and index saved papers, notes, and figures. The …

page 1 / 2 next →
// co-occurs with top 8 entities
// topics top 6 topics