cd/entity/GGUF· home entities GGUF
grep -l @gguf /news/*.json | wc -l → 57

GGUF

mentions 57 type Organization page 1/3 feed RSS

// recent coverage 57 mentions

17:00
2026-08-11
promptcube3.com
artificial-intelligence

Unsloth Desktop finally lets us train models locally without a

Unsloth Desktop, a new local AI training tool from Unsloth, enables users to train models locally without a cloud backend, supporting NVIDIA, AMD, Intel, and Mac hardware. It claims a 70% reduction in…

00:00
2026-08-11
mindstudio.ai
artificial-intelligence

How to Run fuse-1 Lite Locally: VRAM, Setup, and Formats

Fuse-1 Lite, a 5.72B parameter mixture-of-experts coding model from LiquidAI, can run locally with VRAM needs ranging from 3.36 GB in 4-bit quantized form to about 12 GB in full bfloat16 precision, ac…

09:00
2026-08-05
fratepietro.com
artificial-intelligence

Building a Rust Inference Engine That Matches Llama.cpp

Developer Antonello F. released Ferrox, a pure-Rust inference engine that runs open LLMs locally on CPU, Apple Metal, or CUDA, achieving performance parity with llama.cpp on an Apple M2 Pro: 26.9 tok/…

19:27
2026-07-29
empero.org
artificial-intelligence

Qwythos-27B-v1: the long-awaited 27B

Empero released Qwythos-27B-v1, an open-weights reasoning model under Apache-2.0, built on Qwen3.5-27B with native multi-token prediction, full vision, and a 1,048,576-token context via YaRN. The 27B …

18:28
2026-07-27
promptcube3.com
artificial-intelligence

Kimi K3 Weights: Initial Deployment Notes

A developer deploying the Kimi K3 model encountered a CUDA out-of-memory error caused by KV cache allocation during initial inference passes, not the model weights themselves. The developer resolved t…

05:18
2026-07-26
kraghavan.ca
large-language-models

Introduction to LLM Inference

A senior engineer with 11 years of distributed systems experience explains the full LLM inference pipeline, from request arrival to text output, detailing the GGUF file structure and the distinction b…

17:03
2026-07-25
promptcube3.com
artificial-intelligence

Open-Weight AI: Model Wars vs Ecosystem Wars

Open-weight AI models offer freedom but require significant effort to deploy, according to a technical guide that argues the real value lies in deployment pipelines and developer ecosystems rather tha…

17:04
2026-07-24
promptcube3.com
ai-infrastructure

AI Infrastructure

Vendor lock-in in AI infrastructure creates technical debt and forces companies into a one-size-fits-all approach, warns a technical guide. To avoid this, the guide recommends implementing a Gateway P…

00:34
2026-07-24
github.com
ai-agents

Show HN: CLI Coding Agent that runs on llamafile and/or GGUF

Kdeps, a CLI coding agent that runs on llamafile and/or GGUF, enables building and deploying AI agents in YAML with workflow (DAG pipelines) and agent (autonomous LLM loop) modes. The open-source tool…

23:20
2026-07-17
github.com
artificial-intelligence

Show HN: Qwen3.6-35B-A3B on a 16 GB M1 Pro with SSD-streamed MoE

A fork of the DwarfStar inference engine, andreaborio/ds4, aims to run large Mixture-of-Experts models like Qwen3.6-35B-A3B on 16–64 GB Apple Silicon Macs by using adaptive SSD streaming and Metal res…

page 1 / 3 next →
// co-occurs with top 8 entities
// topics top 6 topics