Help with qwen and Heretic
A community derivative of Qwen2.5-1.5B-Instruct, processed with Heretic to reduce refusal behavior, is available as a Q4_K_M GGUF file on Hugging Face from user saidutta69. The derivative is intended âŠ
A community derivative of Qwen2.5-1.5B-Instruct, processed with Heretic to reduce refusal behavior, is available as a Q4_K_M GGUF file on Hugging Face from user saidutta69. The derivative is intended âŠ
OpenBMB's MiniCPM5-2B GGUF quantization, tested hands-on with llama.cpp, shows degraded performance on complex tasks: the Q8_0 model invented a nonexistent column in a SQL debugging task and failed a âŠ
Tinfoil (YC X25) proposes verifiable privacy for cloud AI inference pipelines using cryptographic proofs, but the implementation gap lies in inspecting model artifacts before deployment. Local model fâŠ
A developer calculated the actual VRAM requirements for running Llama 3 8B and Gemma 2 9B locally, revealing that the KV cache can consume far more memory than the model weights, especially at longer âŠ