GLM-5.2 – How to Run Locally
Z.ai released GLM-5.2, a 744B-parameter open model with 40B active parameters and a 1M context window, claiming it matches or exceeds proprietary models like Claude 4.8 Opus and GPT-5.5. Unsloth relea…
Z.ai released GLM-5.2, a 744B-parameter open model with 40B active parameters and a 1M context window, claiming it matches or exceeds proprietary models like Claude 4.8 Opus and GPT-5.5. Unsloth relea…
Unsloth launched Unsloth Studio, a desktop application for Mac and Windows that runs AI models offline, supporting GGUF and Safetensors formats with tool-calling, web search, and an OpenAI-compatible …
Seven open-source AI projects—Ollama, Open WebUI, Browser Use, vLLM, Unsloth, CrewAI, and Continue—are reshaping production software development in June 2026. Ollama, with 174,000+ GitHub stars, now o…
A developer built AETHER, a fully offline, voice-controlled AI assistant that runs three local language models on a laptop to control a PC, send WhatsApp messages, and perform desktop automation witho…
A guide explains how to run the Qwen 3.6 35B A3B model on an RTX 3080 with 16GB VRAM using Llama.cpp, offloading most layers to CPU to fit within memory constraints. The author details steps for insta…
NVIDIA has optimized Google DeepMind's new DiffusionGemma model to run up to four times faster on its GeForce RTX GPUs, RTX PRO workstations, and DGX Spark systems. Unlike traditional language models …
A Reddit user reported that llama.cpp build b9455 achieved 67-81 tokens per second on a dual RTX 3090 setup running Unsloth's Qwen3.6-27B-UD-Q8_K_XL model, matching the speed of vLLM for multi-GPU inf…
Developer Matt Coles has built lgtmaybe, a provider-agnostic PR reviewer supporting six model backends with a single --provider flag, shipping as a PyPI CLI and GitHub Action. He also maintains a home…
A developer benchmarked the Qwen3.6 27B model on Modal using llama.cpp, deploying a serverless pipeline that downloads GGUF shards from Hugging Face and runs perplexity evaluation on an A100-80GB GPU.…
A developer's project to modernize 40-year-old COBOL mainframe code into Python microservices using a fully offline, local AI agent. The developer used the open-source Gemma 4 model, loaded via Unslot…