Unsloth: Introducing AMD support
Unsloth announced AMD GPU support for local LLM training and inference, enabling 500+ models to run up to 2× faster with 70% less VRAM on AMD Radeon, Instinct, Ryzen, and data center GPUs. The release…
Unsloth announced AMD GPU support for local LLM training and inference, enabling 500+ models to run up to 2× faster with 70% less VRAM on AMD Radeon, Instinct, Ryzen, and data center GPUs. The release…
Thinking Machines' open Inkling model, with 975 billion total parameters and 41 billion active, fits on a single high-memory box at 2-bit or 3-bit quantization, unlike the datacenter-scale Kimi K3. Re…
A developer deploying a fine-tuned Llama 3.2 3B medical AI model offline on Android for Yoruba, Hausa, Igbo, and Nigerian Pidgin encountered three distinct crashes during GGUF conversion—tokenizer mis…
Researchers fine-tuned LLaMA 3 (8B) as a drop-in reranker for RAG pipelines, achieving gains of 14% in answer relevancy, 16% in context precision, 19% in answer similarity, and 21% in answer correctne…
A student fine-tuned Meta's Llama 3.1 8B model for multi-step mathematical reasoning using Unsloth, LoRA, and a 'Silent Coder' approach, all within the RAM limits of a free Google Colab instance with …
A new course chapter introduces reinforcement learning (RL) and its application to training large language models (LLMs), explaining core concepts such as agent, environment, action, reward, and polic…
Amazon Web Services announced a partnership with Unsloth to deploy quantized foundation models on Amazon SageMaker AI, reducing memory usage and serving costs while maintaining accuracy. The collabora…
UmarTransit-1B, the first open-source large language model fine-tuned for public transit systems and GTFS data, has been released. Built by fine-tuning Qwen2.5-1.5B-Instruct using QLoRA, the model spe…
A practitioner is running Google's Gemma 4 E2B model on a single 4 GB VRAM card to handle screen watching, voice-memo and meeting transcription, and chat simultaneously, consolidating three previously…
A community-built fine-tune of Google's Gemma 4 31B model has beaten the base model by 290 Elo points on the EqBench3 benchmark for marketing copy, according to a Reddit post. The fine-tune leverages …
A developer explains how to run large language models locally, recommending Gemma 4 and Qwen 3.6 families for most users with 24-64GB memory, and noting that Gemma 4 31B is best for non-coding tasks w…
A developer details practical QLoRA fine-tuning using Axolotl and Unsloth, explaining how parameter-efficient methods like LoRA and QLoRA enable training multi-billion parameter models on a single con…
On May 16, 2026, llama.cpp merged Multi-Token Prediction (MTP) support, enabling 1.7x to 2.4x faster local inference for Qwen3.6 27B models with no accuracy loss or extra downloads. The MTP head is em…
A developer recommends Unsloth as the most cost-effective method for fine-tuning small language models in 2026, citing its ease of use and low VRAM requirements compared to the theoretically cheaper b…
A developer is seeking advice on the optimal prompt format for training the Unsloth/Phi-3.5-mini-instruct model, currently using a custom template with JSON input and output fields. The choice of form…
Unsloth released a guide on June 18 to run Z.ai's 744-billion-parameter GLM-5.2 model on local hardware using aggressive quantization, compressing the model from 1.51 TB to as low as 217 GB. The tooli…
Developer Torgeir Helgevold fine-tuned a 600-million-parameter local LLM (Qwen 3:0.6B) to classify household questions into metadata categories, achieving 92% accuracy on a test set—up from 10% with p…
A developer fine-tuned a tiny 0.6B-parameter Qwen 3 model to categorize household questions into metadata categories like pool, car, and HVAC. The baseline model achieved only 10% accuracy via prompti…
A developer tested agentic AI coding with DeepSeek V4 Flash on GMI Cloud, completing a data processing task in 3 minutes at $0.034 with two mistakes, compared to a human attempt taking an hour with fo…
Microsoft, Hugging Face, Meta's PyTorch team, NVIDIA, and others launched OpenEnv, an open protocol for agent learning environments that standardizes how agents practice and improve. The protocol aims…