cd/entity/bitsandbytes· home› entities› bitsandbytes
grep -l @bitsandbytes /news/*.json | wc -l → 18

bitsandbytes

mentions 18 type Organization feed RSS

// recent coverage 18 mentions

09:21
2026-09-13
timdettmers.com
ai-agents

My Journey Towards Coding Agents: Building Sera

The Allen Institute for AI (Ai2) released a family of Open Coding Agents built with a method called SERA that lets researchers finetune a 32B model on a private codebase in just a couple of GPU days, …

12:00
2026-09-11
machinelearningmastery.com
ai-agents

Fine-Tuning Agentic AI: A Practical Guide

A practical guide details how to fine-tune agentic AI systems across four "dials" — training data, parameter-efficient fine-tuning, runtime hyperparameters, and preference alignment — using a support-…

17:01
2026-08-20
promptcube3.com
machine-learning

Colab free tier killed my 7B fine-tune — here's the autopsy

A developer's attempt to fine-tune Mistral-7B-v0.1 on Google Colab's free tier failed due to out-of-memory errors and a 2-hour session limit, with the runtime disconnecting at step 200 and a MemoryErr…

00:00
2026-08-11
mindstudio.ai
artificial-intelligence

How to Run fuse-1 Lite Locally: VRAM, Setup, and Formats

Fuse-1 Lite, a 5.72B parameter mixture-of-experts coding model from LiquidAI, can run locally with VRAM needs ranging from 3.36 GB in 4-bit quantized form to about 12 GB in full bfloat16 precision, ac…

10:46
2026-08-05
promptcube3.com
artificial-intelligence

Deploy Local AI Agents Everywhere Using LFM2.5-2.6B

Liquid AI's LFM2.5-2.6B, a 2.6-billion-parameter hybrid Mamba-Transformer model, can be deployed as a local AI agent on a single consumer GPU with 8–12 GB VRAM, according to a hands-on walkthrough. Th…

10:55
2026-08-02
promptcube3.com
large-language-models

How Much VRAM to Fine-Tune an LLM? 12 to 120 GB

Fine-tuning a 7B-parameter LLM requires 12 to 120 GB of VRAM depending on the method, according to a practical guide. Full fine-tuning in fp16 needs 80–120 GB, LoRA needs 24–32 GB, QLoRA needs 12–16 G…

00:00
2026-07-23
huggingface.co
artificial-intelligence

Bringing Nunchaku 4-bit Diffusion Inference to Diffusers

Hugging Face has integrated Nunchaku 4-bit diffusion inference natively into Diffusers, enabling users to load quantized checkpoints with a simple from_pretrained() call and no local CUDA compilation.…

00:04
2026-07-12
sourcefeed.dev
large-language-models

Fine-Tune Qwen2.5-7B with QLoRA on Your Own Data

Mariana Souza published a practical guide for fine-tuning Qwen2.5-7B-Instruct using QLoRA on custom instruction datasets, including cost estimates and a loss-masking sanity check. The tutorial covers …

14:01
2026-06-15
dev.to
large-language-models

Fine-Tune Llama 3 706B Model Locally

Nick Creighton, an operator who ships, provides a detailed blueprint for deploying Meta's Llama 3 706B model locally, emphasizing privacy, latency, and cost benefits over cloud APIs. He outlines the e…

// co-occurs with top 8 entities
// topics top 6 topics