cd/entity/Qwen2.5-0.5B-Instruct· home entities Qwen2.5-0.5B-Instruct
grep -l @qwen2.5-0.5b-instruct /news/*.json | wc -l → 11

Qwen2.5-0.5B-Instruct

mentions 11 type Organization feed RSS

// recent coverage 11 mentions

22:37
2026-09-01
github.com
machine-learning

Moe expert offloading on a 2-core Celeron with 2.7GB RAM

A developer's on-hardware test shows that prefetching mixture-of-experts (MoE) expert weights from disk ahead of matmuls can improve token generation speed on a severely resource-constrained machine, …

20:17
2026-08-17
lancedb.com
machine-learning

Data Loading for AI/ML: A Comprehensive Guide

LanceDB's guide to data loading for AI/ML explains the process of moving data into algorithms, focusing on PyTorch and LanceDB, and covers the I/O and CPU stages, noting that I/O is rarely the bottlen…

11:41
2026-06-28
dev.to
large-language-models

I Benchmarked Speculative Decoding — a = 3.5 Wasn't Enough

A developer benchmarked speculative decoding using Qwen2.5-0.5B-Instruct as the draft model and Qwen2.5-1.5B-Instruct as the target model on a CPU. Across code, JSON, and story generation tasks, specu…

17:26
2026-06-02
kyrieblunders.bearblog.dev
machine-learning

I made a kernel 2.2x faster. It made my training loop 3x slower

A developer wrote a fused decode-attention kernel that ran 2.2× faster than the baseline in microbenchmarks, but when integrated into a HuggingFace `generate` call for an RL training loop, the decode …

// co-occurs with top 8 entities
// topics top 6 topics