cd/entity/Qwen3-8B· home› entities› Qwen3-8B
grep -l @qwen3-8b /news/*.json | wc -l → 99

Qwen3-8B

mentions 99 type Organization page 2/5 feed RSS

// recent coverage 99 mentions

04:16
2026-09-07
pub.towardsai.net
large-language-models

Latency Optimization Levers for Open-Weight LLM Inference: Part-1

Serving an open-weight large language model in a user-facing product is primarily a latency problem, and open weights enable engineers to control inference speed through levers such as weight quantiza…

14:00
2026-08-31
kdnuggets.com
large-language-models

Speed Up LLM Inference with DSpark Speculative Decoding

DeepSeek's DSpark speculative decoding technique, which combines parallel drafting with a lightweight sequential component, can improve local LLM generation speed on the same GPU, with DeepSeek report…

13:34
2026-08-27
alphaxiv.org
artificial-intelligence

A general tensor-structured compression scheme for efficient LLMs

Researchers introduced Tensor Mixture (MixT), a general tensor-structured compression scheme that replaces dense linear layers in large language models with natively executable mixtures of tensor oper…

03:01
2026-08-19
dev.to
artificial-intelligence

How I Cut AI API Costs 95% — A Data Scientist's Field Guide

A data scientist at an unnamed company cut AI API costs by 95% by analyzing six months of logs and implementing a model-routing pipeline that matches each request to the cheapest adequate model. The a…

16:32
2026-08-18
dev.to
artificial-intelligence

The Cheapest AI APIs in 2026: A Bootcamp Grad's Deep Dive

A bootcamp graduate's deep dive into the cheapest AI APIs of 2026 reveals that ultra-budget models like Qwen3-8B and GLM-4-9B cost as little as $0.01 per million output tokens, while the sweet spot ti…

← prev page 2 / 5 next →
// co-occurs with top 8 entities
// topics top 6 topics