cd/entity/Qwen3-8B· home› entities› Qwen3-8B
grep -l @qwen3-8b /news/*.json | wc -l → 99

Qwen3-8B

mentions 99 type Organization page 1/5 feed RSS

// recent coverage 99 mentions

04:00
2026-10-07
arxiv.org
artificial-intelligence

WavePrune: One period is often enough for RoPE

WavePrune, a method that restricts each RoPE channel to its first rotation period, raises the HELMET long-context score on four of five tested models without extra tuning, including 35.7 to 40.0 on Qw…

23:58
2026-10-02
runtimewire.com
artificial-intelligence

Ai2 releases an open 8B model for cited science reports

The Allen Institute for AI released AstaBrief 8B on October 2nd, an open model built on Qwen3-8B that turns research questions and retrieved scientific papers into cited reports and now powers Fast mo…

15:19
2026-10-02
huggingface.co
artificial-intelligence

Open-sourcing AstaBrief, the fast report-generation model in Asta

The Allen Institute for AI (Ai2) open-sourced AstaBrief 8B, an 8-billion-parameter model built on Qwen3-8B that turns a research question and retrieved literature excerpts into a cited scientific repo…

04:00
2026-10-01
arxiv.org
artificial-intelligence

DEdit: Iterative Draft Editing for Speculative Decoding

DEdit, a diffusion-based speculative decoding drafter from an arXiv paper (arXiv:2609.38510v1), achieves macro-average speedups of 5.72x on Qwen3-4B and 5.97x on Qwen3-8B over autoregressive generatio…

21:26
2026-09-25
cryptobriefing.com
artificial-intelligence

Stanford and Nvidia’s CLM-8B model runs up to 9x faster than Jev

Stanford and Nvidia released CLM-8B on September 23, an open-weights Contrastive Language Model that makes real-time agent decisions up to 9x faster than TypeSafe AI's proprietary Jev model, clocking …

00:00
2026-09-18
seangoedecke.com
large-language-models

Two techniques for working with System One models

Software engineer Sean Goedecke outlined two techniques for programming with "System One" language models — models that output only decisions from user-provided multiple-choice questions — in a post d…

12:31
2026-09-13
pub.towardsai.net
large-language-models

Latency Optimization Levers for Open-Weight LLM inference:Part-2

A second-part technical article on latency optimization for open-weight LLM inference details four techniques measured on a fixed deployment of Qwen3-8B (Apache-2.0) served on vLLM through SageMaker's…

21:41
2026-09-10
promptcube3.com
artificial-intelligence

Osprey boosts speculative decoding acceptance rates by 16% to 22%

Osprey, a target-agnostic speculative decoding method detailed in arXiv:2609.09338v1, raises mean acceptance length by 16.1% on Qwen3-8B, 21.2% on Llama-3.3-70B-Instruct, and 22.7% on MiniMax-M2.5 (22…

03:43
2026-09-09
eiln.github.io
ai-infrastructure

Getting 50 GB/S Back Out of the Neural Engine

An RTL performance erratum in the Apple M3 Neural Engine throttles DRAM weight streaming throughput to 17–19 GB/s from the nominal 45–60 GB/s when total weight size is an integer multiple of 1 MiB, af…

02:35
2026-09-08
github.com
machine-learning

Show HN: Zero downtime embedding model upgrades

A developer has released EmbedFlow, an open-source tool that enables zero-downtime upgrades of embedding models by retrieving top-K documents with the old model and reranking them with the new model, …

page 1 / 5 next →
// co-occurs with top 8 entities
// topics top 6 topics