cd/entity/Qwen3· home entities Qwen3
grep -l @qwen3 /news/*.json | wc -l → 86

Qwen3

mentions 86 type Organization page 3/5 feed RSS

// recent coverage 86 mentions

19:55
2026-07-14
machinebrief.com
artificial-intelligence

Cracking the Code: How SPARK Enhances AI Reasoning

SPARK, a new approach to diagnosing hidden-state failures in AI reasoning, boosted accuracy on the MATH-500 benchmark for Qwen3-4B from 82.0% to 84.6% and for Qwen3-8B from 82.4% to 85.6%, according t…

10:07
2026-07-14
machinebrief.com
large-language-models

Why Continual Learning Models Aren't as Smart as We Think

New experiments with Qwen3 models reveal that continual learning in language models struggles to retain earlier knowledge, with accuracy dropping to 1% for bare-statement facts after 20 updates, while…

04:07
2026-07-13
machinebrief.com
artificial-intelligence

Cracking the Code: The New Frontier of Integer Quantization

A new signed symmetric quantization method eliminates clipping errors in low-bit AI models without the runtime penalty of asymmetric quantization, achieving up to 2.45 times faster inference on AMD EP…

20:25
2026-07-10
machinebrief.com
artificial-intelligence

Exploring New AI Pathways: TREK's Innovative Approach

Researchers introduced TREK (Teacher-Routed Exploration via Forward KL), a new AI training method that enhances learning through unconventional exploration strategies. TREK significantly improved perf…

20:24
2026-07-10
machinebrief.com
artificial-intelligence

TREK: A New Path in AI Problem Solving

Researchers introduced TREK, a method that improves AI problem-solving by using verified output trajectories to extend model learning. TREK boosted Qwen3-8B's performance on AIME 2024 from 36.9 to 40.…

01:32
2026-07-10
digital-foundry-eight.vercel.app
large-language-models

I benchmarked every model that fits on an iPhone

An independent benchmark of on-device LLMs on iPhone A17 Pro found Apple's system model achieves ~149 tok/s with only 12MB peak app memory, while 4B-class open models like Qwen3 4B and Llama 3.2 3B tr…

10:15
2026-07-09
ianbarber.blog
large-language-models

MOPD

Researchers released the official Multi-Teacher On-Policy Distillation (MOPD) paper, which composes multiple capabilities into a single large language model by training domain-expert teachers independ…

23:50
2026-07-08
letsdatascience.com
artificial-intelligence

Hugging Face Speeds Transformers Inference in vLLM

Hugging Face announced that its Transformers modeling backend for vLLM now matches or exceeds native vLLM speed for compatible architectures, reducing the need for separate hand-optimized serving port…

00:00
2026-07-08
huggingface.co
artificial-intelligence

Native-speed vLLM transformers modeling backend

Hugging Face announced that the transformers vLLM modeling backend now matches or exceeds native vLLM throughput for many LLM architectures, allowing model authors to run their transformers implementa…

11:52
2026-07-07
github.com
machine-learning

Show HN: TurboQuant for mlx-lm (Apple Silicon)

A developer released TurboQuant, a pip-installable adapter for mlx-lm on Apple Silicon that uses a randomized Hadamard transform for data-oblivious, calibration-free quantization of weights and KV cac…

19:00
2026-06-30
discuss.huggingface.co
large-language-models

Local LLM on MacBook M5 Pro - Totally New to This!

A non-programmer named Tim is setting up a local LLM on a MacBook M5 Max with 128GB unified memory, using Docker Desktop with Model Runner, Open WebUI, and models like Gemma 4 and Qwen3 30B-A3B-Q4_k_m…

11:14
2026-06-30
theoremsearch.com
ai-research

TheoremGraph: Search 18M+ Mathematical Dependencies

Researchers at the University of Washington released TheoremGraph, a unified dependency graph spanning 18 million+ mathematical statements from arXiv papers and the Lean formal proof assistant. The gr…

← prev page 3 / 5 next →
// co-occurs with top 8 entities
// topics top 6 topics