cd/entity/Qwen3· home› entities› Qwen3
grep -l @qwen3 /news/*.json | wc -l → 123

Qwen3

mentions 123 type Organization page 5/7 feed RSS

// recent coverage 123 mentions

10:07
2026-07-14
machinebrief.com
large-language-models

Why Continual Learning Models Aren't as Smart as We Think

New experiments with Qwen3 models reveal that continual learning in language models struggles to retain earlier knowledge, with accuracy dropping to 1% for bare-statement facts after 20 updates, while…

04:07
2026-07-13
machinebrief.com
artificial-intelligence

Cracking the Code: The New Frontier of Integer Quantization

A new signed symmetric quantization method eliminates clipping errors in low-bit AI models without the runtime penalty of asymmetric quantization, achieving up to 2.45 times faster inference on AMD EP…

20:25
2026-07-10
machinebrief.com
artificial-intelligence

Exploring New AI Pathways: TREK's Innovative Approach

Researchers introduced TREK (Teacher-Routed Exploration via Forward KL), a new AI training method that enhances learning through unconventional exploration strategies. TREK significantly improved perf…

20:24
2026-07-10
machinebrief.com
artificial-intelligence

TREK: A New Path in AI Problem Solving

Researchers introduced TREK, a method that improves AI problem-solving by using verified output trajectories to extend model learning. TREK boosted Qwen3-8B's performance on AIME 2024 from 36.9 to 40.…

01:32
2026-07-10
digital-foundry-eight.vercel.app
large-language-models

I benchmarked every model that fits on an iPhone

An independent benchmark of on-device LLMs on iPhone A17 Pro found Apple's system model achieves ~149 tok/s with only 12MB peak app memory, while 4B-class open models like Qwen3 4B and Llama 3.2 3B tr…

10:15
2026-07-09
ianbarber.blog
large-language-models

MOPD

Researchers released the official Multi-Teacher On-Policy Distillation (MOPD) paper, which composes multiple capabilities into a single large language model by training domain-expert teachers independ…

23:50
2026-07-08
letsdatascience.com
artificial-intelligence

Hugging Face Speeds Transformers Inference in vLLM

Hugging Face announced that its Transformers modeling backend for vLLM now matches or exceeds native vLLM speed for compatible architectures, reducing the need for separate hand-optimized serving port…

00:00
2026-07-08
huggingface.co
artificial-intelligence

Native-speed vLLM transformers modeling backend

Hugging Face announced that the transformers vLLM modeling backend now matches or exceeds native vLLM throughput for many LLM architectures, allowing model authors to run their transformers implementa…

11:52
2026-07-07
github.com
machine-learning

Show HN: TurboQuant for mlx-lm (Apple Silicon)

A developer released TurboQuant, a pip-installable adapter for mlx-lm on Apple Silicon that uses a randomized Hadamard transform for data-oblivious, calibration-free quantization of weights and KV cac…

19:00
2026-06-30
discuss.huggingface.co
large-language-models

Local LLM on MacBook M5 Pro - Totally New to This!

A non-programmer named Tim is setting up a local LLM on a MacBook M5 Max with 128GB unified memory, using Docker Desktop with Model Runner, Open WebUI, and models like Gemma 4 and Qwen3 30B-A3B-Q4_k_m…

11:14
2026-06-30
theoremsearch.com
ai-research

TheoremGraph: Search 18M+ Mathematical Dependencies

Researchers at the University of Washington released TheoremGraph, a unified dependency graph spanning 18 million+ mathematical statements from arXiv papers and the Lean formal proof assistant. The gr…

08:16
2026-06-30
sebastianraschka.com
artificial-intelligence

Build a Reasoning Model From Scratch Is Out

Sebastian Raschka announced the release of his new book "Build a Reasoning Model (From Scratch)", a 440-page full-color guide that teaches readers how to implement modern reasoning techniques on a Qwe…

00:00
2026-06-30
aclanthology.org
large-language-models

HW-TSC’s Submission to the IWSLT 2026 Subtitling Track

HW-TSC submitted a cascaded system to the IWSLT 2026 Subtitling track, using a large-model-based streaming speech recognition framework with VAD, sliding-window context caching, long audio chunking, a…

20:16
2026-06-27
github.com
machine-learning

GitHub DeepSeek-AI/DeepSpec

DeepSeek-AI released DeepSpec, an open-source codebase for training and evaluating draft models for speculative decoding, supporting three draft model algorithms (DSpark, DFlash, Eagle3) and requiring…

← prev page 5 / 7 next →
// co-occurs with top 8 entities
// topics top 6 topics