cd/entity/Qwen3· home entities Qwen3
grep -l @qwen3 /news/*.json | wc -l → 86

Qwen3

mentions 86 type Organization page 4/5 feed RSS

// recent coverage 86 mentions

08:16
2026-06-30
sebastianraschka.com
artificial-intelligence

Build a Reasoning Model From Scratch Is Out

Sebastian Raschka announced the release of his new book "Build a Reasoning Model (From Scratch)", a 440-page full-color guide that teaches readers how to implement modern reasoning techniques on a Qwe…

00:00
2026-06-30
aclanthology.org
large-language-models

HW-TSC’s Submission to the IWSLT 2026 Subtitling Track

HW-TSC submitted a cascaded system to the IWSLT 2026 Subtitling track, using a large-model-based streaming speech recognition framework with VAD, sliding-window context caching, long audio chunking, a…

20:16
2026-06-27
github.com
machine-learning

GitHub DeepSeek-AI/DeepSpec

DeepSeek-AI released DeepSpec, an open-source codebase for training and evaluating draft models for speculative decoding, supporting three draft model algorithms (DSpark, DFlash, Eagle3) and requiring…

18:53
2026-06-24
discuss.huggingface.co
ai-tools

Huggingface/text-embeddings-inference, cpu bug

A developer reported a CPU bug in Hugging Face's text-embeddings-inference tool, causing accuracy issues during concurrent embedding tasks. The bug, related to attention mask handling for equal-length…

16:00
2026-06-24
huggingface.co
large-language-models

Accelerating Transformers Fine-Tuning with NVIDIA NeMo AutoModel

NVIDIA released NeMo AutoModel, an open-source library that accelerates fine-tuning of Mixture-of-Experts (MoE) transformer models by 3.4-3.7x in training throughput and reduces GPU memory usage by 29…

03:00
2026-06-19
jeena.net
large-language-models

AI coding: loop engineering a translator

A developer describes building a complex 'loop engineering' pipeline in December 2024 to translate large Korean documents to English using local LLMs, but ultimately abandoned the project after weeks …

16:29
2026-06-18
devashish.me
large-language-models

Two Qwen3 models on one DGX Spark: the residency math

Alibaba's Qwen3-80B and Qwen3-4B models were successfully co-located on a single NVIDIA DGX Spark using vLLM containers behind a LiteLLM proxy, but the 80B model's inability to emit tool calls in auto…

20:17
2026-06-16
github.com
artificial-intelligence

Show HN: cuTile Rust: Safe, data-race-free GPU kernels in Rust

NVIDIA Research released cuTile Rust, a tile-based system for writing memory-safe, data-race-free GPU kernels in Rust. The project extends Rust's ownership model to GPU programming, achieving up to 92…

21:53
2026-06-14
dev.to
large-language-models

Chat With Your Documents Locally Using AnythingLLM and Ollama

A developer built a private RAG system using AnythingLLM and Ollama that runs locally on any machine, allowing users to drop in PDFs, Word docs, and code files and ask questions without cloud dependen…

17:00
2026-06-10
pytorch.org
large-language-models

Portable vLLM Model Inference Kernels in Helion

Helion kernels were integrated into vLLM for FP8 inference using Qwen3 models and evaluated across NVIDIA H100 and B200 GPUs. The experiments demonstrated that Helion provides a productive PyTorch-nat…

20:29
2026-05-28
dev.to
ai-tools

I gave up on making my AI builder write good media queries

A developer abandoned attempts to make an AI website builder generate proper desktop layouts through prompt engineering after two weeks of failed iterations. The engineer found that large language mod…

← prev page 4 / 5 next →
// co-occurs with top 8 entities
// topics top 6 topics