cd/entity/DeepSeek-V3· home› entities› DeepSeek-V3
grep -l @deepseek-v3 /news/*.json | wc -l → 56

DeepSeek-V3

mentions 56 type Organization page 2/3 feed RSS

// recent coverage 56 mentions

16:02
2026-08-04
cursor.com
machine-learning

Mixture-of-Kittens: our open-source MoE megakernel for NVL72s

Cursor, the AI coding company, open-sourced Mixture-of-Kittens (MoK), a deterministic MoE training megakernel for NVL72s that fuses communication and computation into a single kernel, achieving a 1.41…

09:02
2026-08-02
github.com
large-language-models

AirLLM: Inference 2.8T Kimi K3 on a single 4GB GPU

AirLLM, an open-source inference library, now supports running the 2.8T-parameter Kimi K3 model on a single 4GB GPU, using only 3.72GB of VRAM on an RTX 6000 Ada, by streaming one expert at a time for…

10:01
2026-07-26
promptcube3.com
artificial-intelligence

Agentic Coding: Benchmarks and Test Process Insights

Claude 3.5 Sonnet is the current gold standard for agentic coding benchmarks, with superior ability to parse stack traces and fix bugs, while GPT-4o often gets stuck in looping behavior and DeepSeek-V…

19:05
2026-07-25
promptcube3.com
machine-learning

MoE Capacity Factor: Why Your Tokens are Being Dropped

Token dropping in mixture-of-experts (MoE) layers occurs when a capacity factor limits each expert's buffer, causing excess tokens to bypass the expert MLP and degrade model quality under production l…

17:46
2026-07-25
promptcube3.com
artificial-intelligence

Hardening AI-generated code for legacy systems: A practical guide

A practical guide recommends hardening AI-generated code for legacy systems by using a multi-model verification loop, cross-model validation, and isolated test generation. The guide suggests using Cla…

13:06
2026-07-25
promptcube3.com
artificial-intelligence

DeepSeek vs Claude for coding, which is better

Claude 3.5 Sonnet is generally superior for complex architectural planning and nuanced refactoring, while DeepSeek-V3 offers industry-leading performance in raw logic and competitive benchmarks at a s…

20:47
2026-07-24
promptcube3.com
artificial-intelligence

Prompt Engineering: A Complete Guide for Beginners

A guide for beginners on prompt engineering outlines techniques such as role prompting, few-shot prompting, and output format control to improve AI responses, noting that GPT-4o excels at few-shot pro…

04:00
2026-07-24
arxiv.org
artificial-intelligence

JAXBench: Benchmarking Autonomous TPU Kernel Optimization

A new benchmark suite called JAXBench, comprising 50 JAX workloads for AI-generated kernel optimization on Google Cloud TPUs, shows that target-specific documentation context matters more than model s…

14:54
2026-07-23
dev.to
machine-learning

Mixture of Experts (MoE) Explained

Mixture of Experts (MoE) is a machine learning architecture that divides a model into specialized subnetworks called experts, with a gating network routing each input token to only the most relevant s…

15:00
2026-07-21
developer.nvidia.com
artificial-intelligence

Setting a World Record for MoE Pre-Training on NVIDIA GB300 NVL72

NVIDIA GB300 NVL72 set a world record for pre-training DeepSeek-V3 671B at 1,648 TFLOPs per GPU, demonstrating how advances across the AI platform—from silicon to networking to software—push training …

18:01
2026-07-20
dev.to
artificial-intelligence

Is China's Open-Weights AI Strategy Actually Winning?

China's open-weights AI strategy is winning, as labs like DeepSeek and Z.ai release competitive models under permissive licenses, building developer ecosystems and bypassing the proprietary API moat o…

← prev page 2 / 3 next →
// co-occurs with top 8 entities
// topics top 6 topics