cd/entity/Mixtral· home› entities› Mixtral
grep -l @mixtral /news/*.json | wc -l → 26

Mixtral

mentions 26 type Organization page 1/2 feed RSS

// recent coverage 26 mentions

11:40
2026-09-22
discuss.huggingface.co
artificial-intelligence

SplitMOE- similar to deepseek-MOE `same same but different`

Developer Priyanshu-5257 released SplitMoE, an open-source Mixture of Experts routing architecture that splits wide feed-forward layers into smaller, fine-grained expert partitions, published on GitHu…

16:46
2026-09-14
promptcube3.com
machine-learning

NVIDIA Transformer Engine makes JAX MoE training actually fast

NVIDIA Transformer Engine is now integrated with JAX to accelerate dropless Mixture of Experts (MoE) training, targeting the conditional computation bottleneck in architectures used by DeepSeek, Qwen,…

22:36
2026-09-11
github.com
artificial-intelligence

Show HN: Bypassing Transformer Softmax via Static Contraction

A Proof-of-Concept repository published on GitHub under the project jax-softmax-bypass claims to bypass the Transformer Softmax's transcendental exponential function through a static algebraic contrac…

05:05
2026-08-12
dev.to
large-language-models

Self-Hosted LLM on a $5 VPS in 2026: What Actually Works

A developer's guide to self-hosting LLMs on budget VPS plans in 2026 finds that quantized 3B and 7B models are now viable on servers under $6 per month, with RAM bandwidth as the key bottleneck. The a…

14:01
2026-08-03
pub.towardsai.net
artificial-intelligence

Mixture of Experts: How Sparse Gating Enables 1T-Parameter LLMs

Mixture of Experts (MoE) architectures enable 1-trillion-parameter large language models by activating only a subset of expert networks per token, achieving 8× parameter count with only ~2× compute. M…

04:00
2026-07-24
arxiv.org
artificial-intelligence

JAXBench: Benchmarking Autonomous TPU Kernel Optimization

A new benchmark suite called JAXBench, comprising 50 JAX workloads for AI-generated kernel optimization on Google Cloud TPUs, shows that target-specific documentation context matters more than model s…

14:54
2026-07-23
dev.to
machine-learning

Mixture of Experts (MoE) Explained

Mixture of Experts (MoE) is a machine learning architecture that divides a model into specialized subnetworks called experts, with a gating network routing each input token to only the most relevant s…

01:33
2026-07-23
petals.dev
large-language-models

Run large language models at home, BitTorrent‑style

Petals, a project from the BigScience research workshop, lets users run large language models like Llama 3.1 (up to 405B parameters) at home by distributing model parts across a peer-to-peer network, …

07:12
2026-07-10
dev.to
large-language-models

Large Language Models Demystified: A Visual and Practical Guide

A developer published a visual and practical guide to large language models, explaining that they are computer programs trained on massive text datasets to predict the next word. The guide covers how …

page 1 / 2 next →
// co-occurs with top 8 entities
// topics top 6 topics