17:27
2026-08-22
runtimewire.com
artificial-intelligence
FreeToken reports 22-25 tok/s for a 284B DeepSeek model on one RTX 5090
FreeToken, an open-source inference engine developed by UC Berkeley researchers including Shuo Yang and Xiaoze Fan, reports decoding DeepSeek-V4-Flash, a 284B-parameter Mixture-of-Experts model, at 22โฆ