08:16
2026-07-28
byteiota.com
artificial-intelligence
Neutrino-1 8B: 763 tok/s Without Standard Quantization
Fermion Research released Neutrino-1 8B, a 3.88 GB language model trained natively with ternary weights ({-1, 0, +1}) that achieves 763 tokens per second on an H100 with speculative decoding and 24.9 โฆ