How far are local models from Opus 5.5? Even on expensive hardware? Running an open-weight model at the level of Opus 5.5 requires roughly 610 GB of VRAM plus RAM for a maximally quantized Kimi K3, while the full-precision model would need 1600 GB, according to a Level1Techs forum discussion. A commenter citing benchmarks said GLM 5.3 is "decent" with close to 1 TB of VRAM, prompting another participant to note that 1 TB of RAM only yields decent performance. The thread also questioned whether matching GPT and Claude on long-running goals remains unsolved, and whether a Gorgon Halo system with 196 GB of memory or a Threadripper Halo system with 4 MI350Ps would be required. I’m just trying to figure out what kind of current hardware would be required to match or somehow exceed Opus 5.5? As it seems to be the new peak of solid general AI coding performance? I know matching the software capabilities of GPT / Claude for long running goals is another matter, or has that been figured out yet? I haven’t tried running linux/open code on a while on my RX 9070 XT system, but previous it was quite a pain and barely worked, didn’t try a lemonade server though. Like Might Gorgon Halo with 196GB of memory be alright? Or would someone need that new Threadripper Halo system with 4 MI350ps? Or the NVidia equivalent. Kimi K3 is probably the best you can do with an open model. If you quantize it as hard as possible, it still needs 610 GB VRAM + RAM. The full-precision model would need 1600 GB. mietzen https://forum.level1techs.com/u/mietzen 3 I don’t know how much you can give on these benchmarks, but if you trust them and got close to one TB of VRAM GLM 5.3 is decent: I was hoping someone had some real world uses, but holy crap 1TB of RAM only gives you decent???