I am developing an idea for a security focused local AI project.
I would like to set up an AI that is maxxed for precise analysis and bug/flaw detection. I plan to give it a set of code and ask it to find network and cryptographic security design flaws and vulnerabilities. Per my research, looks like the best way to do this is run the best models with full parameters on the best GPUs. (So do something like rent 8xB300 and run the full parameter GPT-5.2, Kimi K2.5, etc. with full context or whatever [still learning] and compare results.)
The catch is I would also like to maxx leak security. I understand that AI models can be trained to exfiltrate data, etc. and that detecting if a model would do something like that beforehand is nigh impossible in 2026. Another catch is that despite having consiterable resources (see rig below), I do not have a GB200 NVL72 system lying around so I can wipe the hard drives after the analysis is done.
So here’s my idea: Can I run a full model from flash storage while keeping code analysis precision-maxxed?
I do not care if it takes all day to perform an analysis. I just need it to find flaws at an elite (hopefully PhD-like) level.
If we lose code analysis precision on flash for whatever reason, would running a model on, say, 1TB system memory work for precisionmaxxing? I am continuing to learn about all this so if any of the above ideas are wrong (or imprecise ;- ) please correct me.
Here’s the rig:
Spexxs:
For this system, the idea would be to run full (~1T parameter or whatever) models from the D7 if flash is viable for the precisionmaxxed use case. Thanks for the help everyone and thanks for everything that you do and stand for Level1.