Mana: 2-3 Seconds to Feeling Human A developer has built Mana, a voice AI assistant that runs entirely on a local machine with an 8GB VRAM GPU, achieving 2-3 second response times by using a single 4B model for reasoning, code generation, and explanation in one pass. The system, inspired by fictional AIs like Alice from Sword Art Online and JARVIS, has been running 24/7 for three months without crashes. so I shipped a voice AI assistant that runs entirely on my machine. no cloud, no APIs, no latency nightmares. the original idea came from Alice in Sword Art Online — an AI that feels like an actual person, not a chatbot. mixed with JARVIS's anticipation and Neuro-sama's quirky personality. here's what actually went into getting from "wouldn't it be cool" to "this runs 24/7 without issues." most voice assistants are cloud-first: you speak → sent to server → processed → response → back to you. each hop adds latency. you're looking at 3-6 seconds before you hear anything. for a voice interaction, that's dead. it kills the feeling of talking to something intelligent. I wanted something faster. something that responds . the constraint: do it locally. use an 8GB VRAM GPU, run everything on-device, no external APIs except for the live2d avatar bits because that's hard to render locally and still look good . here's the reality: I have a GPU with 8GB VRAM. no budget to experiment with better cards or more models. so every architecture decision was forced by what actually fits. naive approach: chain multiple specialized models. User speaks → Transcription model Whisper → Planning model 3B: what should I do? → Coding model 7B: generate implementation → Verification model 4B: is this correct? → TTS speak the answer math: 1s + 2s + 3s + 1.5s = 7.5s of latency before the user hears anything. nope. the problem isn't just that each model is slow. it's model loading overhead . every time you swap from one model to another, you: with only 8GB, this gets gnarly fast. the constraint was hardware. 8GB VRAM. no more, no less. that forced clarity: pick one model that does everything, or pick nothing. so I went with a single model 4B by default, with 7B/8B quality modes available that does reasoning + code generation + explanation in one pass . latency: ~2-3 seconds total. actually conversational. this is the difference between a chatbot and a companion. JARVIS doesn't pause for 6 seconds before responding. neither does Mana. turns out, when you're forced to optimize for latency because you only have 8GB to work with , you accidentally build something that feels human. the tradeoff: a 4B model is weaker than larger models, but fits in 8GB VRAM and keeps latency down. for voice queries, that accuracy loss is negligible. I can bump to 7B or 8B for quality mode when latency isn't critical. context preservation — the LLM reasons internally "user wants me to find X in their data" , then codes, then explains. no information loss at model boundaries. VRAM efficiency — load the 4B model once ~2-3GB in INT8 . keep it there. reuse it for every query. upgrade to 7B/8B only when you want quality over speed. simple output format — use XML tags to split the LLM response: