# Mana: 2-3 Seconds to Feeling Human

> Source: <https://dev.to/yuuzulight/mana-2-3-seconds-to-feeling-human-3e0b>
> Published: 2026-08-04 23:42:09+00:00

so I shipped a voice AI assistant that runs entirely on my machine. no cloud, no APIs, no latency nightmares.

the original idea came from Alice in Sword Art Online — an AI that feels like an actual person, not a chatbot. mixed with JARVIS's anticipation and Neuro-sama's quirky personality. here's what actually went into getting from "wouldn't it be cool" to "this runs 24/7 without issues."

most voice assistants are cloud-first: you speak → sent to server → processed → response → back to you. each hop adds latency. you're looking at 3-6 seconds before you hear anything. for a voice interaction, that's dead. it kills the feeling of talking to something intelligent.

I wanted something faster. something that *responds*.

the constraint: do it locally. use an 8GB VRAM GPU, run everything on-device, no external APIs except for the live2d avatar bits (because that's hard to render locally and still look good).

here's the reality: I have a GPU with 8GB VRAM. no budget to experiment with better cards or more models. so every architecture decision was forced by what actually fits.

naive approach: chain multiple specialized models.

```
User speaks
  → Transcription model (Whisper)
  → Planning model (3B: what should I do?)
  → Coding model (7B: generate implementation)
  → Verification model (4B: is this correct?)
  → TTS (speak the answer)
```

math: 1s + 2s + 3s + 1.5s = 7.5s of latency before the user hears anything. nope.

the problem isn't just that each model is slow. it's *model loading overhead*. every time you swap from one model to another, you:

with only 8GB, this gets gnarly fast.

the constraint was hardware. 8GB VRAM. no more, no less. that forced clarity: pick one model that does everything, or pick nothing.

so I went with a single model (4B by default, with 7B/8B quality modes available) that does reasoning + code generation + explanation in *one pass*.

latency: ~2-3 seconds total. actually conversational.

this is the difference between a chatbot and a companion. JARVIS doesn't pause for 6 seconds before responding. neither does Mana. turns out, when you're forced to optimize for latency (because you only have 8GB to work with), you accidentally build something that feels human.

the tradeoff: a 4B model is weaker than larger models, but fits in 8GB VRAM and keeps latency down. for voice queries, that accuracy loss is negligible. I can bump to 7B or 8B for quality mode when latency isn't critical.

**context preservation** — the LLM reasons internally ("user wants me to find X in their data"), then codes, then explains. no information loss at model boundaries.

**VRAM efficiency** — load the 4B model once (~2-3GB in INT8). keep it there. reuse it for every query. upgrade to 7B/8B only when you want quality over speed.

**simple output format** — use XML tags to split the LLM response:

```
   <reasoning>what I understood</reasoning>
   <code>implementation</code>
   <explanation>what to say via TTS</explanation>
```

then execute code silently, speak only the explanation.

```
[Voice input]
  ↓ Whisper (local transcription)
[Text]
  ↓ Qwen 4B (reasoning + code + explanation)
[Structured output]
  ├─ Code (execute silently, log results)
  └─ Explanation (TTS via Kokoro/Chatterbox/Fish Speech)
  ↓
[Audio output to user]
```

total latency from "hey Mana" to hearing the response: ~2-3 seconds. feels like talking to something intelligent.

with an 8GB GPU, the budget is tight:

total: actually fits (barely). I quantize aggressively, drop smaller models, and flush unused ones. the tradeoff is worth it — every millisecond of startup or response latency costs the feeling of talking to something alive.

it's been running 24/7 for about 3 months now. no crashes. no "processing never finished" hangs. it just works.

**earlier profiling** — I spent weeks optimizing things that didn't matter (TTS buffer sizes), then found the real bottleneck (model loading) in one afternoon with a profiler.

**accept quantization earlier** — INT8 quantization costs ~5-10% accuracy on reasoning. for voice queries, that's negligible. I fought it for weeks.

**screen context is expensive** — OCR-ing the screen every query tanks latency. I ended up making it opt-in ("hey Mana, look at my screen").

currently working on:

the long-term goal (as hardware allows):

that's it. local-first voice AI that ships on consumer hardware. the constraint (8GB VRAM) forced good decisions: use one model, quantize aggressively, separate concerns (code vs. explanation), measure everything.

the rest is just execution.

at the core, the goal was Alice. an AI that feels like talking to a real person. turns out, the technical constraints *create* that feeling. instant response. screen awareness. a quirky personality through code decisions. JARVIS's anticipation through context preservation. Neuro-sama's realness through being unfiltered and responsive.

you can't fake that with prompts. you have to build it into the architecture.

want to try it? it's open source: [github.com/Yuuzulight/Mana](https://github.com/Yuuzulight/Mana)
