# Munshi: A Local AI Clerk for Businesses Receiving Orders by Phone

> Source: <https://dev.to/coder_nesi_c5c15cf349b1e4/munshi-a-local-ai-clerk-for-businesses-receiving-orders-by-phone-p29>
> Published: 2026-10-05 05:36:31+00:00

*This is a submission for the [Hacktoberfest Weekend Challenge: Build for a Friend](https://dev.to/challenges/hacktoberfest-weekend-2026-10-01)*

Every night around 10 pm, my friend Iqbal's father sits down with a notebook and does his whole day a second time.

I've always called him Uncle. He runs A2Z general store, a kirana (grocery) store in Ranchi, and most of his orders don't arrive politely. They come as a phone call while his hands are busy weighing something for someone else. As a WhatsApp voice note from a customer who recorded it while walking. As a text where half the items are in Hindi, half in English, and one is a brand nickname only the two of them understand. *"Do bori cheeni, paanch kilo chai patti, aur wo wala biscuit jo pichli baar bheja tha."* (two bags of sugar, 5 kg of tea leaves and the biscuit delivered earlier) 

He can't stop and type that properly at 2 pm, so he scribbles. At night he replays the voice notes, squints at his own handwriting, and turns it all into a clean list for the packer. There are times when mistakes happen. For instance, a customer asked for a 15-rupee biscuit, but it was misheard as 15 packets, leading to an incorrect order.

That was the part I understood when I started. The part I hadn't understood is that writing the order down is where his job *begins*.

After that, someone has to pack it. Someone has to deliver it. And someone has to notice when either of those isn't happening, and in his shop that someone is Uncle, by phone, between customers. *Packing ho gaya? Nikla kya? Kahan tak pahuncha?* There's no board and no list of what's stuck. If an order is slow, he usually finds out when the customer calls to ask where it is. And that call lands on the same phone that takes new orders, so every "mera order kahan hai?" (where is my order) is one more interruption for the person who is the clerk, the manager and the cashier at once.

Then there's the khata, the credit ledger. Whether to accept a big order on credit isn't something he looks up. It's a running total in his head plus a feeling about the customer. The feeling is mostly right. once, he had to stop and double-check because the amount was getting uncomfortably close to what he was willing to let that customer owe.

When I asked what he'd want gone first, he said the repeated checking and rewriting: if Munshi could take the order properly, keep track of what was happening to it, and only bring him the things that actually needed his attention, that would save him the most time.

So I built **Munshi** (an AI Agent which does the exact tedious job), named after the old clerks who kept the books in shops like his. It takes over the routine part of both jobs, the 10 pm copying and the all-day chasing, and it pulls him in only when something actually needs him.

**A normal order handles itself.** A regular customer sends a voice note. Munshi transcribes it, works out the items (even when it's "cheeni" one time and "chini" the next), checks stock and the customer's credit, writes the bill, and sends it back in the chat. A packing slip goes to the least busy packer. Uncle never touches it.

**A risky order stops and asks him.** If an order would push a customer past their credit limit, Munshi doesn't decide. It pauses and shows him a card: what was ordered, what the customer already owes, the limit, and exactly why it stopped, with Approve and Decline buttons. The customer gets a polite "owner se confirm karke batata hu".

**A wrong count stops the truck.** If the packer's counts don't match the order, dispatch locks and both the packer and Uncle are told. He can accept the partial order, ask for a recount, or cancel.

**It watches the order, not just the message.** If packing runs past the time he's set, Munshi nudges the packer. If it keeps running, Uncle is told, and the customer hears from Munshi first, with a new approximate time and a link. The link opens a small page showing exactly where the order is, so the customer doesn't need to call.

**Each morning owner gets a short summary** in Hinglish: how many orders Munshi handled on its own, what needs him, who is close to their credit limit, and what will run out soon.

The rule I kept coming back to: **the AI understands, the rulebook decides, and humans handle exceptions.** The language model reads the order. It never decides anything about money. Credit limits, approvals, stock and dispatch locks live in a plain config file Uncle can read and change. One of the rules is his own: credit orders that cross his configured limit need his approval.

I didn't try to turn it into a polished testimonial. His reaction was basically that the useful part was not "AI"; it was that he could see the order, see what Munshi was waiting for, and only step in when something actually needed him. His next question was the practical one: whether this could eventually sit on the same phone workflow he already uses for WhatsApp.

WhatsApp is simulated by a chat page in this MVP, because the business API needs account approval I couldn't get in a weekend. The voice notes I tested with are a mix of real shop-style recordings and controlled recordings made for testing. The 30-day order history in the demo i seed data, and it's labelled as such in the code

**Munshi is a local AI clerk that turns a small wholesale shop’s customer messages into checked orders, bills, packing work, delivery updates, and an auditable activity log.**

It is designed around a practical boundary: AI can help interpret language and investigate an order, but deterministic shop rules and the order engine control money, stock, approval, billing, and dispatch. If something is unclear or risky, the system asks the customer or the owner instead of silently guessing.

Small wholesale and kirana shops coordinate repeat orders across informal channels: typed messages, voice notes, familiar product names, quantities, credit, available stock, packing, and delivery. A conventional order form expects structured input up front; a free-form chatbot can be difficult to audit and unsafe to trust with business decisions.

Munshi combines a familiar chat-like customer page with an explicit operational workflow. It aims to:

The README has setup (one command to run the demo). The rulebook is `rules.yaml`, and `PROMPTS.md` holds every prompt I gave my coding agent.

Everything runs on one MacBook Air M5 with 16GB of memory. That was a useful constraint, because it forced every choice to stay small.

**Hearing.** Voice notes are converted and transcribed locally with mlx-whisper (large-v3-turbo). I also tested Gemma 4's own audio input through Ollama as an alternative and compared both on a small set of real voice notes: mlx-whisper was the production choice because it fit the Apple Silicon setup cleanly and kept the ASR path local and lightweight.

**Understanding.** Gemma 4 E4B, running in Ollama, pulls the items, quantities and units out of the transcript as structured JSON. Plain code then matches them against the catalog with fuzzy matching over Hindi, Roman and Devanagari aliases. Nothing is guessed. If it isn't sure whether "cheeni" means 5 kg or 50 kg, Munshi asks the customer.

**Deciding.** A state machine with explicit transitions and a rulebook in YAML. Credit, stock, approvals and the dispatch lock are ordinary Python, covered by 49 automated tests in the current Phase 3/4B baseline. The model has no say here.

**The agent.** Gemma also runs a small tool-calling loop: look up customer history, search the catalog, check stock, then propose a next step. The engine validates everything. A proposal that is less cautious than the rulebook gets overridden and logged, and if tool calling fails, a deterministic path takes over, so an order never gets stuck. Agent handled the clean test cases without fallback; real fallback rate is not being claimed until a larger real-order evaluation is run.

**The supervisor.** This part deliberately uses no model. It compares time-in-stage with the limits in the rulebook, nudges, escalates, and sends the customer message from templates. Late orders are exactly where I didn't want anything unpredictable.

The stack is FastAPI, SQLite, vanilla JS and rapidfuzz. I built it with OpenCode as my coding agent, and Gemma is the agent inside the product. They're separate things.

| Metric | Result | 
|---|---|
| Automated tests | **60/60 passing** | 
| Synthetic voice tests | **6 clips** | 
| Exact end-to-end accuracy | **66.7% (4/6)** | 
| Incorrect orders committed | **0/6** | 
| Catalog match threshold | **≥88 score + 8-point margin** | 
| Agent limit | **6 tool-calling turns/order** | 
| Audio chunking | **28s** | 
| Concurrent Ollama models | **1** | 

| Model | Avg. order latency | Memory | Decision | 
|---|---|---|---|
| Gemma 4 E2B | **~2.4s** | **~5.1 GB** | Faster, weaker on ambiguity | 
| **Gemma 4 E4B** | **~4.8s** | **~8.3 GB** | **Selected** | 
| Gemma 4 12B | **~11.7s** | **~13.5 GB** | Too heavy for 16GB | 
| Whisper Large-v3-Turbo (MLX) | **~1.6–2.3s / 10s audio** | **~2 GB** | **Selected STT** | 

The key result wasn't raw accuracy: **0 incorrect orders were committed** in the synthetic test. Uncertain extraction goes to clarification, while credit, stock and billing remain controlled by deterministic Python rules.

**LLM proposes → Python validates → rules decide → engine commits.**

The important result for me was not a single accuracy number. It was that when the model was uncertain, the system could detect that uncertainty and stop before turning it into a wrong bill.

A few failures were useful because they exposed exactly where the safety boundary needed to be. Noisy audio could produce a transcript with a product name slightly wrong, so catalog matching and clarification had to catch the uncertainty before billing. A brand nickname that wasn't in the alias list stayed unresolved instead of being silently mapped to the nearest product. And one mixed Hinglish phrase was transcribed correctly enough to read but not confidently enough to commit, so Munshi stopped and asked instead of turning an uncertain interpretation into a bill.

A kirana (grocery) store's khata (ledger) is one of the most sensitive things it owns. It records who owes what, but the privacy issue goes beyond the ledger. Every order also contains customer information, their name, phone number, address, order history, and sometimes voice recordings. Sending all of that to a cloud LLM just to understand a voice note didn't feel right to me. With open-weight models running locally, the customer's personal details, voice notes, orders and balances stay on the shop's laptop instead of being sent to a server the shopkeeper doesn't control.

Then there's money. Uncle runs on thin margins, and a per-call fee that looks tiny to a developer looks like a recurring bill to him. Once the models are downloaded, Munshi costs nothing per order, and it keeps working when his internet doesn't.

The one I didn't expect to matter so much is control. Because the models are open, I could treat them like parts. I ran the same voice notes through Whisper and through Gemma's own audio input and compared them. I tried three Gemma sizes and picked the one my 16GB laptop could run comfortably alongside everything else. When something misheard an item, I fixed my prompt and alias list instead of filing a support ticket. And because the rulebook sits outside the model, Uncle can change how Munshi behaves by editing one sentence.

It wasn't free of cost. I'm not claiming a small local model beats the best cloud model at messy Hinglish. The honest limit is that the system still depends on careful aliases, conservative matching and clarification when speech is ambiguous, especially for brand nicknames and noisy recordings. For this shop, privacy, cost and control were worth that trade, and a system that asks when it's unsure and escalates when it's risky matters more than raw accuracy.

I built Munshi with OpenCode, and I used Entire alongside it to capture and keep track of the development sessions.

Munshi was built in phases rather than as one giant prompt. I would give OpenCode the project instructions, work through one phase, test it, fix what broke, and then move to the next phase. Entire gave me a persistent record of those coding sessions and the context behind the changes, so I could go back and understand not just what changed in the code, but how the implementation evolved.

This was especially useful for a project like Munshi because the important decisions were not only about writing code. I had to keep track of why the system uses a deterministic rulebook, why the AI is not allowed to make credit decisions, why STT and the LLM run sequentially on the 16GB Mac, and why failures should fall back instead of leaving an order stuck.

The combination worked well for me: OpenCode was the coding agent, Entire captured the development context, and Gemma was the agent running inside Munshi. They each had a different job.

The earlier implementation prompts and project instructions are also kept in "PROMPTS.md", so the development process is reproducible rather than just a final code dump.
