This is built for a friend - a small business owner who runs a bubble-tea shop. Every month she exports her bank statement to reconcile her books, and every month it's the same wall of numbers: hundreds of rows of transactions, no structure she can act on. She asks the same question in three different ways - "Is the shop actually making money? Where does it all go? What should I do differently?"
Her bank statement is also exactly the kind of file she should never upload to some AI service. It's her income, her rent, her suppliers, her staff wages. A closed API would mean shipping all of that to someone else's server.
So I built Finance Friend for exactly this person: upload an Excel/CSV export, and a small open-weight model running on her own machine writes a plain-language report - where the money comes from, where it goes, and what to watch. No cloud. No API keys. No data leaving the laptop.
bank statement (xlsx / xls / csv)
|
v
analyzer.py -- pandas, on this machine
| auto-detects date / description / income / expense columns
| monthly income & expenses, net flow
| top income sources, top expense categories, biggest transactions
| compact JSON summary
v
reporter.py -- llama.cpp, on this machine
| Qwen2.5-3B-Instruct (open weights, GGUF)
| 250-350 word plain-language business report
v
Gradio UI - overview table, monthly table, AI report
The parser is forgiving on purpose - Chinese and English bank exports name their columns differently, so it matches common names loosely (Chinese exports use columns like income/expense in Chinese, which the same matcher handles):
IN_KEYS = ["income", "deposit", "credit"]
OUT_KEYS = ["expense", "withdrawal", "debit"]
def _pick_columns(df, keys):
cols = {_norm(c): c for c in df.columns}
for k in keys:
if _norm(k) in cols:
return cols[_norm(k)]
...
The report generator is just a chat completion against a local llama.cpp model:
from llama_cpp import Llama
llm = Llama(model_path="qwen2.5-3b-q4.gguf", n_ctx=8192, n_threads=8)
report = llm.create_chat_completion(
messages=[
{"role": "system", "content": SYSTEM_PROMPT},
{"role": "user", "content": prompt}, # the compact JSON summary
],
temperature=0.3,
)["choices"][0]["message"]["content"]
That's the whole AI surface. One model file, one local process, zero network calls.
Feed it three months of a bubble-tea shop's transactions and it produces a report like this (I asked it for English here; for her, the same model writes in Chinese - the language is part of the prompt):
In the period from July 1 to September 28, 2026, this small bubble-tea shop is profitable overall, with a net surplus of 6,536.07 yuan. Revenue comes mainly from in-store WeChat payments and Alipay takeout orders, which together make up most of income, and both are trending upward month over month. Spending is dominated by dairy suppliers and staff wages. Rent and utilities also stand out as large monthly costs. Overall, income exceeds spending and revenue is growing, but rent and utilities deserve attention. Recommendation: negotiate with suppliers to lower purchase costs, and review whether the current rent and utility terms can be improved.
Which is exactly the answer to her three questions: yes, it's profitable; the money goes to dairy and staff; negotiate with suppliers and review rent.
This is the "Build for a Friend" challenge, and the friend is real - a small business owner, exactly the person I have in mind at every design decision. The demo data is synthetic on purpose: even a demo shouldn't run on anyone's real financial data, which is the product philosophy in a nutshell.
Next step: hand the app to her, tune the report to how she actually reads it, and let her real statements run only on her own machine - never anywhere else.
Repo: github.com/fengyuGbt/finance-friend
python3 -m venv .venv && source .venv/bin/activate
pip install llama-cpp-python pandas openpyxl gradio
curl -L -o qwen2.5-3b-q4.gguf \
https://hf-mirror.com/Qwen/Qwen2.5-3B-Instruct-GGUF/resolve/main/qwen2.5-3b-instruct-q4_k_m.gguf
python app.py # -> http://127.0.0.1:7860
Want a demo file without using your own data? python sample_data.py generates a synthetic three-month statement.
Built with open-source AI at its core: an Apache-2.0 open-weight model, running locally through llama.cpp - because for this friend, privacy isn't a feature, it's the whole product.