This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend My girlfriend is an analogue collage artist. She cuts up old and new magazines and comes up with all sorts of weird and wonderful pieces that are full of colour and humour. She exhibits them in cafés and libraries around Madrid.
The problem with getting her work onto a wall is another job entirely:
She does all of this by hand, from her private Gmail and an Excel sheet.
To make it easier for her to manage the whole process, I built an Art Exposition Manager, an app that takes her from "I'd like to exhibit my work at that library" to a ready-to-send email:
prospect to pitched once she sends it.
She can also do all of this from Telegram: "add the X library in the Y neighbourhood of Madrid", "what's the status with Café XYZ?".
The AI drafts, and she modifies as needed. Nothing is ever sent to a venue automatically.
After the MVP was built, we went for a walk in Madrid city centre, to find a real venue she'd consider. I handed my mobile to her (she didn't have Telegram installed), and when we found an interesting coffee shop we passed by, she sent a message to the bot. It replied asking about confirmation and afterwards stored the venue.
She was quite amazed by the fact that the evening before we were just chatting about the idea of building it, and the next afternoon it was ready to be tested, and we could chat with it through Telegram.
What turned out not to work quite right was the bot (with Hermes Agent - the AI chat framework behind it) trying to write the draft itself, instead of invoking the API that would then do it as part of the application logic. But that was quite easy to address later via prompt and MCP updates.
The app keeps basic track of the artworks:
It keeps track of venues where expositions can potentially take place:
Adding the venue even with as little as a name and a city will store the details and look up the rest via web search. Over time you can keep the details about the venue up to date:
You click a button to generate a "pitch":
And then in a minute or so you get a draft of an email, and a set of suggested works to be exposed, with an explanation of why a certain piece should end up in this exposition:
gitlab.com/mjarosie/exposition-manager Her originals, the database and her emails are git-ignored. The repo only contains synthetic test fixtures.
After having an initial chat with the target user (the artist!), I had a brief chat with Claude, who helped me brainstorm the ideas on what features should be included in the initial MVP (I didn't have much time to work on it, so I had to cut down scope significantly).
Then I've built the service using Claude Code.
Everything runs on my local machine. There are no paid APIs anywhere in the loop.
The model: Gemma 4 26B-A4B, 4-bit, served by llama.cpp's llama-server on my Mac's GPU, using its OpenAI-compatible API. It's a mixture-of-experts model with about 4B parameters active per token, so it's fast enough for interactive use. It also accepts images, so one model covers both analysing the artwork and writing emails.
The agent framework: Mastra. I'd never used it before, so that was an interesting learning experience. Each AI task is a Mastra agent with a Zod schema for its output:
artwork-describer: looks at each collage once and writes a description, motifs, palette, mood, likely source materials, and an explanation of the title's wordplayvenue-profiler: turns a venue's website, photos and notes into a structured profilework-picker: chooses works for a venue and explains why pitch-writer: writes email drafts
These are chained into Mastra workflows. The non-AI steps live in the workflows too: fetching pages, OpenStreetMap lookups, and a code-only check on every draft that rejects placeholders like [Name], wrong languages, drafts that don't mention the venue, and drafts outside 120–350 words. If a draft fails the check, it's regenerated once with the reasons attached.
The vision pass was the part I was most curious about. Here's what Gemma wrote about one of her pieces, Viaja por un ojo de la cara:
The title is a play on the Spanish idiom 'costar un ojo de la cara' (to cost an arm and a leg). The collage literalizes the phrase by showing a creature physically interacting with the woman's eye.
Those descriptions are what make the pitches specific. The picker doesn't choose "work #12", it chooses "the muted blue one that suits a reading room".
Venue research. Web search goes through a self-hosted SearXNG container, and place lookup through OpenStreetMap's Nominatim. Only the venue's name, neighbourhood and city ever leave the machine. The local model decides which search results are the right venue, and it may only return links that actually appeared in those results. Research only fills empty fields and records a source for each one.
A chat front end: Hermes Agent (Nous Research), running in Docker with a Telegram gateway. The app exposes its features as an MCP server, and Hermes talks to the same local Gemma.
Tracing: Arize Phoenix, through Mastra's @mastra/arize OpenTelemetry exporter. Every agent call, workflow step and LLM request shows up as a span, tagged with prompt_version, whether thinking was on, and the venue/artwork/exposition IDs. With a local model you're tuning prompts constantly, and being able to see exactly what Gemma received (including the image) is what makes that workable.
One of the first things I noticed was that the model was not always making use of the tools, especially when researching venues. I decided to make use of Mastra's createWorkflow to guide the agent through a more predefined set of steps for achieving it.
Stack: TypeScript end to end. Hono + Vite/React, SQLite via Drizzle, sharp to shrink her 5000–7000 px scans to 1024 px before the model sees them. Everything except the model runs in Docker Compose (app, Phoenix, SearXNG, Hermes). I decided to run the model on the host, because I couldn't make Docker on macOS make use of the GPU.
How I worked with my coding agent. I planned milestones with Claude, then handed it over to Claude Code. Every acceptance criterion had to have a test before a milestone counted as done. Because LLM output changes from run to run, the tests come in layers:
@live suite runs against the real Gemma and checks This set of tests was what let the agent iterate on its own and know when a feature actually worked.
chat_template_kwargs.enable_thinking on each request, so every agent sets its own value. With --reasoning-format deepseek, the reasoning comes back in a separate field and never breaks JSON parsing.clarify, memory and vision fixed that.null, plus a live test that checks it.
First, more of a general thought: before the LLM boom you used to be able to build anything with just a laptop and internet access, but that slowly stops being the case. Developers get more and more accustomed to the speed at which the AI writes code, and sometimes take it to extremes, where planning and architecture are completely left to the model itself (which I personally don't find to be a good long-term practice). As a result, unless you can afford running a GPU rack in your basement, you're more and more dependent on services of Anthropic, OpenAI, SpaceXAI's Cursor etc. Being able to spin up a capable model on a laptop, with no API keys, no usage meter and no data leaving the machine, gives that independence back, but I think we still have to reach the point where it's nice and easy to use them to work on complex programming tasks.
Open innovation lets others learn too - I would not be able to learn how to serve an LLM if it wasn't for the fact that there are so many tools available for easily down and spinning one up, which I really appreciate - we're truly standing on the shoulders of giants.
Some other things open made possible, specific to this little weekend project:
To be honest about where it was weaker: a frontier closed model would write a slightly more polished first draft, and faster. But every draft would be edited anyway, and good example emails matter more for quality than model size does. So the trade was easy.
The planning conversation, from brain dump to MVP plan: PLANNING_SESSION.md <!-- TODO: check the branch name in this link -->
@mastra/arize.