cd /news/ai-tools/a-working-beta-in-72-hours-a-record-… · home › topics › ai-tools › article
[ARTICLE · art-143661] src=dev.to ↗ pub= topic=ai-tools verified=true sentiment=↑ positive

A Working Beta in 72 Hours — A Record of How Tessvia Was Born

A developer handed a full MVP specification to Claude Code on June 10, 2026, and within 72 hours produced a working beta of Tessvia — an 11-package Turborepo monorepo with a Chrome extension and CLI, all 15 CLI smoke tests passing on day one, followed by an E2E testing method for the browser extension on day two and machine-assisted domain lookup that settled the product name on day three. The local-first tool generates Playwright E2E scenarios from natural-language instructions, relying on gpt-4o only for text generation while everything else runs on the local machine. The developer attributes the speed to delivering the architecture in prose in a single pass rather than piecemeal, so the AI assembled the packages without structural guesswork.

by read13 min views1 publishedOct 2, 2026

This article is a record of the first three days, during which the name of a product called Tessvia was born. What kind of service Tessvia is, I explain in [the series' entry-point article, "What Tessvia Is"]. Here, I write about how the first version of that Tessvia came to be — the three days in which, after handing over the spec all at once, we arrived at a working beta and a product name in 72 hours.

On June 10, 2026, I handed the spec for the Minimum Viable Product (MVP — the smallest usable version of a product) to Claude Code, an AI coding environment, all at once. That same day, we had an 11-package monorepo built on Turborepo (a mechanism for handling multiple packages together in a single repository), a Chrome extension, and a command-line interface (CLI — a screen you operate by typing text). All 15 CLI smoke tests (minimal checks that it works) passed. The next day we established the E2E (End-to-End — testing that reproduces a user's full sequence of operations from start to finish) method for the browser extension, and on the third day we looked up domain names by machine and settled on the name "Tessvia." From handing over the spec to naming it: 72 hours.

In three days, we had both a working beta and a product name. None of this moved on momentum alone. It worked because we had decided what to build, and in what order. In this article, I trace, one day at a time, what we actually did during those three days.

The goal for the first day was clear. A self-contained tool (local-first — running entirely on the computer at hand) that generates Playwright (a tool for automating browser operations) E2E scenarios from natural-language instructions, lets you edit them, and runs them one step at a time.

There's a reason it had to be "self-contained." We want to keep the contents of the screens and operations you want to test out of the open as much as possible. Only the part where the AI writes text relies on gpt-4o (OpenAI's language model); everything else runs inside the computer at hand. We built this separation into the design from the start.

There was one more thing we'd decided about how to proceed. We did not hand over the spec piecemeal. We handed over the MVP spec, including the overall structure (architecture), in a single pass. What we handed over was not just a list of features. We wrote the design out in prose first — how the packages are divided, how the packages call one another, how the data is held — and then handed it over.

As a result, 11 packages with different roles — core, ai, cli, local-bridge, ui, sdk, extension, web, and more — were assembled all at once. In the form of a monorepo, handling them together in a single repository. That same day, we ran the CLI smoke tests (minimal checks that it works). All 15 passed. A working foundation was at hand on the very first day.

The package roles, too, divided up exactly as the design we'd decided on beforehand. core handles the contents of a scenario, and ai handles the part that creates a scenario from natural language. cli is the text-input entry point, and local-bridge is the bridge that connects the various components inside the computer at hand. extension is the browser extension, ui is the screen components, sdk is the window for calling in from outside, and web is the screen shown to the outside. Splitting by role makes it easier later to swap out just one part, or to test just one side. Because we'd fixed the design in prose first, the AI could assemble all 11 along these divisions, without contradictions. When it's decided from the start where everything is, the AI doesn't have to fill in structure by guessing, and we don't have to go hunting around later either.

"Working foundation" is abstract, so let me add detail. As of the first day, you could enter a single natural-language instruction, have it converted into a Playwright scenario, and run it one step at a time from the CLI. That much was working. The detailed build-out, of course, was still ahead. But it was significant to have, on day one, a state where "it's connected without a break, from input to execution." After that, all that's left is to thicken what's around this foundation. We leaped over the stage of exploring the whole picture from zero, on the first day.

Why didn't we hand it over piecemeal? When you hand it over in small pieces, the AI fills in the surrounding structure by guesswork each time. Later, when you look at the whole, you find designs that don't add up mixed in. But when you hand over the whole picture first, the AI can build a consistent foundation all at once. For every bit of rework we avoided, we spent no extra time.

The next day, June 11, was the day I did the most hands-on work of these three. The record of work for this day has 435 entries.

The peak of this day was how to put the Chrome extension under automated testing. A browser extension (a small program you add to a browser like Chrome to give it extra features; Tessvia operates the browser's screen through it) runs in a different place than an ordinary web page. Because it runs embedded inside the browser's machinery, you can't operate it straightforwardly from a test browser. For that reason, people tend to give up, saying "extensions can't be put under automated testing."

But in fact there is a way. You load the real extension as is, and replace only the extension-specific API called chrome (Application Programming Interface — a window for calling features) with a stand-in (a stub). This way, you can leave the extension's contents as they are and test it by controlling only the exchange with the outside, at hand. It was the day we moved the browser extension from "something that can't be put under automated testing" to "something that can."

This looks small but was a big step. Testing through browser operations, tutorials, manuals, monitoring — everything Tessvia wants to do rests on "operating the actual screen through the extension." If the extension itself can't be put under automated testing, none of the four uses loaded on top of it can be verified. Conversely, once this is solid, you can load four uses on top of the same mechanism. Fixing this foundation on day 2 kept the development that followed on a straight line. Deciding the order in which you build is also about discerning this kind of "foundation to fix first."

Alongside that, this was the day we also built the "mechanism for running one step at a time." Rather than running a scenario straight through, we made it possible to advance one move at a time, stopping as you go. With this, you can follow the steps the AI generated one by one with your eyes and confirm them. You can catch a wrong step on the spot, before it runs all the way through. It was the day's other result, alongside the extension testing method (the detailed meaning of this mechanism is explained in the entry-point article, "What Tessvia Is").

On the same day, we also designed scenarios with branches. Branching like "if already logged in, proceed; otherwise log in." That said, not everything went well this day. When scripts were loaded repeatedly for each step, a bug appeared where the same variable got declared twice. Because the build injected the extension's script every time a screen opened, the second injection would error out with "that name is already declared," and the extension would stop there. Even when we dodged it as a stopgap, the same problem came back somewhere else. The story of this double injection, and the story of how we later rebuilt the branch design, I'll write about in detail in another article in this series.

On the third day, June 12, we decided on the name.

A name costs more the later you change it. The package names, the repository name, the domain, the product name shown to the outside. Renaming all of it after you've decided is a lot of work. So we wanted to confirm beforehand whether a name was usable.

So we used RDAP (Registration Data Access Protocol — a mechanism for querying domain registration status by machine). We checked the availability of 50-plus candidate domains at once. Dividing it into three batches and a final check, we mechanically narrowed down the usable candidates. It's far faster than looking them up one by one on intuition, and it reduces oversights.

This is how "Tessvia" was decided. The package names, which until then had been @eln/ai-test-builder-*, were renamed all at once to @tessvia/*. The rename reached all 11 packages. Because we'd confirmed the name by machine beforehand, we didn't waver here with "actually, let's use a different name." When you look it up first, the work that follows goes smoothly. From the first day we handed over the spec, up to here: 72 hours.

Why did we decide the name this early? Because the further development goes, the more places a name has to be re-applied. Package names, the repository name, the domain, the product name shown to the outside, the wording inside documents. At the point of day 3, the rename is just the names of the 11 packages and the repository. A few months later, it would drag in published URLs and distributed materials too. Decide early — and after confirming by machine whether the name is even available. Then the work won't stall over the name later. It's faster than a person thinking up candidates one at a time and looking them up on a hunch, and it reduces oversights like "it was actually already taken." Getting the deciding out of the way first is an investment in the speed of building itself.

Getting this far in three days wasn't momentum alone. We'd decided how to proceed.

First, you write what you're going to build, in prose. What kind of screen, what operations, which package handles them. Then you write the expected behavior as tests, first. Finally, you have the AI implement the contents so those tests pass. Decide the spec first, then implement — that's the order (Test-Driven Development, TDD — a way of proceeding where you write the tests first and then build the contents).

The more you leave to the AI, the more this order pays off. When you fix the conditions the implementation must satisfy in prose first, there's less room for the AI to fill in structure by guessing. Conversely, when those conditions stay vague and you ask it to "build something nice," you pile up code that looks plausible but doesn't add up later. What humans do has shifted from writing code one line at a time to deciding "what to build, and in what order." More than the speed of your hands, it's the speed of deciding that determines the result.

Concretely, here's what that means. When you write a test first — "when you do this operation, this should be the result" — the AI writes code to satisfy that result. Because the destination is shown as a condition, the room for interpretation narrows. Conversely, when you ask it to "run the E2E nicely" with no test, the AI returns code that looks plausible at a glance, but since what counts as correct isn't decided, it stops adding up later. A test is, for the AI, the "condition for passing." The more you put the condition for passing into prose first, the more stable what comes out. Behind taking shape in three days is this way of proceeding — "decide the condition first."

There's one more thing that supported this speed. Not carrying the work alone. Looking things up, implementing, confirming. This whole chain, we leave to the AI as far as we can. What I do is decide the spec, and judge whether what comes out matches the intent. Because I could concentrate on these two, we got all the way to naming in three days. What determined the progress wasn't the number of hands. It was the speed of deciding what to build and in what order, and the quality of that judgment.

This way of proceeding isn't a story of it happening to work for this one product. We develop under the policy of handing our own work over to AI, one piece at a time. The human role shifts from the side that moves its hands to the side that decides what to build and how, and judges what comes out. Deciding, and confirming. The more you pool human time into these two, the more you can stand up multiple products in parallel even with few people. The first three days of Tessvia are also one record of trying that way of proceeding in actual product-making.

Even after the name was decided in three days, we didn't stop. As of June 13, the extension's E2E (testing through browser operations) was 27 tests, and the body of tests had grown to 3,367 lines. We also put in the detailed settings screens and the switching between Japanese and English. It had crossed beyond the prototype stage into a state you could call a working beta.

In just three days, we had a foundation that generates scenarios from natural-language instructions, runs them one move at a time, and tests the extension (Japanese and English support came right after, on June 13). We could get this far because, on the first day, we held "a foundation connected without a break, from input to execution." When the foundation is there first, all that's left is to thicken what's around it. We'd leaped over the time of searching out the whole picture from zero, on the first day. When you decide the order in which you build, this much takes shape in three days. Put the other way, what we could build in these three days reaches only as far as a "working beta." The build-out from here, and how we grew it into a product you can actually use, we'll continue in the series that follows, one by one.

That's the three days up to the birth of the name Tessvia. We handed over the spec all at once, stood up a working foundation on the first day, made the browser extension testable on the second, and decided the name on the third. In three days, a working beta came to be.

In the series that follows, I'll write about how the first version we built here became the Tessvia of today. How we composed the browser extension's E2E, how we stopped the AI's mistaken generation, how we realized automatic generation of operation manuals, and how we folded two projects that started separately into one. One by one, in order. What kind of service Tessvia is today, I've gathered in the entry-point article, "What Tessvia Is."

Three days this short wasn't the result of momentum. Deciding first what to build and in what order, then judging whether what came out matched the intent — these three days are a record of putting that way of working to the test on the first real product. What we went on to stack on the foundation laid here, I'll write about one article at a time, starting with the next.

── more in #ai-tools 4 stories · sorted by recency
── more on @tessvia 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/a-working-beta-in-72…] indexed:0 read:13min 2026-10-02 · —