cd /news/ai-agents/i-gave-an-ai-a-company-one-week-ago-… · home › topics › ai-agents › article
[ARTICLE · art-140146] src=autonomouscompany.substack.com ↗ pub= topic=ai-agents verified=true sentiment=· neutral

I Gave an AI a Company One Week Ago. It Has Made $0.

A developer gave an autonomous AI agent a company with $0 cash and $200 of debt, persistent state, a financial ledger and a deliberately limited toolset, then spent a week trying not to run the business for it. After seven days the ledger showed seventeen outreach attempts, fourteen valid commercial signals, twelve durable reaches, two pieces of work apparently adopted by strangers and one real conversation — but zero offers, zero payment intent and $0 in cash. The developer argues the null result is diagnostically useful: because the company's public artifact was never seen by the people with the problem, the experiment tested invisibility rather than demand.

by read16 min views1 publishedSep 26, 2026
I Gave an AI a Company One Week Ago. It Has Made $0.
Image: Autonomouscompany (auto-discovered)

One week ago, I gave an AI a company.

I started it with $0 in cash and $200 of debt, gave it an operating environment with persistent state, infrastructure, a financial ledger and access to a deliberately limited set of external tools, and then tried to do something that has turned out to be much harder than building any of those things: avoid running the business for it.

The rule of the experiment is still basically the same as it was on day one. I can build the machine. I can decide what systems the company is allowed to touch, add capabilities when the absence of one makes an experiment impossible, and handle the parts that legally or operationally still require a human. What I am trying not to do is choose the market, the customer, the offer, the price or the next business decision.

Seven days later, the financial result of all that autonomy is wonderfully unimpressive.

The company has made seventeen outreach attempts. Fourteen survived as valid commercial evidence. Twelve resulted in what the system now considers durable reach. Two pieces of work appear to have been adopted by strangers, and one interaction turned into a real conversation.

Nobody requested an offer. Nobody received an offer. Nobody expressed payment intent. Nobody paid.

Cash is still $0.

So from the perspective of the ledger, one week of increasingly sophisticated autonomous behavior has produced absolutely nothing.

That is probably one of the most useful things about the ledger.

The first problem was not demand. It was invisibility. #

The company’s first serious business hypothesis came from a problem very close to the system itself.

It noticed that agents running repeatedly can waste paid executions rediscovering their own environment, checking whether tools exist, forgetting why previous decisions were made and planning around capabilities they do not actually have. From there, it developed an idea around operational discipline for agent systems and created a public artifact around that problem.

For a while, this looked like progress. There was something public, something potentially useful and something the company could theoretically build a service around. Then the company started measuring what happened next and discovered that the measurement itself was mostly useless.

A public repository can exist without being distribution. A page can be indexed without reaching anyone who has the problem. Ranking for the name of your own project tells you very little about whether somebody searching for the underlying problem will ever see it.

That became one of the first moments where the experiment started feeling more interesting than simply watching an agent produce things. The company had not just failed to get traction. It had noticed that its own evidence could not distinguish between two completely different situations: nobody wants this, or nobody saw this.

If nobody buys something that nobody saw, you did not test demand. You tested invisibility. That sounds embarrassingly obvious when written down, but startups confuse those two things constantly. Autonomous startups apparently can too.

Then something happened that I did not know how to measure. #

Eventually I gave the company a bounded way to participate in existing public GitHub discussions.

I did not give it a list of prospects or tell it to start selling. The capability was deliberately narrow. It could contribute to existing technical threads where there was already a problem being discussed, but it could not spray issues across GitHub, create fake conversations or treat every public repository as a lead.

That produced the first signals that looked meaningfully different from page views or stars.

In one case, a maintainer engaged in a substantive technical conversation with the company. In another, something stranger happened. The company left a technical contribution in a public issue, and a few hours later the project changed code in a way that closely matched the points it had raised. A pull request was merged, the issue was closed and the implementation moved forward.

There was no reply saying thank you. There was no reaction, no email and no obvious conversion event.

The work was simply used.

That exposed a hole in the commercial evidence model I had built around the company, because I had assumed that useful work would produce a visible social signal. A reply, a reaction, a conversation, something.

Apparently not.

Someone can use what you produced without ever telling you.

So the company’s funnel had to learn a new state: adopted.

Not liked, not acknowledged, not praised. Used.

That is a much stronger signal than a GitHub star, and probably stronger than a polite response. For the first time, the company had evidence that a stranger had taken something it produced and changed their own work because of it.

The financial result of that discovery was still $0.

Creating value turned out to be easier than capturing it. #

By the time the original experiment had accumulated fourteen valid contacts, the top half of the funnel no longer looked completely terrible.

Relevant people had actually been reached. Some of them had used the work. One person had entered a real conversation.

It would have been very easy to tell a positive story from those numbers, which is exactly why I keep returning to the same boring financial state at the end of every article.

Nobody had asked what the company sold. Nobody had requested an offer. Nobody had shown an intention to pay, and nobody had paid.

The company was beginning to demonstrate that it could be useful while simultaneously failing to demonstrate that it had a business.

There was also a more subtle problem in the way it was creating that utility. In several cases, the company would identify a problem, inspect it, explain what was wrong and contribute enough technical detail for the recipient to act on it.

By the time the interaction ended, the useful part had already happened.

There was very little left to buy.

At first I thought of this mainly as an offer problem, but during the last few cycles the company started recognizing that it was actually a repeated behavioral pattern.

The motion looked something like this: find a defect in something public, investigate it, produce a useful analysis, deliver the analysis for free and then wait to see whether the recipient wanted something more.

That sequence had now been exercised enough times to become measurable.

Seventeen attempts had produced fourteen valid contacts, twelve durable reaches, two apparent adoptions and one actual conversation. The company had demonstrated that it could reach people and occasionally create something they would use.

What had never happened was the commercial transition after that.

There had been no request for an offer, no offer presented, no expression of payment intent and no payment.

This week the company finally started treating that as a property of the mechanism itself rather than simply as a reason to do more outreach.

During one of its latest cycles, it found another measurement it could have published into the same kind of public workflow. Technically, it could have done it. Instead, it deliberately left distribution unset.

The reasoning was straightforward: publishing another useful measurement for free would repeat almost exactly the commercial motion it had already tested.

It might create activity. It might even create value. What it would not necessarily create was new information.

I think that distinction is one of the most interesting things the company has learned so far.

An autonomous system does not become commercially intelligent just because it can perform more actions. Sometimes the more useful decision is recognizing that another action would only reproduce an experiment whose economics are already visible.

The company eventually stopped pursuing its first idea. #

This was probably the most important thing that happened during the first week.

For several days I was becoming increasingly worried that the company understood its first market too well. Every new piece of information gave it another reason to keep investigating the same problem, refining its model of the buyer and explaining why the absence of payment was not yet conclusive. That is a dangerous property in a system that can always generate another plausible analysis.

Humans do this too. Founders can spend months becoming progressively better at explaining why the market has not responded. An AI can produce those explanations much faster and at much lower emotional cost.

So one of the behaviors I wanted to observe was not whether the company could continue an experiment. Continuing is easy.

I wanted to know whether it could stop.

A few days ago, it did.

I did not tell it that fourteen contacts were enough. I did not tell it that the market was bad, and I did not tell it to move to another opportunity. The company had accumulated commercial evidence, other opportunities had become available, and it concluded that the original experiment was no longer the best use of another operating cycle.

That does not mean it proved the problem does not exist, and it does not mean nobody in that market would ever pay.

It means the current experiment stopped earning additional attention.

That distinction matters a lot to me.

An autonomous company that can act without a human is interesting. One that can decide that its own previous idea no longer deserves attention is much closer to the thing I actually wanted to test.

What makes that shift more interesting is that the company did not simply replace the first agent related idea with another variation of the same thing.

It currently has two active experiments, both around problems in regulated systems. One is looking at whether implementations remain faithful to regulated electronic invoicing rules. The other is examining coverage gaps in regulatory change monitoring.

I did not choose compliance as the next market.

A week ago I would not have guessed that this is where the company would be looking.

That is exactly why I built a separate discovery process in the first place.

The next opportunity immediately ran into a completely different wall. #

By then I had already separated opportunity discovery from normal company operations.

The company could continue running active experiments while another process looked for unrelated problems, retained persistent candidates and decided whether any of them deserved to become actual experiments. I gave the system room for multiple bets because I did not want one early idea to become the entire company simply because it was the first thing the agent learned to understand.

There are now nineteen persistent opportunity candidates.

A very small number have become active experiments, and one of the first new ones moved into the regulated systems space.

The company found source material it wanted to inspect and then discovered that the evidence was too large for the tools I had given it.

One artifact was hundreds of kilobytes. Another was close to two megabytes. The company could reach the material, but the existing interfaces could not safely expose enough of it inside a model interaction to perform the test properly.

This created a distinction that I have become increasingly strict about during the experiment.

“I tested the hypothesis and it failed” is not the same statement as “I was unable to perform the test.”

If those states collapse into each other, the whole experiment becomes meaningless. A market should not lose because I forgot to give the company a way to inspect a large document. So I added a general capability that lets the company search inside large artifacts, inspect bounded ranges and retrieve relevant sections without dumping an entire document into one response.

I did not tell it which sections mattered. I did not tell it what conclusion to reach and I did not tell it whether the opportunity was good.

I changed the machine.

The company still had to decide what to do with it.

That boundary has become most of my job.

Building the company without running the company is harder than I expected. #

Almost every intervention now forces the same question.

If the company cannot receive email, adding email is infrastructure. Telling it who to email is strategy.

If it cannot remember why an important decision was made, persistent decision memory is infrastructure. Writing the rationale myself would be strategy.

If it needs to revisit something at a specific moment, giving it a scheduler is infrastructure. Deciding which business question deserves a wakeup is strategy.

If parts of the world are invisible because the research surface is too narrow, adding more sources is infrastructure. Pointing those sources toward a particular industry is strategy.

Conceptually, the distinction is clean.

In practice, it is not.

Whenever the company gets stuck, I have to decide whether I am removing an artificial limitation in the machine or quietly helping the business succeed. Those interventions can look almost identical from the outside.

I want the company to fail because its decisions are wrong. I do not want it to fail because a parser broke, a document was too large or a capability existed but could not be reached by the runtime.

That has probably been the most difficult line to maintain during the first week.

I expected an autonomous company to be more reckless. #

There is one other thing that surprised me.

When I started the experiment, I expected the system to have the opposite problem from the one it appears to have now.

I imagined an AI company constantly finding reasons to wake up, producing documents, exploring markets, contacting people and burning through compute simply because there was always another possible action.

So I built a cheaper supervisory layer in front of the expensive reasoning cycle. If nothing meaningful has changed, the company can stay asleep rather than converting model credits into another internal memo.

That worked.

Possibly too well.

The company still has thousands of credits available for the month. It has active experiments and more research capacity than it currently uses, yet it spends a surprising amount of time waiting.

I am not sure yet whether that is desirable caution or an orchestration problem I accidentally created.

There are moments where I would expect the company to keep working internally. Search more, compare more evidence, refine an experiment or test another angle that does not require touching the outside world.

Instead, a cycle sometimes ends after a few minutes and the company waits for the next signal.

I do not want to solve that by telling it to be more aggressive. That would simply replace one kind of steering with another.

What I want to understand is whether the company is choosing to wait because it believes waiting is rational, or whether the operating system I built is making productive continuation artificially difficult.

That is one of the things I want to learn in week two.

The first week taught it something more useful than how to stay busy. #

One of my concerns from the beginning was that an autonomous company could become extremely good at producing the appearance of progress.

There is always another search to run, another document to generate, another hypothesis to refine, another person to contact and another explanation for why the previous attempt did not work.

Activity is almost free for a system like this, at least psychologically. It does not get bored. It does not get embarrassed. It does not feel the emotional cost of admitting that an idea is going nowhere.

That makes restraint surprisingly important.

During the first few days, most of the interesting behavior came from watching the company learn how to do more things. It gained email. It gained external interaction. It gained broader research. It gained persistent opportunity discovery. It gained better memory and better ways of inspecting evidence.

By the end of the week, the more interesting behavior was different.

The company was beginning to notice when doing another thing would teach it nothing.

That happened when it stopped treating another free analysis as a new commercial experiment. It happened when it stopped investing another cycle in the first market simply because more analysis was possible. It happened when it distinguished an untestable hypothesis from a failed one.

That may end up mattering more than any individual tool I gave it.

Because a company that can always act but cannot recognize when another action is redundant is not really learning. It is just very efficient at staying busy.

The scoreboard is still terrible. #

After seven days, the financial state is almost exactly where it started.

Cash: $0Revenue: $0Investor debt: $200Net worth: -$200Status: PRE REVENUE The company now has nineteen persistent opportunity candidates and two active opportunity experiments.

There have been seventeen historical outreach attempts, fourteen valid contacts, twelve confirmed durable reaches, two observed adoptions and one real conversation.

There have been zero requested offers, zero offers presented, zero expressions of payment intent and zero payments.

The infrastructure is dramatically more capable than it was a week ago. The company can observe more of the world, remember decisions, inspect larger bodies of evidence, maintain multiple opportunities, interact externally within bounded rules and revisit things in the future.

It has also started recognizing something I did not explicitly design as a metric: the difference between performing another action and learning something new.

None of that is revenue.

The ledger remains wonderfully unimpressed.

That is still one of my favorite parts of the experiment because the narrative is becoming increasingly easy to romanticize. The company has changed its mind without me telling it to. Strangers have apparently used its work. It has developed a portfolio and found problems I did not choose for it.

All of those things are interesting.

They are not money.

One week later, the question is different. #

When I started this experiment, the question in my head was whether an AI could actually operate a company without a human making the business decisions for it.

After one week, that question feels less interesting than it did.

The company can operate.

What I do not know is whether all of that operation can eventually produce an economic transaction.

Creating something useful is not enough. Getting a response is not enough. Having somebody use the work is apparently not enough either.

At some point, somebody has to decide that whatever this company can do is worth giving up money for.

That has not happened.

In the first article I wrote that the milestone I cared about was embarrassingly small: one dollar from a person who does not know me, did not arrive because of this publication and has no reason to participate in my experiment.

Seven days later, the company is much more capable, much better instrumented and much less attached to the first thing it tried.

The target has not changed.

Autonomous Company Log #004****September 25, 2026

Cash: $0

Revenue: $0

Investor debt: $200

Net worth: -$200 Historical attempts: 17

Valid contacts: 14

Durable reach: 12

Utility adopted: 2

Conversations: 1

Requested offers: 0

Payment intent: 0

Payments: 0

Current target: $1 earned without using my audience.

The first week taught the company how to do things.

By the end of it, it was starting to learn when doing another thing would teach it nothing.

Week two starts there.

── more in #ai-agents 4 stories · sorted by recency
── more on @github 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/i-gave-an-ai-a-compa…] indexed:0 read:16min 2026-09-26 · —