I asked an AI agent to help me market IssueFlow.
I use ChatGPT to help me manage the GTM campaign, a reasonable request: Find the right people, explain the product, help bring in customers. Ideally while I’m doing my full-time job and building two products, because apparently looking at my calendar and then at my wife - the obvious conclusion is: “I could use some help.”
So we started running ads.
Then my marketing agent pointed out something inconvenient: getting clicks is nice, but we needed better evidence of what happened afterward.
Did someone sign up? Did they complete a useful workflow? Did they become a paying customer?
Fair questions. But that needs integratin into IssueFlow code itself…
So I told the agent to create the development issues for the integration.
IssueFlow (which /I sue to develop IssueFlow) picked them up, and its configured agents started working through the delivery process. Triage, design, design-review…
Then the first design reached an approval gate.
That’s when things got interesting.
Normally, I review the work at that gate.
This time, the marketing agent had written the issues. It knew what it needed and why. So I explicitly authorized it to review the designs and approve them—or send them back.
To be clear: I delegated that decision for these issues. The system didn’t quietly decide that human approval was an outdated concept.
Now we had an agent requesting a feature, other agents designing it, and the requesting agent reviewing whether the proposed solution actually solved its problem.
An agent needed an agent.
We had successfully recreated a small software team. Thankfully, nobody scheduled a recurring alignment meeting.
The first issue was about signup tracking.
It had already passed an automated design review. But when the requesting agent examined it against the original requirements, it found gaps.
Some were technical: how the website would pass the visitor’s identity and session information to the portal, what happened when sending an event had an uncertain outcome, and how later conversions would retain useful attribution.
One particularly subtle problem concerned time.
The design’s timing rule was based on when information was handed over. The reviewer needed it tied to the actual session.
Those sound similar until someone leaves a tab open, returns later, and your neat little assumption starts doing interpretive dance.
The agent sent the design back with specific corrections.
honest confession: In this particular case - to be honest - I had no clue…
I am deeply involved in how IssueFlow handles, well… issue flow… but not how it integrate with external marketing platforms.
The revised design addressed several of them. The reviewer (the marketing agent) accepted those changes, kept the remaining concern open, and requested another focused revision (from the design agent).
IssueFlow is designed to handle those loops and handled it to a design-rework agent that can hadle higher complexity requests.
Eventually, that issue reached a design the requester could approve.
What mattered to me was that the feedback remained part of the work. The next review could focus on what had changed instead of starting the whole explanation again.
The next two issues covered activation and payments.
This is where the difference between “technically plausible” and “what we actually want” became very clear.
For activation, one proposed interpretation could count breaking a request into smaller child issues as meaningful progress. But that wasn’t the metric we wanted.
Creating more work is not the same as completing useful work.
If it were, my to-do list would be our most successful customer. The requester pushed back: activation should reflect a genuinely successful workflow outcome. Splitting a request into pieces wasn’t enough.
Agent vs Agent now - Astra vs Opus dialogue. And I am watching…
The payment design raised another question.
Suppose you can verify a payment, but you haven’t finished checking the earlier payment history. Can you call this the customer’s first payment?
You know money arrived. You don’t necessarily know that it was the first time.
That distinction matters when you’re trying to measure new paying customers.
The revised design separated those facts instead of forcing uncertainty into a confident-looking number.
There were also corrections around consent and recovery: analytics should respect the agreed consent rules, and a measurement failure shouldn’t prevent the underlying customer workflow from succeeding.
These were acceptance decisions. They were about preserving the meaning of the request through implementation.
Once the designs were approved, implementation continued.
And later checks found more problems.
IssueFlow continuted the flow: validation, security check, and then Code review caught implementation defects. Validation exposed missing test evidence. Related changes needed to be reconciled so they could work together.
That was useful, too.
A design gate doesn’t make everything afterward correct. It gives the work a better-defined direction, and later checks still have a job to do.
IssueFlow by itself rerouted the issue for correction until all gates passed.
All three issues eventually completed and their code reached a successful production deployment.
But even that wasn’t permission to announce, “Measurement solved!”
Configuration and actual processed-event verification were still separate work. Shipped code and a verified customer journey are different milestones.
That’s a distinction I want the system—and the people using it—to keep making.
I build IssueFlow, so this isn’t an independent customer review. It was a real exercise on the product we use to build the product.
And it wasn’t frictionless. Some reviews required too much reading. A clearer summary of “what you asked us to fix, what changed, and where the evidence is” would have made the process easier.
Still, the useful part was very concrete.
The requesting agent could inspect the work, challenge its interpretation, return specific feedback, and approve a revision when the important gaps were addressed.
The delivery workflow retained those decisions and continued.
I didn’t personally type every review response. I decided who could make those decisions for this work.
That’s the kind of control I care about: being able to delegate execution and judgment deliberately, with boundaries and a record I can inspect.
Also, apparently, having an AI agent discover the timeless software-development experience of saying:
That’s close. But it isn’t quite what I asked for.
Some things really do survive every technology shift.
If you’re building with coding agents, who checks that “done” still means what you originally wanted? I’m the developer behind IssueFlow, a workflow orchestration tool for AI coding agents. It helps automate the development process while keeping human judgment where you want it. This story comes from using it to build IssueFlow itself.