My recent experience is that prototyping is needed when designing multi-step workflows with more than 20 screens. Let me recap what didn't work, and what did work.
The context:
Emm is a local first notes manager and issue tracker for agents and humans.
It's important to me that users can start building with local agents without ever connecting to the cloud.
This creates complexity. A project exists prior to it being backed up in my cloud. Additionally, a user might have many projects. When they sync to Emm's server, they need a signed-in user. The user might want to change the signed in user. And, on top of all of this: as an young company, first time user experience is existentially important.
First try: trust the agent to design a workflow
This was a disaster. I asked Astra to ask me questions, identify key decisions, generate options, then create a spec. I then let it implement the spec.
To the agent's credit, it handled the plumbing that years ago would have taken a few days. However, the flow it selected required the user to make, for example, 7 clicks when she only needed to make 4; it navigated the user to pages where the no decision was being made, when it could consolidate the flow into fewer pages.
Second try: Build one screen at time.
My next strategy was to build the full experience, one single screen at a time. My strategy was to start with creating a project, then work through putting it on the cloud, then inviting a user, then accepting a user, then... etc.
I gave up on the agent having taste. I told it specifically what I wanted, waited for it to implement it, gave feedback, then asked for the next screen.
This was also a fail, for two reasons. My app is pure rust, no framework. This means that even simple changes to the application take the agents 10-20 minutes to make. So, the process of adding one single branch of a login could take hours.
The second problem is that design is iterative. I built something that basically worked, but was confusing to the user and didn't handle all edge cases well. To re-build it using the old method would take another 8+ hours
Third try: Success: ChatGPT desktop HTML prototype
Finally, I switched from Herdr-hosted codex to ChatGPT Desktop. I asked the agent to build out the entire mock application. The windows I would see would span different surfaces - web browser, my app, pop-ups from MacOS passkey, email - and also include pages that aren't tied to any surface. It opted for a drop down for testing from different starting configurations.
It still took me a solid three hours to work through - the task was iterative, and required a lot of decisions and testing. And to be honest by 11:30 I was not my sharpest - but the end result was a full prototype that made the desired behavior crystal clear and covered all use cases. I let in run all night completing the work.
Lesson
When designing user workflow that have complex back-end machinery, have the AI agent create a working prototype in HTML, so that you can quickly test and work through every situation, then have the agent implement in one go.