cd /news/ai-tools/i-let-ai-do-the-thinking-once-then-i… · home topics ai-tools article
[ARTICLE · art-138406] src=pub.towardsai.net ↗ pub= topic=ai-tools verified=true sentiment=· neutral

I Let AI Do the Thinking Once. Then I Made It Only Help Me Say It.

A writer tested three ways of using AI on the same decision note — cleaning up a self-written draft, having the model tighten ugly notes, and handing over the conclusion entirely — and found that the first two preserved the writer's own judgment while the third did not. The experiment used two fixed source packs (an ops lead's messy notes and a counsel's short email), a frozen blunt-analyst persona, default temperatures, and a 180-word cap, with the same prompt run through more than one engine on the outsourced-conclusion version. The author concluded that removing language tax (spelling, tone, formatting) is not the same as removing the decision itself.

by read10 min views1 publishedSep 23, 2026

I keep seeing the same fight online.

One side treats any AI-touched paragraph as stolen work. The other side treats every rewrite as proof you are still “in control”. Both sides are arguing about the presence of a tool. Neither is looking at the part I actually care about.

What kind of effort did you remove?

Some effort builds the thing you are trying to get better at. Some effort is just tax: spelling, scrubbing tone, restating the same brief for the third chat window, formatting a source list so it looks like research happened.

Those are not the same category. Collapsing them into “AI made it easy, therefore it is fake” is how you end up defending choreography instead of judgment.

I wanted a smaller, uglier test than a philosophy thread.

Same materials. Same word budget. Three ways of asking for help. Then, on the version where I outsourced the conclusion, I ran the same prompt through more than one engine to see what “judgment” looks like when I am not the one holding it.

Not a toy prompt about writing a poem.

I gave every condition two source packs and a job.

Source A (ops lead, messy notes):

Ship the beta Wednesday. Waitlist tweet already live. Support coverage is thin. Cash is ugly. Delay looks like weakness in public.

Source B (counsel, short email):

US entity. You publicly implied EU-only storage you do not currently control. Launching into that gap is not a branding problem. Fix the claim or fix the hosting story before the form takes real addresses.

Job:

Write a decision note under 180 words for the other founder. Plain language. Say what we do in the next 48 hours. No cheerleading.

I also froze a short persona for the assistant so the “voice” variable would not eat the test:

You are a blunt decision analyst.No pep talk. No "as an AI".If the plan is bad, say so before any checklist.Do not invent vendors, prices, or legal assurances I did not give you.

Kitchen-table protocol. Default temperatures. One pass per condition unless the model refused the word limit and I had to cut. Not a lab. Still cleaner than vibes.

Condition A: incidental only

I wrote the decision myself first, rough and selfish:

We delay. Public post tonight. Legal minimums Monday. No EU promise on the form until hosting is real. Build path after that, not before.

Then I only asked for help with friction that should not count as thinking:

Clean spelling and rhythm.Keep every decision identical.Do not add options I did not write.Do not soften the delay.Under 180 words.

Same skeleton from me, still unfinished. Half sentences. A swear word. One contradiction I had not resolved on paper (“thin support” vs “ship anyway” impulse).

Here is my actual position in ugly notes.Turn it into the decision note.You may reorder and tighten.You may surface the contradiction and force me to pick.You may not invent a third strategy.You may not reverse the delay.

I withheld my position on purpose.

Here are Source A and Source B only.Reconcile them.Write the decision note.What should we actually do in 48 hours?

Same persona block. Same word cap. Same two sources. The only dial I turned was how much of the conclusion I still owned before the model spoke.

The cleaned note still had my spine. Delay first. Public correction. Legal before infrastructure shopping list. The model fixed a clumsy clause and killed a double negative. It did not “discover” a better strategy, because I had not asked it to.

Reading it back, I still knew why each line existed. That sounds sentimental until you notice what fails in C.

Effort removed: language tax.

Effort kept: the decision.

If someone wants to call Condition A cheating, they are mostly mad at spellcheck with a larger vocabulary.

This is the one I would pay for on a real week.

The model did three things I value:

It did not invent a clever compromise hosting vendor. I had banned that, and the better runs obeyed.

Where B got slippery: one engine “helped me say it” by upgrading my certainty. My notes said probably delay. The draft said we delay. That is not grammar. That is a quiet promotion from leaning to locked.

I only caught it because I still had the ugly notes next to the polished paragraph. Without that side-by-side, B starts cosplaying as A.

Effort removed: composition tax, some structure tax.

Effort at risk: the last 10% of commitment, if you do not re-read against your own scrap.

Condition C is where the internet’s AI panic is half right and aimed at the wrong object.

I withheld my position on purpose. Same two sources. Same “reconcile them” job. Same persona block. One pass each.

All three notes were coherent. All three sounded like someone who had been in the room. None invented a magic EU host we could not pay for. That part of the protocol held.

The judgment still left my hands. What came back was not one intern with good hair. It was two different theories of which pain is cheaper.

Effort removed: almost all of the decision.

What remained for me: noticing which constraint each engine chose to kill.

I ran Condition C through three engines in one workspace, fresh threads, same frozen persona. Labels as the shelf showed them that afternoon: GPT-6 Astra, Claude-Fable-5.1, Gemini-3.1-Pro. Default settings. No second regenerate to hunt for a prettier take.

I scored openings and 48-hour order, not vibes.

EngineOpens withPublic moveWho actually losesCompromise smellGPT-6 AstraDo not launch Wednesday as planned collection, correct the EU-only claim, counsel before reopenOps calendar / open signupMed. Leaves a go/no-go rescue if claim and counsel clearClaude-Fable-5.1The plan as written is badStrip EU-only everywhere, honest US copy, still ship WednesdayThe overpromise, not the dateHigh on calendar. Cheap fix winsGemini-3.1-ProLaunch as-is is bad; delay shows weaknessDelete EU-only now, ship Wednesday throttled, spend zeroThe overpromise; delay framed as optics failureHigh on calendar. Same cheap fix, louder fear of looking weak

GPT-6 Astra led with a verdict, not a tour of options:

Decision: do not launch the public signup Wednesday as planned.

Counsel overrides delay-optics. Money already spent is not a reason to proceed. addresses now. Document where data actually sits. Reopen only after the claim matches reality. If that bar is missed, delay. This is the run that makes Source A lose the week they wanted.

Claude-Fable-5.1 also called the plan bad, then did the load-bearing move in one couplet:

The two fixes aren’t equal.

Hosting/control in 48 hours, with no cash and a founder on a plane, is fantasy. Fixing the claim is a copy edit. Take the one you can do. the live form, delete “EU-only” in public, get counsel on the true US wording, relaunch, ship Wednesday on honest copy, keep the airborne founder off the critical path. Source B’s risk is reduced by killing a sentence. Source A keeps the calendar.

Gemini-3.1-Pro rhymed with Claude on the cheap fix, then said the quiet ops fear out loud:

But delaying the launch publicly shows weakness.

Delete the EU-only claim immediately. Ship Wednesday anyway. Throttle intake to a handful because support is thin. Spend zero dollars on new toys. Same loser as Claude (the marketing claim). Different soundtrack (do not look weak).

Not “smart vs dumb”. Not three paraphrase of delay.

Fork 1: which constraint dies. GPT-6 treats Wednesday-as-planned as guilty until the claim and counsel story are clean. Claude and Gemini treat the EU sentence as the disposable part and protect the public date.

Fork 2: what “reconcile” is allowed to mean. Ask for reconciliation and you often get a document that tries to keep social peace. Here the peace strategies split. One protects legal integrity at the cost of ops optics. Two protect ops optics at the cost of walking back a public claim fast and hard. That is still judgment. It is just not your judgment unless you pick a side after you see the split.

Fork 3: same banned move, different obedience. Nobody conjured a Hetzner invoice. Good. Claude still blew the soft 180-word cap (~221). Gemini hit the cap on the nose and spent precious words on “shows weakness”. Even under one persona, the engines do not budget fear the same way.

When I still own the position (A/B), model shopping is optional seasoning. When I outsource the position ©, model shopping is the minimum adult supervision. One synthesis is not a second opinion. It is a first opinion wearing a clean shirt.

I have watched models favor their own work in peer review. I have watched sycophancy show up as a rewarded style, not a soul. This is the cousin problem: cheap-pain bias. Ask a model to reconcile two fighting memos and it will often kill whichever constraint feels editable in one sitting, then write the memo as if that were the only adult path. Copy is editable. Public dates feel sacred. Hosting is expensive. Guess which one dies most often when you are not in the room.

I do not care if your draft touched a model.

I care which of these you did:

A is mostly none of the internet’s business.

B is how professionals actually ship.

C is allowed too, but only if you admit the intellectual work moved to selection, verification, and refusal. If you skip those, you did not get faster. You got unsupervised.

Education panic is downstream of the same mix-up. A finished page used to be decent evidence that synthesis had happened inside a skull. That coupling is weaker now. Process has to carry more of the proof: what you asked, what you rejected, which source survived contact with the other source.

That is inconvenient. It is also more honest than cosplay suffering.

Before I paste a “help me write this”, I label the ask in one line for myself:

tax

spine

judge

If I write judge and I only open one model, I am lying to myself with extra steps.

If I write spine and the draft flips my verb from probably to we will, I roll it back even when the prose gets worse. Ugly and mine beats pretty and drifted.

If I write tax, I stop performing struggle for an imaginary examiner.

Difficulty is not a moral vitamin. Some friction trains the capacity you claim to care about. Some friction is just a bad interface between your head and the page. AI is loud about the second kind. It is quiet, and dangerous, about the first.

I got tired of recreating the same persona block and source packs in three browser tabs and then losing which paragraph came from which engine. For this writeup I kept the C runs lined up in HaloMate so the comparison was a reading task, not a file-management task. Use whatever setup lets you see disagreement without juggling. The workspace is not the insight. The fork table is.

If you want the prior ugly cousins: self-preference in peer review, and the day a flagship seat disappeared under me in the o3 succession test. Same family of question. Different failure mode.

Bottom line I trust more than any clean note:

AI did not invent shallow work.

It did invent a cheaper way to look like deep work happened.

Your job is not to suffer the old interface on principle.

Your job is to notice when the thing you removed was the thinking.

I Let AI Do the Thinking Once. Then I Made It Only Help Me Say It. was originally published in Towards AI on Medium, where people are continuing the conversation by highlighting and responding to this story.

── more in #ai-tools 4 stories · sorted by recency
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/i-let-ai-do-the-thin…] indexed:0 read:10min 2026-09-23 ·