cd /news/artificial-intelligence/how-to-get-better-every-week-with-ai · home topics artificial-intelligence article
[ARTICLE · art-84553] src=theaithinker.com ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

How to get better every week with AI

A former product leader's practice of weekly self-review, inspired by James Clear's 'Atomic Habits' 1% improvement rule, can be supercharged by using large language models like Claude to analyze verbatim transcripts of interviews and meetings, providing specific, evidence-quoting feedback that human reviewers no longer offer. The article, authored by an unnamed writer, details three feedback loops—interviews, meetings, and Friday retros—with exact prompts, and addresses the challenge of preventing AI flattery, aiming to help product leaders improve their craft weekly.

read23 min views1 publishedAug 3, 2026
How to get better every week with AI
Image: Theaithinker (auto-discovered)

Paste a real transcript into any LLM. Ask it to review you, not the user. Debrief your hardest meetings. Automate the Friday retro with a skill.

A few years ago, at a previous company, I sat through a leadership training I’ve mostly forgotten. One idea survived. The most senior product leader in the room told us to end every week the same way: write yourself three or four honest lines about how the week went, and what you’d do differently on Monday. His pitch was simple math, borrowed from the idea James Clear made famous in Atomic Habits: get 1% better every week, and compounding does the rest. I’ve written that little Friday retro almost every week since, and it works.

For years, though, the habit had a flaw I couldn’t see. The reviewer was me. My memory of the week is a friendly narrator: it remembers the demo that landed, and it quietly forgets the question I dodged in a tense meeting. Every Friday I graded my own homework, and every Friday I passed. A retro written from memory is a review of the week you’d like to have had. Better than nothing, but it plateaus fast. Then one afternoon I pasted the verbatim transcript of a user interview I’d just run into Claude. Not to summarize it (I do that all the time). I asked it to review my interviewing: how I phrased questions, what I missed, how much I talked. It found three leading questions in my first ten minutes and quoted them back word for word. Nobody had reviewed my work at that level of detail in years.

Everyone I know uses AI to do the work: draft the spec, summarize the research, write the update. Almost nobody uses it to review how they work. The first use saves you hours this week. The second one compounds, because it makes every next interview, meeting, and decision slightly better than the last. That’s the 1% from the training, and AI finally makes it cheap to collect.

The whole practice fits in one picture.

The gap. I’ll show you why nobody reviews your actual craft anymore, and why the feedback you do get says more about the giver than about you.The loops. The three I run (interviews, meetings, the Friday retro) with the exact prompts I use, plus six more worth trying, from your forecasts to your 360 packet.The honesty problem. Here’s how to stop the model from flattering you, because by default it will.The team angle. You’ll build a skill that asks your PMs your own 1:1 questions before they ever reach you.

By Friday you can have run your first loop on a real artifact from your own week. You get the kind of specific, evidence-quoting feedback nobody has given you since you got senior, and you can get it every single week. The first loop needs nothing new: any LLM chat works (Claude, ChatGPT, Gemini), and the raw material is the transcripts your meeting tools already produce (Zoom AI Companion, Gemini in Google Meet, Granola, or a voice memo on your phone). And if you lead PMs, you’ll leave with two skills worth an afternoon each: one that reviews your week on a schedule, one that preps your team for 1:1s.

Your reviewer is ready. Let’s put it to work.

Nobody reviews your work anymore #

Think about the last time someone watched you work and told you something true about it. Not your outcomes: your craft. The way you ran the interview, the way you handled pushback in the room. If you lead product teams, the honest answer is probably “years ago.” Your manager sees headlines and results, not how you got them. Your team won’t critique you upward. The more senior you get, the less anyone reviews how you actually work. What’s left is the official cycle: twice a year, compressed into themes like “communicate more.”

And hold even that loosely. In The Feedback Fallacy (Harvard Business Review, 2019), Marcus Buckingham and Ashley Goodall assembled the research on how well humans rate other humans. The headline finding has a name, the idiosyncratic rater effect: more than half of any rating of you reflects the rater, not you, and no amount of training fixes it. In their words, feedback is “more distortion than truth.” Your manager’s read on your “communication” says as much about their definition of communication as about yours, and that stays true for the best managers you’ll ever have. The little feedback that reaches you arrives filtered through someone else’s lens, and the research says the filter is most of the signal. Not a reason to dismiss your manager. A reason not to make two filtered data points a year your only mirror.

Performance-heavy fields treat this as a solved problem. Athletes and musicians review tape after every game and every recital. Atul Gawande wrote the canonical essay on this, Personal Best (2011): a surgeon at the top of his field plateaued after eight years, then started improving again when he put a reviewer back in his operating room. The research on expertise agrees. Anders Ericsson’s work on deliberate practice (Peak, 2016) found that experience without feedback doesn’t accumulate into skill; it plateaus. Twenty years of interviews don’t make you better at interviewing if nobody ever shows you what you did.

What changed isn’t the theory. It’s the tape. Your meeting tools record by default now: Zoom AI Companion, Gemini in Google Meet, Granola quietly taking notes on your laptop. Verbatim records of you working have been piling up in your drive for a couple of years. We skim the summaries for action items and move on. The evidence of how you work is exhaust of a normal week, and almost nobody reads it in the review direction.

One precision before you touch that pile. The AI notes your tools generate are summaries, and a summary is already the friendly narrator’s cut, machine-made this time. The loop needs the verbatim transcript. In Google Meet, that’s the separate “Transcripts” option, not the Gemini notes doc. In Zoom, it’s the transcript file attached to the recording. Granola keeps the full conversation one click behind its polished notes. Feed the model what was actually said, not what the notes decided to remember.

And the reviewer for all that unread tape? You already have it. You’ve just been pointing it the other way.

You already run AI as a production engine. It drafts the spec, summarizes the research, writes the launch note. Point the same model at a transcript of how you worked, and it becomes a feedback engine. Same tool, opposite direction. In that direction it has qualities no human reviewer can offer: the patience to re-read your entire week verbatim, no stake in your ego, and availability at 6 pm on a Friday. The only reviewer with time to watch all your tape has been sitting in your browser tab all along.

The direction flip is the whole method. Here’s what it looks like on a real week.

Start with my three loops #

I run three loops on my own weeks, and they share one shape. Take a verbatim artifact your week already produced. Ask the model to review how you worked, not what was decided. Extract one change for next time. Write that change down. The artifact does the honesty, the model does the patience, and the written line makes it stick. Here’s the shape before the details.

Three isn’t a magic number, and my three aren’t the menu. They’re worked examples, mapped to where my own weeks leave a record: user interviews, hard meetings, and the week itself. Treat them as illustrations of the shape, not the full list; your calendar will suggest loops mine can’t. A catalog of other places to point the same shape comes right after these three.

Loop 1: Review your interviews

After a user interview, I copy the full transcript into the LLM and ask it to review the interviewer. Not the user’s answers, not a summary of insights: me. The phrasing of my questions, the follow-ups I didn’t ask, how much of the conversation was my own voice. The distinction matters because “summarize this interview” is production work, and every PM already does it. ”Review how I ran this interview” is feedback work, and it’s the version almost nobody asks for.

Here’s the prompt I use, lightly cleaned up. Yours will look different; the load-bearing parts are the named rubric and the quoted evidence.

Here is the verbatim transcript of a user interview I ran today. Review the interviewer (me), not the user.

Grade me against The Mom Test’s three rules:talk about their life, not my idea; ask about specifics in the past, not hypotheticals about the future; talk less, listen more.For every finding, quote the exact line from the transcript.List every question I asked, label each one open or closed, then tally the list. End with the one change that would most improve my next interview.

The rubric is The Mom Test, Rob Fitzpatrick’s 2013 book on customer conversations, and spelling its rules out inside the prompt matters. The critique works even if you (or the model) never opened the book, and you learn the rubric by reading your own violations of it. And the leading questions were only the loudest finding. The same session flagged the moment a user mentioned a workaround and I moved on instead of digging. It listed my questions, and over half were closed. None of that was in my notes, because my notes were about the user.

The critique is the start, not the verdict. I discuss it, and the discussion is where the learning happens. The rule I follow: seek evidence, don’t argue. “Show me the line” is a safe question. “Rewrite that question the way I should have asked it” is a great one. Protesting (”I don’t think that was leading”) is the one move to avoid, because models tend to fold when you push back, and a critique that folds is worthless. Treat the chat like a film session with a coach, not a negotiation over your grade.

Interviews are the easiest tape to start with: low stakes, famous rubric, fast wins. The harder tape is the meeting where you had something to lose.

Loop 2: Debrief your hard meetings

Some meetings matter more than others. The technical review where the architecture pushback derailed your agenda. The leadership meeting where your update ran long and the ask got lost. Your calendar is full of them, and that’s not the failure it’s usually made out to be: meetings are where product work happens. Spending most of your week in meetings isn’t the problem. Letting them end with nothing actionable is.

I used to replay the hard ones in my head on the way home, which felt like reflection but was really just rumination. Now I feed the transcript (or my detailed notes, when there’s no recording) to the LLM and ask for a debrief of my own performance. How I presented, where I lost the room, how I handled the pushback, what to change next time. The hours were already spent; the debrief is how they pay something back.

This move has a precedent you already practice at the team level. Product teams close every sprint with a retrospective: what went well, what didn’t, what changes next time, run by the team, for the team. Norman Kerth, who wrote the practice’s handbook (Project Retrospectives, 2001), even gave it a safety rule, the Prime Directive: whatever we discover, everyone did the best they could with what they knew at the time. Yet the ritual usually stops at the team boundary, and nobody runs one on their own performance in a single meeting. A meeting debrief is a one-person retrospective, and AI makes it cheap enough to run after every meeting that matters. The prompt is shorter than the interview one, because there’s no famous rubric for meetings. I anchor it on my goal instead.

Here is the transcript of a technical review I ran today. My goal was alignment on shipping the migration in two releases instead of one.

Review my performance, not the decisions:how I opened, how I responded to pushback, where I talked past someone, where I lost the thread.Quote the line for every finding.Then give me the strongest case that this meeting went badly for me. End with what I should do differently in my next meeting with this group.

The debrief pays out twice. Once right after the meeting, when it shows you the exact moment you started answering a different question than the one asked. And once before the next meeting, when you reopen the same chat and prep: same audience next month, help me open differently, here’s what I’m walking in with. One honest limit from experience: this loop needs an artifact. Hallway conversations leave no tape, and a debrief of your own recollection is the friendly narrator reviewing itself.

Interviews and meetings are single events. The third loop is what strings them into actual improvement.

Loop 3: Review the whole week

This is the original habit from that training, and it’s still the anchor of everything above. Every Friday, three or four honest lines: what went well, what didn’t, what changes on Monday. Five minutes, no template, one running file. The one-change outputs from the interview and meeting loops land here as lines. Without the written close, the critiques evaporate by Tuesday.

But run the first two loops for a few weeks and you meet their limit: they only cover the events you picked. A product week is fifteen meetings, and you reviewed two. Nobody has the hours to chat-review a whole week one transcript at a time, and pasting fifteen transcripts into one window isn’t a workflow, it’s a punishment. The weekly review doesn’t need a better prompt. It needs to become a system. Same folder of inputs, same rubrics, same output format, every single week: that’s not a conversation, that’s a job. And a repeated job with fixed rules is exactly what Claude’s skills are built for: a skill is a folder with an instructions file that teaches Claude a task once, so it can run it every time after.

Here’s the shape of the system to build, and it’s smaller than it sounds.

The skill’s instructions fit on one page: read every transcript in this week’s folder, review my performance in each one against my rubrics (The Mom Test rules for interviews, my stated goal for everything else), quote the line for every finding, then draft my three-or-four-line retro and append it to the running file.

Collecting the folder is lighter than it sounds. Zoom transcripts download as text files, Google Meet saves them as Docs in a Drive folder, and Granola keeps the full conversations behind its notes. And scale is where the review gets interesting: a reviewer that reads all fifteen meetings can tell you the same tic showed up in four of them, which no single-meeting debrief will ever surface. If you’ve never built a skill, it’s an afternoon, not a project: I walked through the full method in How to build an AI helper for your team with Cowork, and the same steps apply here with a different instructions page.

Then remove the last manual step. Claude Cowork now supports scheduled tasks: set the skill to run every Friday afternoon and it executes remotely, even with your laptop closed. You open the draft with your coffee, check the quoted lines, and spend your five minutes editing instead of reconstructing.

The honest writing stays yours; the system makes sure it starts from evidence instead of memory, at the scale of your whole week.

Now, about that 1% pitch. Clear’s canonical version is 1% per day, which compounds to 37x in a year. The training I got was humbler: 1% per week, which works out to about 68% better over a year (1.01^52 is 1.68, if you want to check the math). Treat both as the spirit of the practice, not a forecast. Skill isn’t a number, and nobody can measure 1% of interviewing. The idea’s lineage runs through British Cycling’s ”aggregation of marginal gains”, but the version that changed my weeks came from a trainer with a marker and a flipchart. Small weekly changes stack, and a written retro is what keeps them from resetting to zero.

Three loops, one of them running itself. But the shape doesn’t stop where my calendar does.

More loops to try #

Once you see the shape (artifact in, critique against a standard, one change out), you start spotting loops everywhere. I went looking for the ones with the strongest cases behind them, and this is the shortlist I’m working through next. If a week of your work leaves an artifact, there’s a loop that can read it.

The decision journal. When Michael Mauboussin asked Daniel Kahneman what single thing would most improve decision-making, the answer wasa cheap notebook:write down each big decision and what you expect to happen. Months later, the LLM grades your calibration against outcomes: “you were overconfident on every timeline that depended on another team.”The forecast scorecard. Philip Tetlock’sSuperforecastingis blunt about how forecasters improve: “Forecast, measure, revise: it is the surest path to seeing better.” Put numbers on your own calls (”70% we ship by March”), then ask the model each quarter where your numbers run hot.The spec-versus-shipped review. Feed January’s PRD and June’s launch numbers into one chat and run a retrospective on the pair: what was supposed to happen, what happened, why, and what changes next time. The gap between those two documents is feedback nobody schedules a meeting for.The rehearsal review. Toastmasters assigns a dedicated person, theAh-Counter, just to count filler words. Record one run-through of your next big presentation, and the model plays that role free, then tells you how many minutes pass before your headline number shows up.The hidden-ask audit. Paste a week of your sent emails and Slack requests and ask which sentence carries the actual ask. Consultants are trained onBarbara Minto’s Pyramid Principle, answer first and reasoning after; most of us put the answer somewhere in paragraph three.The 360 synthesis. Paste every review packet you’ve received and ask what three or more people said independently. Any single rating is mostly rater (that’s the idiosyncratic rater effect again), but a theme that survives five raters and three years is genuinely yours.

Whichever loops you pick, they all lean on one fragile assumption: that the reviewer tells you the truth. By default, it won’t. The model is trained to be nice to you.

Make the critique honest #

“The LLM just flatters me.” If that’s your objection, you’re right, and it deserves a serious answer. Sycophancy is a documented, trained-in behavior of assistant models across vendors, not a paranoid suspicion. Anthropic researchers showed that state-of-the-art assistants are consistently sycophantic, likely because human preference data rewards agreement (Sharma et al., published at ICLR 2024). You’ve probably felt it yourself: in April 2025, ChatGPT got so flattering that OpenAI rolled the model back within days. Ask a naive “how did I do?” and you’ll get praise plus three soft suggestions. By default, the model is a cheerleader wearing a reviewer’s badge.

The fix isn’t finding a more honest model. It’s writing prompts that leave no room for flattery. Five mechanics do that work for me.

Give it the artifact, never your self-report. The model can’t flatter its way past a transcript, but it will happily flatter its way through your own account of the meeting. You already know which narrator wrote that one. Asking about yourself is the worst case:one 2025 studyfound LLMs preserve the asker’s self-image far more than human advisers do. The tape is your protection.Name a rubric.“Grade this against The Mom Test” beats “was this good?” every time. A named standard gives the model permission to find faults and a scale to grade on. No famous rubric for your artifact? State your goal for the meeting and let that be the standard.Require quoted evidence. Every finding must cite the exact line. This converts vague praise into checkable findings, and it screens out invented criticism too. The same rule governs the discussion afterward: ask “show me the line,” never argue “I don’t think that was leading.” The same Anthropic research found assistants wrongly walk back correct answers when users express doubt.**Evidence-seeking keeps the critique standing; arguing makes it fold.**Ask for counts, but make it list first. Talk ratio, open versus closed questions, minutes before the first real question. Useful numbers, and a known weak spot: a one-shot count is an estimate wearing a number’s costume. So make the model quote every question, label each one, then tally the list. The list is checkable against the tape, and the tally of a list is reliable.Make it argue the prosecution.“Give me the strongest case that this interview went badly.” The model is excellent at arguing a side when you tell it which side to argue. Flattery has nowhere to hide in a prosecution brief.

One more objection deserves a straight answer, because legal will raise it before you do. Transcripts carry other people’s words. User interviews sit under whatever consent your research process already runs, so check that first. The load-bearing rule is the workspace: run the loop inside your company’s approved AI workspace, the same environment where your meeting tool already stores the recording (hand-anonymizing a 40-minute conversation is fiction, so don’t build your compliance on it). Strip names where it’s quick to do. And keep other people’s performance out of scope: the loop reviews you, not your colleagues.

Honest critique, ten minutes a week, grounded in what was actually said. The last move turns a personal practice into a leadership one.

Bring it to your team #

Everything so far works for one person: you. But if you manage PMs, you’re also the person whose questions and standards your team is trying to guess at. The manager’s version of the practice has three moves: model it yourself, build it into a tool your team owns, and keep your hands off their critiques.

Show your own findings first

You don’t roll this out with a deck. The practice spreads in an embarrassingly simple way: you talk about your own findings first. In your next team meeting, mention what your review found. “It went through Tuesday’s interview and I did 40% of the talking. I’m working on it.” That sentence costs you thirty seconds of vulnerability and buys the whole room permission. Naming your own findings first is what makes the practice safe to copy.

Build the 1:1 prep skill

Then give the practice a tool. Every manager carries a mental list of the questions they always ask in a 1:1. What moved the metric this week? What’s the impact of the thing you shipped? What’s next, and what’s blocking it? Your team knows the list exists, but they only meet it one meeting at a time, live, when it’s too late to prepare. Put that list into a skill, and your questions become something your team can use instead of something they sit through. The skill’s context page holds three things: the company’s OKRs, the metrics your team owns, and your standing questions.

Here’s how it runs once it exists.

A PM runs their weekly update through the skill before your 1:1. It asks your questions before you do, checks the update against the real OKRs, and flags what’s thin: an impact claim with no number, a next step with no owner. They arrive having already answered the routine layer, on their own terms, with no audience. And the skill improves the same way you do: every time you catch yourself asking a question it missed, add it to the context page after the 1:1. I’ve built this exact shape before for AI-tooling questions, and the full walkthrough carries over here: same build, different context page.

Be honest with the team about what it’s for. The skill will never replace the conversation; it changes where the conversation starts. The routine questions get handled before the meeting, so the meeting spends its thirty minutes on what actually needs two humans: the trade-off they’re unsure about, the stakeholder they can’t read, the call that needs your judgment.

Keep the loop theirs

Hand over the loop, not a verdict. For the PM whose interviews keep coming back thin, the move isn’t “your interviews are weak.” It’s “run your next transcript through this prompt before our 1:1, and bring what surprised you.” The critique stays theirs. You never read their review; you only ever see the change they choose to make. That’s the oldest rule of the retrospective, the one Kerth wrote his Prime Directive to protect: the review belongs to the people who did the work, not to an evaluator. The moment you ask to see their critiques, the practice turns into surveillance, and people quietly stop running it.

What shifts is the shape of your 1:1s. You were never going to give transcript-level feedback to six PMs; there aren’t enough hours in your week, and there never were. Now that feedback exists without you. The conversation moves from “here’s what I think you did wrong” to “what did your review find, and what are you changing?” You stop being the bottleneck on your team’s craft feedback without becoming absent from it. You coach the change instead of hunting for the evidence.

And if you keep one line from all this, keep this one: you get a performance review twice a year; your transcripts can give you one every Friday. It works in a 1:1 with the PM who wants to grow faster. It works on your own tape, this week.

Start this Friday #

You don’t need a new tool, a habit system, or anyone’s permission. You need one artifact from this week and ten minutes on Friday. Pick the tape where you already suspect the truth lives.

Your last user interview. Start here if you run any: the rubric is famous and the wins come fast.The meeting that went sideways. The transcript is already in your drive. Ask for the debrief you’ve been replaying in your head anyway.Your old retros, if you keep any. Paste the last two months and ask for the pattern you keep repeating.A decision you’re about to make. Open the journal with one dated entry: what you decided, what you expect to happen. Future you collects the feedback.

The training that started this for me lasted one day, and I’ve forgotten almost all of it. The habit it left keeps compounding. That’s the honest promise of this whole practice: no transformation by Friday, just next-quarter you asking better questions, running tighter meetings, and knowing exactly what you’re working on. Start with one transcript. The 1% is already in there.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @james clear 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/how-to-get-better-ev…] indexed:0 read:23min 2026-08-03 ·