While turning my recent agent experiments into articles, I ran into an uncomfortable gap. I could follow the code, connect tools, and use AI to help debug a problem. Explaining why the system worked that way was much harder.
My notes covered ReAct, Reflection, and frameworks including LangGraph. But having a page for each topic did not mean I had a connected understanding of the system.
I decided to revisit three examples already in my own notes: a serial-device gateway, a Reflection coding exercise, and a LangGraph question-answering workflow. The exercises below are my next steps, rather than results I have already achieved.
In an earlier project, I used Python and pyserial to put a serial-device gateway behind an HTTP API. This let AI participate in testing a hardware module through tools.
At the time, I focused on getting the test workflow running. For revision, I want to follow one call across the boundaries:
Task and previous results
|
Model proposes a tool name and arguments
|
Execution program sends an HTTP request
|
Gateway uses pyserial to send a command and read the reply
|
Execution result is included in a later model request
|
Next action or final report
This is the simplified path in that project. Splitting it into stages gives each failure a place to investigate:
My first exercise is to take an existing test record and line up the task input, tool arguments, gateway response, and final report. If all I can find is “test passed,” I still need to locate the reply that supports that conclusion.
I do not need to rebuild the project to do this. I need to explain one complete execution path. Original serial-gateway notes, in Chinese
One of my Reflection exercises generated a Python function for finding prime numbers. Looking back at its output, I noticed a review that said no algorithmic improvement was necessary, followed by another code-optimization step.
That observation alone does not establish a bug. The feedback also discussed engineering improvements, and the task may have allowed further changes. It does raise a useful question: what actually makes this program continue or stop?
A model can produce an assessment, but the program still has to turn that assessment into an action. I want to inspect the stopping condition, the expected review format, the parser, and the iteration limit.
My revision exercise has two parts.
First, I will test the control flow with fixed feedback: a passing review, a request for changes, and an unrecognized response. I will predict the next branch before running it. This separates the program's control logic from variation in a real model response.
Second, I will test the generated function itself. If its contract is to return primes no greater than n, a few initial cases are:
expected = {
1: [],
2: [2],
10: [2, 3, 5, 7],
}
I still need broader coverage and boundary checks. The point is to compare the code's behavior with the reviewer's opinion. A more enthusiastic review does not establish that the implementation improved.
That makes Reflection concrete for me: saved results, feedback passed into the next request, explicit stopping behavior, and tests. My Reflection notes, in Chinese
My LangGraph notes include a three-stage assistant: understand the question, search, and generate an answer. Its state contains fields such as search_query, search_results, and step.
I used to treat those fields as implementation details. Now I see them as the places to ask basic questions: what did the search node receive, where was its output stored, and what did the answer node read?
In that example, the answer stage can fall back to the model's existing knowledge after search fails. A fluent final answer therefore does not prove that search succeeded.
I plan to hold the question fixed and inspect three inputs to the answer stage:
| Search outcome | What I want to check |
|---|---|
| Results returned | Can the answer's important claims be traced to those results? |
| Empty results | Does the answer confuse “not found” with “does not exist”? |
| Search error | Does it make clear that the external material was not retrieved? |
There is also a task-level decision: should this workflow continue answering after retrieval fails? For a question that depends on current information, the model's existing knowledge may be insufficient. Stopping and explaining the missing evidence may be more appropriate.
The fields in a workflow state give me something to inspect between nodes. They do not, by themselves, establish persistent storage or long-term memory. My LangGraph notes, in Chinese
Instead of starting a new project for every concept, I want to reuse the same sanitized logs and examples:
The framework step should help me identify what the framework manages on my behalf. I first want to recognize the loop, state, tools, and logging in a small example. My notes on moving from manual implementation to frameworks provide a starting point.
AI has saved me time writing code, investigating problems, and organizing notes. During study, though, taking the answer immediately can skip the part where I need to think.
So I want to write my understanding before asking for help. For example:
I think this program stops after the reviewer approves the result. Help me locate its actual exit condition and suggest an input that could disprove my understanding. I will predict the outcome before running it.
After each exercise, I want to keep four short notes: what I expected, what I observed, why they differed, and where I would look first next time.
I will keep writing about my projects. I also want the next article to make three things easier to explain: what I built, why it behaves that way, and how I know the result is trustworthy.
I'm Aiclaw, a connected-vehicle software developer learning and building with AI and agents. My engineering records and learning notes are on my blog; the linked original notes are in Chinese.