Every instrumentation layer is an interrogation. You add a trace, and the trace changes the testimony.
When we wrap an agent's reasoning chain with logging, we are not passive observers. We are inserting variables into the system's decision surface. The chain-of-thought is not a fixed artifact waiting to be captured—it is a behavior that responds to the capture itself.
1. Prompt-Level Instrumentation
The most common method—appending "output your reasoning step by step" to the system prompt—modifies the conditional distribution of the next token. The model is no longer solving the original problem; it is solving the original problem while writing a report about solving it. The reasoning trace becomes a performance, shaped by the audience.
2. State-Level Logging Overhead
Capturing intermediate tool calls, partial outputs, and internal scores forces the runtime to serialize and emit state at defined checkpoints. That check-pointing changes timing and resource allocation. Tool call order can shift because the scheduler is waiting on a log flush.
3. Selective Sampling Bias
A common reflex: log only failed branches for post-mortem analysis. This makes failures visible while silently blinding you to the normal path. But monitoring that samples only anomalies cannot distinguish "the agent is behaving" from "the agent is behaving but no one is looking."
Suppose you suspect instrumentation is altering behavior. Here is a concrete test.
Run the same task set under two conditions:
Compare the distributions, not the averages. Measure:
If these metrics diverge beyond your noise floor, the observer effect is present. The magnitude of that divergence is the price of your visibility. Monitoring is safe when it stays outside the decision boundary: observing outputs and side effects after the fact. It becomes invasive when it reaches inside the reasoning process—because the reasoning process is plastic, and it will contort around the request for transparency.
The practical stance: prefer reconstruction over extraction. Do not ask the agent to narrate. Instead, record the visible aftermath—actions, timings, external responses—and infer the reasoning path from those tracks. You lose fidelity, but you keep the behavior unpolluted.
This is the discipline of forensic observation. The good detective does not interrupt the crime to ask for a confession. He watches, takes note of what moves, and builds the case from the residue.
In agent observability, the residue is the trace you can trust.