Agents Need Observability, Not Just Context A developer working in embedded systems argues that AI coding agents need observability tools—debug probes, GDB sessions, serial logs—rather than just better prompting, drawing an analogy to reinforcement learning's verifiable rewards. The engineer contends that agents default to reading code and guessing instead of stepping through firmware or inspecting registers, and that developers must build the verification loop at inference time by giving agents the means to observe system state directly. AI models keep getting more powerful, yet people still struggle to get the most out of them. This post isn't a lesson on the "right" way to use an LLM. It describes one possible approach to developing software with an AI model, one that has worked well for me. Traditional software development already has several methodologies, like Test-Driven Development https://martinfowler.com/bliki/TestDrivenDevelopment.html , popularized by Kent Beck. They all share the same goal: maximize software quality and keep development as smooth as possible, reducing the risk that comes with every change think regression tests . We want to be sure that what we're building is correct, and we need some way to verify it. I think LLMs require us to rethink these approaches for the tools we now have. Let's start with the tool itself. Oversimplifying quite a bit, an LLM is a text predictor: given a sequence of tokens, it predicts the next one. It's produced through what we call "training," which can be roughly split into three phases: - Pre-training: the model is exposed to a massive amount of data and learns the distribution of tokens. - Supervised fine-tuning SFT : the model learns from curated examples to follow instructions and behave like an assistant. - Reinforcement learning: the model is specialized further on different tasks. The two main flavors are Reinforcement Learning from Human Feedback RLHF https://arxiv.org/abs/2203.02155 , where the reward comes from human preferences, and Reinforcement Learning with Verifiable Rewards RLVR https://arxiv.org/abs/2411.15124 , where it comes from an automatic check. The idea is simple: if we can describe a problem so that we can always check whether the model's output is correct, we can let the model learn that skill inside this loop. In practice, a model can learn any problem framed this way, as long as it has enough capacity for that particular task. This is why we're seeing more and more models maxxing out benchmarks: a benchmark is, by definition, something you can verify. Andrej Karpathy sums it up in his post Verifiability https://karpathy.bearblog.dev/verifiability/ : "Software 2.0 easily automates what you can verify." In his framing, a verifiable task can be optimized directly or through reinforcement learning, because the AI gets to "practice" it. Some problems can't be described this way, and so they can't be "hacked." Take creative writing. There's no way for a model to know whether the book or post it wrote is "good." The definition of "good" changes over time and from person to person. Here's the key point. During training, the model learns to observe a problem and iterate toward a solution. If this loop works for updating the weights, nothing stops us from reusing it at a higher level of abstraction, through in-context learning. The weights stay fixed, but the agent can still try, observe, and correct itself within the session. So a good developer can't just be good at writing prompts that narrow down the space of solutions. They also need to give the agent tools to observe the problem: to see how its variables change and to check whether a solution actually holds. In other words, we have to build the verifiable reward ourselves, at inference time. I work mainly in embedded systems, so I know how important it is to attach a debugger to a microcontroller and look at the actual state of the system. I would never try to hunt down a bug just by reading the code. Yet many developers don't push their agents to do the same. Agents also seem pretty lazy about it at least that's my impression : they almost never propose stepping through the firmware, reading registers, or adding traces on their own. They read the code and guess. Giving the agent access to a debug probe, a GDB session, or even just the serial log turns guessing into observing. The second example is a task I worked on recently. A common way to measure a room impulse response RIR is to play a known excitation signal, such as an exponential sine sweep, and record it with a microphone. This works well when playback and recording share the same clock. But what if they run on two different devices, each with its own oscillator? Oscillator accuracy is specified in parts per million ppm , and when the playback and recording clocks differ, the result is a timing and phase error that distorts the measured impulse response. Hannes Gamper's paper Clock drift estimation and compensation for asynchronous impulse response measurements https://www.microsoft.com/en-us/research/publication/clock-drift-estimation-compensation-asynchronous-impulse-response-measurements/ presents a very effective solution: it estimates the drift between the playback and recording clocks directly from the recorded response, and uses that estimate to produce a drift-compensated IR. The catch is that the only reference implementation is a set of MATLAB scripts https://github.com/microsoft/Asynchronous impulse response measurement , and I needed it in Python. First attempt: no way to observe. I started the obvious way and asked the agent to implement the method in Python. It produced clean-looking code along with its own unit tests, and the tests passed. Then I ran it on real measurements, and the results weren't quite right. The tests weren't exactly wrong. They just checked what the agent already believed the code should do. Nothing in that loop let the agent see whether the algorithm actually recovered the drift. Second attempt: make the problem observable. So I changed approach. Before going back to real hardware, I asked the agent to build a simulated environment with pyroomacoustics https://github.com/LCAV/pyroomacoustics . This Python package lets you quickly set up rooms with sources and microphones and generates room impulse responses with a fast implementation of the image source model. In simulation, we know the ground truth: the room's true impulse response and the exact amount of drift. The agent could inject a known clock drift at different ppm values, run its estimator, and compare. Did the estimated drift match the injected one? Did the compensated RIR match the true one? That changed everything. With a measurable error to look at, the agent stopped guessing. It ran the simulation, saw where the estimate diverged, fixed the code, and ran it again. It kept iterating until the solution held across the whole range of drift values. Only then did I go back to real measurements, and this time the implementation worked. The difference between the two attempts wasn't the model or the prompt. It was the environment. In the first, the agent could only check its code against its own expectations. In the second, it could check it against reality, or at least a faithful simulation of it. Not every problem is observable, but a great many are. In control theory, a system is called observable when its internal state can be reconstructed from its outputs. That's exactly what we should aim for when working with agents: give them outputs rich enough to reconstruct what's really going on. A good software engineer today can't just be a context engineer who knows how to feed the model the right files. They also need to be an observability engineer : someone who puts the model in a position to observe the problem, so it can take it apart piece by piece. I think this is exactly where LLMs still struggle and show little initiative, and where an expert with deep vertical knowledge makes a real difference. Knowing which signal to probe, which simulation to build, and which quantity tells you whether you're right is domain knowledge, and the agent rarely reaches for it on its own. So don't despair I'm saying this to myself too : there's still room for human intelligence. References - Kent Beck, Test-Driven Development: By Example , Addison-Wesley, 2002. Overview: Martin Fowler, Test Driven Development https://martinfowler.com/bliki/TestDrivenDevelopment.html - Ouyang et al., Training language models to follow instructions with human feedback https://arxiv.org/abs/2203.02155 , 2022 - Lambert et al., Tülu 3: Pushing Frontiers in Open Language Model Post-Training https://arxiv.org/abs/2411.15124 , 2024 - DeepSeek-AI, DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning https://arxiv.org/abs/2501.12948 , 2025 - Andrej Karpathy, Verifiability https://karpathy.bearblog.dev/verifiability/ , 2025 - Hannes Gamper, Clock drift estimation and compensation for asynchronous impulse response measurements https://www.microsoft.com/en-us/research/publication/clock-drift-estimation-compensation-asynchronous-impulse-response-measurements/ , HSCMA 2017. Code: microsoft/Asynchronous impulse response measurement https://github.com/microsoft/Asynchronous impulse response measurement - Scheibler, Bezzam, Dokmanić, Pyroomacoustics: A Python package for audio room simulations and array processing algorithms https://arxiv.org/abs/1710.04196 , ICASSP 2018