Tracing Agent Harness Behavior with NVIDIA NeMo Relay NVIDIA published a tutorial showing developers how to trace Hermes Agent behavior with NVIDIA NeMo Relay, running two Hermes Agent examples and inspecting the resulting traces for model and tool calls, errors, retries, duration, and token use. The tutorial covers setting up an isolated Hermes Agent runtime with its native NeMo Relay integration, running a terminal tool task and a file-and-web research task, and viewing the OpenTelemetry trace in Arize Phoenix, using NVIDIA Nemotron 3.5 Lightning via an NVIDIA Build API key. A Hermes ToolPerf case study demonstrates the same approach for evaluating harness changes across repeated runs. An agent https://www.nvidia.com/en-us/glossary/ai-agents/ can finish a task and still take an inefficient path. A failed search can trigger another search. A truncated file read can lead to a command fetching the same content again. A correct final answer hides those extra steps, even though they increase latency and consume tokens https://blogs.nvidia.com/blog/ai-tokens-explained/ . Inefficiencies create more chances for failure. To improve an agent’s behavior, developers must understand whether a task succeeded and how the agent completed it. A success check by itself cannot explain why an agent recovered from a tool error, stopped early, or needed extra model calls. In this tutorial, you’ll run two Hermes Agent examples with NVIDIA NeMo Relay https://docs.nvidia.com/nemo/relay/latest/about-nemo-relay/overview . You’ll use the resulting traces to inspect model and tool calls, errors, retries, duration, and token use, then compare that evidence with each task’s verification result. A Hermes ToolPerf https://github.com/NousResearch/hermes-toolperf-evals case study shows how to use the same approach to evaluate harness changes across repeated runs. This tutorial and video explain how to: - Set up an isolated Hermes Agent runtime with its native NeMo Relay integration. - Run a simple terminal tool task and inspect its event stream and trajectory. - Run a file-and-web research task and explore its OpenTelemetry trace in Arize Phoenix https://arize.com/phoenix/ . - Combine task verification with trace evidence to evaluate a change to an agent harness. Prerequisites Before you begin, make sure you have: - macOS or Linux - Git https://git-scm.com/downloads and curl https://curl.se/download.html - Docker Desktop or Docker Engine https://docs.docker.com/get-started/get-docker/ installed and running - An NVIDIA Build API key for NVIDIA Nemotron 3.5 Lightning https://build.nvidia.com/nvidia/nemotron-3.5-lightning-30b-a3b . Open the model page and select Generate API Key . Next is an overview of the technologies used for this tutorial and how they work together. How NeMo Relay works with Hermes Agent harness NeMo Relay https://docs.nvidia.com/nemo/relay/latest/getting-started/about gives agent developers a common way to observe and control model and tool execution. The popular agent harness Hermes Agent https://hermes-agent.nousresearch.com/ includes NeMo Relay natively and represents its sessions, turns, model calls, and tool calls in NeMo Relay’s scope hierarchy. NeMo Relay records lifecycle events as work begins and ends, preserving its timing and parent-child relationships. Understand the trace outputs NeMo Relay is used for agent observability. You will work with three representations of agent execution: | Format | What it contains | When to use it | |---|---|---| | Agent Trajectory Observability Format ATOF https://docs.nvidia.com/nemo/relay/latest/configure-plugins/observability/atof | A JSONL log of scope starts, scope ends, and point-in-time mark, with IDs and timestamps to reconstruct the agent run. | Use ATOF to debug or audit individual events, timing, and parent-child relationships. | | Agent Trajectory Interchange Format ATIF https://docs.nvidia.com/nemo/relay/latest/configure-plugins/observability/atif | A step-by-step JSON record of agent interactions, tool calls, and observations, assembled from lifecycle events. | Use ATIF to review, analyze, or evaluate the agent’s path step by step. | | OpenTelemetry https://opentelemetry.io/docs/concepts/signals/traces/ with OpenInference https://github.com/Arize-ai/openinference/blob/main/spec/README.md | OpenTelemetry records the run as parent-child spans. OpenInference labels agent, LLM https://www.nvidia.com/en-us/glossary/large-language-models/ , and tool spans and defines their attributes. | Use it in OTEL-compliant tools like Phoenix, to inspect model and tool calls, duration, token use, and errors. | Table 1. Trace outputs are represented in three different formats to accommodate varied agent workflows An ATIF tool request shows what the model asked to run, but it does not confirm the outcome. To verify what happened, inspect ATOF for the matching tool start and end events and any recorded errors. Their shared uuid pairs the events, while parent uuid connects the tool call to its parent. Review traces before sharing them. Depending on your configuration, they can contain prompts, model responses, tool arguments and results, file paths, and other application data. For agent safety and security governance, NeMo Relay helps provide the evidence layer: structured traces and trajectories that enterprises, evaluators, and security systems can use to investigate agent behavior, evaluate policies, improve controls, or create specialized security plugins that extend Relay. Let’s get started with the first agent task run. Experiment 1: Run a simple tool-use task with Hermes Agent The first example is intentionally small so you can verify the complete setup before adding web search and Phoenix. Hermes uses its terminal tool to run the included Python script inside an isolated Docker container. The script prints: VALUE=42. That fixed output gives the runner an exact success check. A passing run also confirms that Hermes reached the model, invoked the terminal tool in the sandbox, and produced both Relay trace files. The container cannot access the network, repository checkout, or NVIDIA API key. Hermes also cannot fall back to running terminal commands on the host. Run the following commands in order. After copying keys.env , add your NVIDIA API key to that file before continuing. Clone the tutorial repository. git clone https://github.com/NVIDIA/nemoclaw-community Enter the cloned repository. cd nemoclaw-community/examples/tools/hermes-relay-tracing Create the isolated Hermes Agent and NeMo Relay runtime. ./scripts/setup tutorial runtime.sh Copy the API key template. cp keys.env.example keys.env Add NVIDIA API KEY to keys.env before continuing. Verify that Docker is running. docker version Build the Docker image for the terminal-tool task. ./scripts/build tutorial image.sh Run the task and generate the ATOF and ATIF traces. ./scripts/run tutorial.sh The Hermes Agent and NeMo Relay processes run from this repository’s local environment. Docker is used separately for the terminal tool sandbox and the local Phoenix service. The setup script in the repository creates a self-contained environment under .tutorial-runtime/ with all the appropriate dependencies like Python 3.11 and Hermes 0.21.1 with NeMo Relay 0.8.3. It does not modify your existing Python or Hermes installation. When the task finishes, the runner checks the response and both trace files. A passing run prints the verification result, followed by the ATOF and ATIF summaries. Review the trace summaries After Hermes completes the task, the runner verifies the expected result and the generated traces. It checks that the terminal command succeeded, that the ATOF trace contains completed LLM activity with token usage and no tool errors, and that a nonempty ATIF trajectory was created. The following output comes from one verified run. Token counts, identifiers, and file paths can vary between runs. ATOF summary events: 74 completed llm scopes: 2 llm scopes with usage: 2 prompt tokens: 7239 completion tokens: 96 total tokens: 7335 tool calls: 1 tool errors: 0 correlated events: 74 ATIF summary agent: Hermes Agent model: nvidia/nemotron-3.5-lightning-30b-a3b steps: 3 llm calls: 2 requested tool calls: 1 Task verified: VALUE=42 Artifacts: .../artifacts/runs/