Developer Tools & Techniques

Tracing Agent Harness Behavior with NVIDIA NeMo Relay

Learn how to use execution traces to understand agent behavior and determine whether harness changes improve task outcomes.

Illustration of a developer at code screens, with a camera and Nous logo beside task success panels showing 70% and 81%.

AI-Generated Summary

  • NVIDIA NeMo Relay provides a common observability layer that captures ordered lifecycle events and structured trajectories for agent runs.
  • Hermes Agent includes native NeMo Relay integration, producing ATOF event streams, ATIF step-by-step trajectories, and OpenTelemetry spans with OpenInference labels for inspection in tools like Arize Phoenix.
  • A terminal-tool task demonstrates the setup: Hermes runs a script in an isolated Docker container, and the runner verifies the exact output VALUE=42 alongside completed model scopes, token usage, and zero tool errors.
  • A multi-tool research task requires Hermes to read a travel record, search the web, verify a conference on its official site, write a report, and return the name COLT 2026, with all model and tool spans sent to Phoenix.
  • NeMo Relay traces enable controlled harness evaluation: a Hermes ToolPerf case study compared baseline and fixed revisions across 108 runs, revealing that Qwen Coder 30B recovered more tasks but increased calls, data, and latency.

Next Steps

Powered by NVIDIA Nemotron. AI-generated content may summarize information incompletely. Verify important information. Learn more

An agent can finish a task and still take an inefficient path. A failed search can trigger another search. A truncated file read can lead to a command fetching the same content again. A correct final answer hides those extra steps, even though they increase latency and consume tokens. Inefficiencies create more chances for failure.

To improve an agent’s behavior, developers must understand whether a task succeeded and how the agent completed it. A success check by itself cannot explain why an agent recovered from a tool error, stopped early, or needed extra model calls.

In this tutorial, you’ll run two Hermes Agent examples with NVIDIA NeMo Relay. You’ll use the resulting traces to inspect model and tool calls, errors, retries, duration, and token use, then compare that evidence with each task’s verification result. A Hermes ToolPerf case study shows how to use the same approach to evaluate harness changes across repeated runs.

This tutorial and video explain how to:

  • Set up an isolated Hermes Agent runtime with its native NeMo Relay integration.
  • Run a simple terminal tool task and inspect its event stream and trajectory.
  • Run a file-and-web research task and explore its OpenTelemetry trace in Arize Phoenix.
  • Combine task verification with trace evidence to evaluate a change to an agent harness.
Video 1. A step-by-step code walkthrough for evaluating Hermes Agent traces with NeMo Relay

Prerequisites

Before you begin, make sure you have:

Next is an overview of the technologies used for this tutorial and how they work together.

How NeMo Relay works with Hermes Agent harness

NeMo Relay gives agent developers a common way to observe and control model and tool execution. The popular agent harness Hermes Agent includes NeMo Relay natively and represents its sessions, turns, model calls, and tool calls in NeMo Relay’s scope hierarchy. NeMo Relay records lifecycle events as work begins and ends, preserving its timing and parent-child relationships.

Understand the trace outputs

NeMo Relay is used for agent observability. You will work with three representations of agent execution:

FormatWhat it containsWhen to use it
Agent Trajectory Observability Format (ATOF)A JSONL log of scope starts, scope ends, and point-in-time mark, with IDs and timestamps to reconstruct the agent run.Use ATOF to debug or audit individual events, timing, and parent-child relationships.
Agent Trajectory Interchange Format (ATIF)A step-by-step JSON record of agent interactions, tool calls, and observations, assembled from lifecycle events.Use ATIF to review, analyze, or evaluate the agent’s path step by step.
OpenTelemetry with OpenInferenceOpenTelemetry records the run as parent-child spans. OpenInference labels agent, LLM, and tool spans and defines their attributes.Use it in OTEL-compliant tools like Phoenix, to inspect model and tool calls, duration, token use, and errors.
Table 1. Trace outputs are represented in three different formats to accommodate varied agent workflows

An ATIF tool request shows what the model asked to run, but it does not confirm the outcome. To verify what happened, inspect ATOF for the matching tool start and end events and any recorded errors. Their shared uuid pairs the events, while parent_uuid connects the tool call to its parent.

Review traces before sharing them. Depending on your configuration, they can contain prompts, model responses, tool arguments and results, file paths, and other application data.

For agent safety and security governance, NeMo Relay helps provide the evidence layer: structured traces and trajectories that enterprises, evaluators, and security systems can use to investigate agent behavior, evaluate policies, improve controls, or create specialized security plugins that extend Relay.

Let’s get started with the first agent task run.

Experiment #1: Run a simple tool-use task with Hermes Agent

The first example is intentionally small so you can verify the complete setup before adding web search and Phoenix. Hermes uses its terminal tool to run the included Python script inside an isolated Docker container. The script prints: VALUE=42.

That fixed output gives the runner an exact success check. A passing run also confirms that Hermes reached the model, invoked the terminal tool in the sandbox, and produced both Relay trace files.

The container cannot access the network, repository checkout, or NVIDIA API key. Hermes also cannot fall back to running terminal commands on the host.

Run the following commands in order. After copying keys.env, add your NVIDIA API key to that file before continuing.

# Clone the tutorial repository.
git clone https://github.com/NVIDIA/nemoclaw-community

# Enter the cloned repository.
cd nemoclaw-community/examples/tools/hermes-relay-tracing

# Create the isolated Hermes Agent and NeMo Relay runtime.
./scripts/setup_tutorial_runtime.sh

# Copy the API key template.
cp keys.env.example keys.env

# Add NVIDIA_API_KEY to keys.env before continuing.

# Verify that Docker is running.
docker version

# Build the Docker image for the terminal-tool task.
./scripts/build_tutorial_image.sh

# Run the task and generate the ATOF and ATIF traces.
./scripts/run_tutorial.sh

The Hermes Agent and NeMo Relay processes run from this repository’s local environment. Docker is used separately for the terminal tool sandbox and the local Phoenix service.

The setup script in the repository creates a self-contained environment under .tutorial-runtime/ with all the appropriate dependencies like Python 3.11 and Hermes 0.21.1 with NeMo Relay 0.8.3. It does not modify your existing Python or Hermes installation.

When the task finishes, the runner checks the response and both trace files. A passing run prints the verification result, followed by the ATOF and ATIF summaries.

Review the trace summaries

After Hermes completes the task, the runner verifies the expected result and the generated traces. It checks that the terminal command succeeded, that the ATOF trace contains completed LLM activity with token usage and no tool errors, and that a nonempty ATIF trajectory was created. The following output comes from one verified run. Token counts, identifiers, and file paths can vary between runs.

ATOF summary

events: 74
completed llm scopes: 2
llm scopes with usage: 2
prompt tokens: 7239
completion tokens: 96
total tokens: 7335
tool calls: 1
tool errors: 0
correlated events: 74

ATIF summary

agent: Hermes Agent
model: nvidia/nemotron-3.5-lightning-30b-a3b
steps: 3
llm calls: 2
requested tool calls: 1

Task verified: VALUE=42
Artifacts: .../artifacts/runs/<run-id>

Look for the line with Task verified: VALUE=42, which confirms the expected result. The ATOF summary reports the completed model scopes, token usage, tool calls, and tool errors. The ATIF summary presents the same execution as a three-step trajectory.

The path following Artifacts: identifies the run directory containing the complete ATOF event stream and ATIF trajectory. The companion repository explains how to inspect or summarize either file again.

Experiment #2: Run a multi-tool research task and explore its traces

After verifying the basic setup, this second example uses the same Hermes and NeMo Relay environment for a task that requires several tools. Hermes receives a travel record containing clues about an unnamed machine-learning conference. It must read the record, find a conference matching the subject, dates, and location, confirm the answer on the official website, save the verified information in a report, and return the conference name.

Use NeMo Relay’s OpenInference exporter to send OpenTelemetry spans to Arize Phoenix over OTLP (OpenTelemetry Protocol). Phoenix displays the run as an interactive trace, where you can inspect model and tool calls, timing, token usage, errors, and available inputs and outputs.

You can send the same OpenTelemetry trace to other OTLP-compatible backends, such as LangSmith, by changing the endpoint and authentication settings. The NeMo Relay observability guide describes other available exporters and configuration options.

The example reuses the NVIDIA Nemotron model and NVIDIA API key from the first example. Hermes uses its built-in keyless web search, and the runner starts Phoenix in a pinned local container. Run the conference search example:

./scripts/run_conference_research_with_phoenix.sh

During the run, NeMo Relay saves the ATOF event stream and ATIF trajectory locally.

Before reporting success, the runner checks that Hermes:

  • Identified COLT 2026 as a conference name
  • Saved a report containing the expected conference details and official source
  • Successfully completed the read_file, web_search, web_extract, and write_file calls
  • Produced a nonempty ATIF trajectory
  • Sent model and tool spans with positive token usage to Phoenix

After the checks pass, the terminal output includes a link to the Phoenix project and the local run directory. Open the project in Phoenix to follow the agent run from the initial file read through the web search, source verification, report write, and final response.

Try the task with another model

To see how another model handles the same conference query, follow the model-profile instructions in the companion repository. Keep the query, available tools, execution limits, and verifier unchanged so you can inspect how the execution path differs.

Because the task uses live web search, use these runs to explore behavior rather than rank models. A controlled comparison requires fixed search responses and repeated runs.

Use traces to evaluate an agent harness change

The two examples show how to verify a result and inspect one run. Evaluating a harness change requires the same checks under controlled, repeated conditions.

Steps for agent harness evaluation

  • Choose a fixed task with an exact, automated success check.
  • Define a baseline and one focused change to the prompt, tool, configuration, or harness.
  • Keep everything except that change constant, including the model snapshot, provider, task input, execution budget, and timeout.
  • Run the same number of repetitions for the baseline and candidate with NeMo Relay enabled.
  • Compare verified task outcomes first. Then use the traces to examine model calls, tool calls, retries, errors, elapsed time, token usage, and cost.
  • Repeat the evaluation across the models or workloads that the change is expected to support before generalizing the result.

Call the candidate an improvement only when it produces a repeatable increase in task completion or preserves completion while improving the reliability, latency, or cost measure you intended to change. One faster run or fewer calls can help explain a result, but neither establishes an optimization by itself.

An unchanged result is useful as well. It can show that an apparent improvement depended on a particular model, environment, or sample.

Benchmark case study for evaluating Hermes tool-layer changes

The Hermes ToolPerf benchmark was developed by Nous Research after analyzing production sessions, auditing tool schemas, and mining production session logs for failure classes. Then, NeMo Relay ATOF traces were used for ground-truth turn accounting. By examining nine failure patterns, the results were turned into deterministic benchmark cases and then used to evaluate a batch of Hermes tool-layer fixes.

The August 6 rerun compared the pinned baseline of Hermes and fixed revisions across nine tasks. Each task ran three times per model per arm, giving 108 runs total, with the same prompts, tools, execution limits, and success checks throughout. A task verifier measured completion, and the NeMo Relay ATOF traces captured model calls, tool calls, errors, retries, tool-result data, and timing.

ModelArmTask successMean LLM callsMean tool callsMean tool result dataMean duration
Claude Sonnet 4.5Baseline24/27 (89%)2.92.217 KB16 s
Claude Sonnet 4.5Fixes23/27 (85%)2.82.117 KB22 s
Qwen3 Coder 30BBaseline19/27 (70%)3.82.816 KB27 s
Qwen3 Coder 30BFixes22/27 (81%)4.93.933 KB42 s
Table 2. Results from the August 6, 2026 A/B run comparing pinned baseline and fixes revisions across two models

In this run, Sonnet showed no real change. It finished 24 of 27 runs on baseline and 23 of 27 on fixes, a difference of one run. Its turns, tool calls, and result data were the same on both arms.

Qwen Coder is where the fixes did something, and the effect cuts two ways. It completed three more tasks successfully, going from 19 of 27 to 22 of 27. It also worked harder for them: mean LLM calls rose from 3.8 to 4.9, tool calls from 2.8 to 3.9, tool-result data from 16 KB to 33 KB, and duration from 27 s to 42 s. The fixes brought success to tasks the baseline abandoned, and the price was a slower, chattier agent.

The task-level audit shows where those trade-offs come from.

  • On the blocked-command task, baseline Qwen died at the parser block and scored 33%, while the recovery recipes took the fixes arm to 100%. That completion is the reason the turn count went up: recovering costs turns that giving up never spends.
  • The case-insensitive search task moved the other way. The zero-match probe output pushed Qwen into extra exploratory searches on two of three reps, taking it from 3.3 turns to 9.3, a regression worth its own look.
  • The hidden-file search stayed at 0 to 33% on both arms and both models, the same gap the original run found and still open at these SHAs.

The success or failure of tasks alone could not have produced this conclusion. The NeMo Relay traces recorded every model call, tool call, error, retry, result payload, and timing for all 108 runs. Qwen’s extra turns were recoveries rather than flailing, and that one task’s regression traced back to a specific probe output. The traces are checked into the results directory, so anyone can unpack them and regenerate these tables from the raw records.

Get started with evaluating agent traces

NeMo Relay gives you a consistent way to capture evidence for evaluating changes to optimize your agent harnesses. ATOF preserves the ordered lifecycle events, while ATIF presents the same work as a readable trajectory. Pairing these traces with a deterministic verifier enables you to compare harness changes without confusing fewer calls with better results.

The companion repository provides the sample artifacts, including the runnable tasks, verifiers, NeMo Relay configuration, and Phoenix setup used in this post.

Learn more:

Discuss (0)

Tags