# AI Agent Observability & Tracing Tools: evaluation worksheet

Source: https://dutygraph.com/directory/ai-governance/categories/observability/
Editorial date: 2026-09-07

## Scope

- Organization / team:
- Task and expected output:
- Human owner:
- Product and version:
- Evaluation date / environment:
- Reviewer:

## Questions

### 1. Can a run be followed across orchestration, model calls and external tools with stable correlation IDs?

- Observation (demonstrated / described / unknown):
- Evidence reference:
- Limitation or follow-up:

### 2. Which prompt, workflow, model and data versions are retained?

- Observation (demonstrated / described / unknown):
- Evidence reference:
- Limitation or follow-up:

### 3. Can missing telemetry be distinguished from a step that did not execute?

- Observation (demonstrated / described / unknown):
- Evidence reference:
- Limitation or follow-up:

### 4. How are trace access, retention, redaction and deletion configured?

- Observation (demonstrated / described / unknown):
- Evidence reference:
- Limitation or follow-up:

## Evidence checklist

- [ ] An end-to-end execution trace
- [ ] A failure investigation linked to a specific run
- [ ] Telemetry access and retention settings

## Boundary to check

A recorded trace is evidence of observed execution, not proof that an action was authorized or a result was correct. Instrumentation gaps should remain visible. Tracing does not replace independent evaluation or enforcement.

## Decision

- Fit for the scoped task:
- Unresolved gaps:
- Next action, owner and date:

This is a planning worksheet, not an endorsement, access approval or compliance certification. Keep confidential evaluation notes in your organization's approved storage.
