Observability, evaluation, and evidence for governing autonomous agents at the edge
SGAEIA Research Series — Article 18
Aridio Silva
Independent Researcher, Brazil
Creator of SGAEIA — Secure Governed Autonomous Edge Intelligence Architecture
ORCID: 0009-0008-2411-6995
Copyright: © 2026 Aridio Silva | License: CC BY 4.0

Abstract
As of October 2026, the major AI-agent platforms can record the individual steps of an execution, yet the ecosystem still lacks stable and uniform semantics. This paper combines three contributions: a conceptual framework for agent-observability questions; a comparative review of 13 platforms, with package versions checked on 10 October 2026; and a non-normative interpretation of the implications for SGAEIA in distributed multi-agent and edge environments. The results show that a trace reconstructs the path of execution but does not demonstrate the correctness of the outcome; that essential capabilities remain in beta, preview, or Development status; and that divergent content-capture defaults create direct privacy, governance, and evidence-continuity risks. The central conclusion is that tracing, evaluation, and governed evidence must operate as complementary functions: observing execution, judging outcomes, and preserving a verifiable basis for audit and assurance.
Keywords: AI; AI agents; observability; distributed tracing; OpenTelemetry; spans; agent evaluation; multi-agent systems; edge AI; SGAEIA.
1. Introduction
When an AI agent is activated, operators need to know what it actually did: whether execution ended normally, where errors or deviations occurred, how long and how much it cost, which tools were invoked, and which agents received control. Agents choose tools and delegation paths at runtime, so their behavior is non-deterministic and cannot be reconstructed reliably from flat application logs alone. Engineering practice therefore adapts distributed tracing from microservice systems to the agent lifecycle.
This paper addresses four research questions:
- RQ1. How do traces and spans apply to the execution of one or more AI agents?
- RQ2. Which signals answer questions about completion, error, deviation, duration, cost, and executed steps?
- RQ3. Which platforms offer these capabilities, at what level of maturity, and with which strengths and limitations?
- RQ4. How do tracing, evaluation, and evidence preservation relate to governance of distributed edge agents from the conceptual perspective of SGAEIA?
The scope covers agent runtimes from Anthropic, OpenAI, Google, Microsoft, and AWS; seven independent observability platforms; and the OpenTelemetry standard. The language model alone does not create this operational record. Tracing is produced by the execution layer: the SDK, framework, orchestration environment, or hosting platform surrounding the model.
2. Foundations: trace, span, and session
A span is the basic unit of traced work. Each relevant operation can produce its own span, normally including start and end times, parent relationship, status, and operation-specific attributes. Common examples include a model call, a tool invocation, and a handoff or nested agent invocation. A trace groups causally related spans into the end-to-end journey of an operation, while a session or thread may group multiple traces that belong to the same conversation or longer-running task.
In the OpenAI Agents SDK, spans expose fields such as started_at, ended_at, parent_id, and step-specific data [R4]. AWS AgentCore similarly records status and events [R15]. The vocabulary is not fully uniform: one platform may define a trace as a complete operation, another as one request-response cycle, and a third may use a session to provide the longer boundary.
2.1 Three nuances that change trace interpretation
First, a trace is not always the entire task. OpenAI can connect several executions through a group_id [R4], whereas AgentCore treats a trace as a request-response cycle and groups traces into a session [R15]. A reviewer must therefore identify the platform's execution boundary before comparing duration, cost, or completion.
Second, a handoff is not always represented by a dedicated span. OpenAI defines a handoff span [R4], while OpenTelemetry defines create_agent, invoke_agent, and execute_tool, allowing a subagent to appear as a nested invoke_agent operation [R18, R19]. Microsoft has also documented semantics for communication between agents [R11, R44].
Third, agent traces may include more than model, tool, and handoff spans. Guardrail checks, turns, permission decisions, and time spent awaiting human approval can be observable operations. The Claude Agent SDK, for example, can record a child span for a tool waiting on user authorization [R1].
2.2 Tree structure
Spans form a tree rather than a flat list. Figure 1 represents the transition from a session to traces and nested spans, including model calls, tools, and subagents. The hierarchy supports latency analysis, causal reconstruction, and failure localization without reducing the execution to an undifferentiated stream of messages.

Figure 1 — From a session to verifiable spans. A session can group multiple traces; each trace organizes agent, model, tool, and handoff spans into a causal structure. © 2026 Aridio Silva | Project SGAEIA | CC BY 4.0.
3. Observability questions and their signals
The main analytical distinction is between technical completion and substantive success. A run can terminate without an exception while still producing an incorrect answer, violating an instruction, or acting outside the intended scope. A trace reports operational facts; a separate evaluator, test, schema validator, or human review must judge the result.
| Question | Signal to inspect | Interpretation limit |
|---|---|---|
| Did the task complete? | Final run status plus independent outcome validation | No exception proves only that execution did not crash |
| Were there errors or deviations? | Error spans, failed tools, exceeded limits, denied permissions, guardrail events | Deviation requires comparison with expected behavior |
| When did it finish and how long did it take? | Run and span duration; tokens and derived cost | Boundaries vary by platform |
| What happened at each step? | Hierarchical trace with tool, arguments, and result | Content capture may be disabled or redacted |
| Who requested and authorized the action? | End-user attributes and permission-decision events | Identity does not itself establish authority |
| Was there a loop or waste? | Turn count, tool calls, and inference calls | Thresholds depend on task and policy |
3.1 Completion and outcome quality
The Claude Agent SDK result can report termination subtype, error state, number of turns, duration, and cost, including outcomes such as success, maximum-turn termination, and execution error [R36]. These fields describe how execution stopped; they do not establish that the requested objective was achieved. Semantic failures can therefore remain invisible to ordinary error status [R37]. OpenAI's trace grading illustrates the complementary approach of evaluating traces and outcomes at scale [R5].
3.2 Errors, deviations, and auditability
Errors appear in span status, API error events, and failed tool results. Deviations require broader evidence: permission events, guardrail decisions, expected-path comparison, and independent observation of resulting state. Claude Code permission decisions can be exported as structured events associated with the end user, providing input to an audit or SIEM workflow [R2].
4. Method, sources, and limitations
The review was conducted on 10 October 2026 using three source categories: package-registry data from PyPI and npm [R40]; official vendor and platform documentation; and secondary technical sources only where official documentation did not cover the relevant point. Secondary and competitor-produced comparisons are identified and treated cautiously. Package versions refer to the stable release visible in the registry on the consultation date.
In the comparison matrix, ✔ indicates a capability confirmed in the consulted documentation, ◐ indicates partial or conditional support, and — means that the capability was not found in the reviewed material. Absence from the matrix is not proof that a feature does not exist. SaaS platforms do not have a single platform version, so the version column reports the associated SDK when applicable.
This is a structured comparative review, not a systematic review or experimental benchmark. The work does not measure instrumentation overhead, reproduce vendor evaluations, or establish production effectiveness. Platform behavior, prices, retention policies, and maturity labels can change after the stated date; adoption decisions must revalidate the relevant documentation.
5. Open standard: OpenTelemetry GenAI
OpenTelemetry GenAI semantic conventions are the leading candidate for a common language across runtimes and backends, but they remain under active development [R18–R20]. The defined operations include create_agent, invoke_agent, and execute_tool; attributes cover agent and tool identity; and metrics include agent-run duration, tool-call count, and inference-call count [R18, R19]. Development status means that names and behavior can still change between versions.
The practical value is portability. Datadog accepts traces compatible with OpenTelemetry GenAI 1.37 or later [R28], MLflow can ingest and export the conventions [R33], and Google ADK implements them natively [R7]. Portability is nevertheless incomplete because vendor SDKs may emit proprietary span types, require a distribution such as ADOT, or depend on third-party bridges.
6. Vendor platforms
6.1 Anthropic: Claude Agent SDK, Claude Code, and Managed Agents
The Claude Agent SDK runs Claude Code as a child process and can export spans for interactions, model calls, tools, tokens, cost, permission decisions, and nested subagents [R1, R2]. Its main strength is OTLP export without a mandatory backend and content-disabled-by-default behavior for read/write payloads. Limitations include beta span names, silent export failures by default, and extra beta configuration for hook spans; Managed Agents also uses a beta API header [R1, R3].
6.2 OpenAI Agents SDK
OpenAI tracing is enabled by default and includes run, task, turn, agent, generation, tool, guardrail, handoff, and audio spans [R4]. The platform connects traces through group_id and integrates trace grading and agent evaluations [R5, R6]. Its strengths are low-friction instrumentation and rich agent-specific semantics; limitations include incompatibility with Zero Data Retention tracing, default capture of inputs and outputs, batch export behavior, and dependence on third-party tooling for an OpenTelemetry bridge [R4, R26].
6.3 Google ADK and Vertex AI Agent Engine
Google ADK implements OpenTelemetry GenAI semantics, with an agent execution as the root span and child spans for model and tool operations [R7, R8]. Agent Engine sends traces to Cloud Trace and exposes response times and executed operations [R9]. The design benefits from native OpenTelemetry support and managed visualization, while Cloud Logging size limits and the lack of a documented integrated trace-evaluation workflow remain material constraints [R9, R10].
6.4 Microsoft Foundry and Agent Framework
Microsoft provides generally available tracing for prompt and hosted agents and preview support for workflows and external agents, storing telemetry in Application Insights [R11–R14]. Its multi-agent semantics and Azure Monitor integration are notable strengths. Costs, retention, access-role requirements, and differing maturity labels across documentation require careful operational review; the evaluation-related comparison includes an explicitly identified third-party source [R39].
6.5 AWS Bedrock AgentCore
AgentCore organizes observability around sessions, traces, and spans, including tool inputs, outputs, timing, and errors [R15]. Unified per-agent log groups can combine spans, prompts, and logs under IAM and customer-managed encryption controls [R17]. Agent spans require ADOT instrumentation, CloudWatch Transaction Search must be configured, and the reviewed documentation did not establish integrated outcome-quality evaluation [R15, R16].
7. Independent platforms
7.1 LangSmith
LangSmith treats runs as spans and groups them into projects and threads. It supports OpenTelemetry ingestion, distributed tracing, sampling, sensitive-data controls, token cost, dashboards, and online evaluation [R21–R23]. Relevant limitations include per-trace run limits, Python-only integration for the OpenAI Agents SDK, and enterprise conditions around self-hosting reported by a third party [R41].
7.2 Langfuse
Langfuse offers an MIT-licensed core, self-hosting, nested observations, sessions, agent graphs, human annotation queues, and OTLP-oriented ingestion [R24, R25, R45]. Open source removes license fees but not the operational cost of running and securing the platform [R38].
7.3 Arize Phoenix and AX
Phoenix uses OpenTelemetry and OpenInference semantics for LLM, Tool, Agent, and Retriever spans, and includes evaluations, datasets, experiments, and prompt management [R26, R27]. The Elastic License 2.0 permits broad internal use but limits offering the software as a managed service, while some TypeScript evaluation components remain early-stage [R26, R46].
7.4 Datadog Agent Observability
Datadog accepts OpenTelemetry GenAI traces and uses span links to build an execution graph for multi-agent pipelines [R28, R29]. It monitors latency, error rate, cost, and decisions with automatic instrumentation for selected frameworks. Cross-trace links may be stored without being drawn, and the offering remains SaaS-only in the reviewed scope.
7.5 Braintrust
Braintrust combines production logs with offline and online evaluation and integrates with Claude, OpenAI, and Google agent SDKs [R30–R32]. Its self-hosting model is hybrid and its documentation describes telemetry sent to the vendor control plane, which is relevant to data-governance analysis. A third-party vendor profile used only as supplementary context is identified separately [R42].
7.6 MLflow Tracing
MLflow provides an open-source tracing layer compatible with OpenTelemetry, including ingestion and export of GenAI conventions and more than 30 automatic integrations [R33, R34]. Autologging must be enabled explicitly for each integration, which makes configuration discipline part of the assurance boundary.
7.7 Pydantic Logfire
Logfire builds on OpenTelemetry with Python and JavaScript SDKs and integrates with Pydantic AI [R35, R43, R47]. Evaluation and prompt-management claims were not fully confirmed on an opened official page in the underlying review, so evaluation remains marked as conditional. Self-hosting is limited to enterprise arrangements [R43].
8. Comparative analysis
8.1 Versions and maturity on 10 October 2026
| Platform | Version reported | Maturity note |
|---|---|---|
| Claude Agent SDK / Claude Code | py 0.2.165; npm 0.3.296; CLI 2.1.296 | Tracing beta; 0.x SDKs |
| Claude Managed Agents | API | Beta header |
| OpenAI Agents SDK | py 0.23.1; JS 0.20.0 | Tracing enabled by default; 0.x SDKs |
| Google ADK | py 2.11.0; JS 2.2.1 | 2.x; native telemetry |
| Microsoft Agent Framework | py 1.21.0; azure-ai-projects 2.8.0 | GA and preview scopes differ |
| AWS Bedrock AgentCore | py 1.24.1 | Rapidly evolving managed service |
| LangSmith | py 0.14.7; JS 0.10.11 | Commercial SaaS; 0.x SDKs |
| Langfuse | py 4.17.0; JS tracing 5.13.1 | OTLP-first v4 platform |
| Arize Phoenix | 20.20.0 | Frequent releases; ELv2 |
| Datadog | ddtrace 4.15.6 | SaaS |
| Braintrust | py 0.45.0; JS 3.37.2 | SaaS and hybrid |
| MLflow Tracing | 3.17.0 | Open source |
| Pydantic Logfire | 5.1.1 | Commercial |
| OpenTelemetry Python SDK | 1.45.1 | GenAI conventions in Development |
8.2 Capability matrix
| Platform | LLM/tool steps | Multi-agent | OTel | Sessions/threads | Trace evaluation | Hosting |
|---|---|---|---|---|---|---|
| Claude Agent SDK | ✔ beta | ✔ subagents | ✔ OTLP | ✔ | — | Backend of choice |
| Claude Managed Agents | ✔ console | — | — | ✔ | — | Managed |
| OpenAI Agents SDK | ✔ | ✔ handoffs | ◐ third party | ✔ | ✔ | OpenAI dashboard |
| Google ADK / Agent Engine | ✔ | ◐ | ✔ native | — | — | Cloud Trace |
| Microsoft Foundry / Agent Framework | ✔ | ✔ | ✔ | — | ◐ | Application Insights |
| AWS AgentCore | ✔ with ADOT | ✔ | ✔ ADOT | ✔ | — | CloudWatch |
| LangSmith | ✔ | ✔ | ✔ | ✔ | ✔ | SaaS / enterprise |
| Langfuse | ✔ | ✔ graph | ✔ | ✔ | ✔ | SaaS / MIT self-host |
| Phoenix / AX | ✔ | ✔ | ✔ | — | ✔ | ELv2 / SaaS |
| Datadog | ✔ | ✔ graph | ✔ | — | ◐ | SaaS |
| Braintrust | ✔ | ✔ | ✔ | — | ✔ | SaaS / hybrid |
| MLflow | ✔ | ✔ | ✔ | — | — | OSS / managed |
| Logfire | ✔ | ✔ | ✔ | — | ◐ | SaaS / enterprise |
The symbols reflect only the reviewed documentation. A dash means “not found in the reviewed sources,” not “the feature does not exist.”
9. Discussion: gaps, risks, and engineering practices
The ecosystem is converging toward OpenTelemetry but still exposes incompatible dialects. Claude emits claude_code.* spans, OpenAI uses a native format with third-party bridges, Google and Microsoft follow GenAI conventions more directly, and AWS depends on ADOT. Multi-vendor deployments therefore need an OpenTelemetry collector or another normalization layer, together with explicit version management for semantic changes.
Maturity is uneven. Several high-value capabilities remain beta, preview, or Development, and central SDKs still use pre-1.0 numbering. Long-lived deployments should expect schema migration and avoid binding assurance claims to unstable attribute names.
Privacy risk is architectural rather than incidental. Claude does not capture read/write content by default [R1], while OpenAI tracing captures inputs and outputs by default [R4]. Prompts, tool arguments, retrieved records, and outputs can contain personal data, secrets, or regulated information, so collection, redaction, access control, retention, and deletion rules must be decided before production.
Telemetry can fail silently. When export failure does not interrupt agent execution [R1], an absent trace can mean “nothing happened” or “evidence was lost.” Monitoring the telemetry pipeline is therefore part of observability itself. A defensible deployment should track span volume, exporter health, queue loss, clock quality, and gaps between agent activity and recorded evidence.
Finally, a trace is not an evaluation. The platforms that connect trace data to grading or evaluation cover both operational reconstruction and outcome analysis; others require an independent evaluator. Practical adoption should separate “execution ended without error” from “the result was correct and authorized,” define capture policy explicitly, monitor telemetry export, and group related traces through a stable session boundary.
10. Conceptual implications for SGAEIA in distributed edge environments
10.1 Observability is not governance
From the SGAEIA perspective, tracing is an observation and reconstruction capability, not authorization to act and not automatic proof of compliance. A trace can show that an agent invoked a tool, transferred control, or produced an output, but legitimacy depends on authority, context, and policy that cannot be inferred from telemetry alone. This preserves the distinction that technical capability, authenticated identity, and stated intent do not equal operational authority.
10.2 Evidence must survive distribution
Across edge-cloud environments, execution can cross devices, services, administrative domains, and periods of intermittent connectivity. Tracing is useful only if context propagates coherently, identifiers and time information remain sufficiently reliable, retention is proportionate to risk, and evidence loss can be distinguished from lack of activity. The conceptual requirement is not to capture everything; it is to preserve enough governed evidence to reconstruct decisions and effects without turning the observability plane into another disclosure channel.

Figure 2 — Tracing, evaluation, and evidence in SGAEIA. The trace records the path; evaluation examines outcome quality and conformity; governed evidence supports audit and assurance under explicit assumptions. This is a non-normative conceptual interpretation. © 2026 Aridio Silva | Project SGAEIA | CC BY 4.0.
10.3 Outcome correctness requires an independent signal
The most important finding for autonomous systems is that lack of an exception does not show that the objective was achieved. A run can finish cleanly while producing incorrect content, choosing an unsuitable path, or acting beyond legitimate scope. Assurance must therefore separate at least two axes: execution integrity and outcome quality or conformity, preventing one green status from hiding a semantic failure.

Figure 3 — A trace is not correctness. Technical execution status and outcome quality are independent dimensions; only the combination of tracing and evaluation distinguishes operational completion from substantive success. © 2026 Aridio Silva | Project SGAEIA | CC BY 4.0.
10.4 Research and validation requirements
This interpretation suggests research questions for the governed evolution of SGAEIA, but it does not automatically promote controls or components into normative architecture. Future work should evaluate context continuity across agents and domains, resistance of the evidence chain to suppression and tampering, offline and degraded operation, the effects of sampling and redaction on auditability, and the relation between local events and distributed trajectories. Any requirement, invariant, or mechanism derived from these findings must pass the Research-to-Architecture process and be validated before supporting an assurance statement.
11. Conclusion
Detailed tracing of agent execution is feasible across every platform reviewed, using spans for model calls, tools, handoffs, and related operations, arranged hierarchically within traces and sometimes grouped into sessions. The field is nevertheless in transition: Claude Agent SDK tracing and Managed Agents remain beta, Foundry workflows and external-agent tracing include preview scopes, OpenTelemetry GenAI conventions remain in Development, and several major SDKs retain 0.x versions. Adopters should expect migration cost and semantic change.
The research questions can be answered as follows. For RQ1, distributed-tracing concepts apply directly to agent systems, with material variation in naming and boundaries. For RQ2, trace signals answer questions about completion, errors, duration, and executed steps, but outcome correctness requires independent evaluation. For RQ3, the ecosystem is broad but uneven in portability, privacy defaults, maturity, evaluation, and hosting. For RQ4, SGAEIA should treat tracing as a governed source of evidence that complements rather than replaces authority, evaluation, validation, and assurance.
Future work includes empirical measurement of instrumentation overhead, direct product verification of capabilities marked as absent from the reviewed documentation, longitudinal monitoring of OpenTelemetry GenAI stabilization, and experiments on evidence continuity across intermittent edge environments.
Glossary
- Agent: a system that uses a model to decide at runtime which tools or agents to invoke.
- Evaluation: code-, model-, or human-based assessment of outcome quality.
- Handoff: transfer of control from one agent to another.
- OpenTelemetry (OTel): an open standard for producing and exporting traces, metrics, and logs.
- OTLP: the OpenTelemetry protocol for telemetry transport.
- Session/thread: a grouping of traces belonging to one conversation or extended task.
- Span: a timed, attributed unit of traced work with a parent relationship and status.
- Trace: the end-to-end record of an operation, composed of spans.
- Trace grading: assigning scores to agent traces to identify failures and regressions.
- ZDR: Zero Data Retention; OpenAI Agents SDK tracing is unavailable under ZDR.
Bibliography / References
All sources were consulted on 10 October 2026. Validation markers from the source study are preserved: [A] page opened in full; [B] page recovered through search and matched to the claim; [C] registry API data. Third-party sources are identified explicitly.
Vendors and standards
- R1. Anthropic. Observability with OpenTelemetry (Agent SDK). [A]
- R2. Anthropic. Monitoring (Claude Code). [A]
- R3. Anthropic. Managed Agents: observability. [B]
- R4. OpenAI. Tracing (Agents SDK). [A]
- R5. OpenAI. Trace grading. [B]
- R6. OpenAI. Evaluate agent workflows. [B]
- R7. Google. Agent activity traces (ADK). [B]
- R8. Google. Observability for agents (ADK). [B]
- R9. Google Cloud. Trace an agent (Agent Engine). [B]
- R10. Google Cloud. Observability for AI agent developers. [B]
- R11. Microsoft. Trace agent concept (Foundry). [B]
- R12. Microsoft. Set up tracing in Foundry. [B]
- R13. Microsoft. Quickstart: tracing a hosted agent. [B]
- R14. Microsoft. Enable observability for agents (Agent Framework). [B]
- R15. AWS. Observability for agentic resources in AgentCore. [A]
- R16. AWS. Add observability to AgentCore resources. [B]
- R17. AWS. Amazon Bedrock AgentCore now delivers unified observability with traces and logs in a single log group (23 July 2026). [B]
- R18. OpenTelemetry. OpenTelemetry GenAI Semantic Conventions. Official repository to which the former agent-spans documentation now points. [A]
- R19. Dash0 (third party). OpenTelemetry GenAI semantic conventions explained. [B]
- R20. DEV Community (third party). OpenTelemetry's GenAI semantic conventions are not stable yet. [B]
- R44. Microsoft. Tracing in Microsoft Foundry (earlier page mentioning Outshift). [B]
Independent platforms
- R21. LangChain. Observability concepts (LangSmith). [B]
- R22. LangChain. Trace with OpenTelemetry (LangSmith). [B]
- R23. LangChain. Observability how-to guides (LangSmith). [B]
- R24. Langfuse. Engineering clarifications. [B]
- R25. Langfuse. Data model. [B]
- R26. Arize. Phoenix repository. [B]
- R27. Atlan (third party). What is Arize. [B]
- R28. Datadog. OpenTelemetry instrumentation. [B]
- R29. Datadog. Agent monitoring. [B]
- R30. Braintrust. Self-hosting. [B]
- R31. Braintrust. Release notes. [B]
- R32. Braintrust. Glossary. [B]
- R33. MLflow. Tracing for LLM and agent observability. [B]
- R34. Databricks. Automatic tracing and integrations. [B]
- R35. Pydantic. Logfire: AI observability. [B]
- R45. Langfuse. AI SDK C++ integration. [B]
- R47. Pydantic. Pydantic AI: Logfire. [B]
Third-party sources and version data
- R36. Inference.net (third party). Claude Agent SDK tracing and evaluation in production. [B]
- R37. ClickHouse (third party). What is AI agent observability. [B]
- R38. BenchLM (third party). Best LLM observability tools. [B]
- R39. Jannik Reinhard (third party). Microsoft Foundry observability. [B]
- R40. PyPI and npm registries, accessed through API on 10 October 2026 for version and publication dates. [C]
- R41. Latitude (third party). LLM observability tools compared. [B]
- R42. RFP.wiki (third party). Braintrust. [B]
- R43. Pydantic. Logfire FAQ. [B]
- R46. Inference.net (third party). Arize Phoenix alternatives for agent observability. [B]
About the Author
Aridio Silva is an independent researcher based in Brazil working on the architecture, security, governance, and trustworthiness of autonomous and distributed artificial intelligence systems.
His research focuses on Agentic AI, Multi-Agent Systems, Edge AI, AI Security, Zero Trust, Security-by-Design, AI Governance, Spec-Driven Development, and continuous security assurance.
He is the creator and lead researcher of SGAEIA — Secure Governed Autonomous Edge Intelligence Architecture, a research initiative investigating architectural foundations for secure, governed, auditable, and trustworthy autonomous AI systems operating across distributed edge-cloud environments.
Research & Project Resources
- ORCID: https://orcid.org/0009-0008-2411-6995
- Google Scholar: https://scholar.google.com/citations?user=rPn5O48AAAAJ
- Zenodo — SGAEIA Community: https://zenodo.org/communities/sgaeia
- OpenAIRE: https://explore.openaire.eu/search/find?fv0=Aridio%20Silva&f0=q
- Medium: https://medium.com/@aridiosilva
- DEV Community: https://dev.to/aridiosilva
- GitHub: https://github.com/aridiosilva
- LinkedIn: https://www.linkedin.com/in/aridio-silva-74997111/
- Homepage: https://aridiosilva.com
- SGAEIA Homepage: https://aridiosilva.com/sgaeia
- SGAEIA LinkedIn: https://www.linkedin.com/company/sgaeia/
Figures
The cover is not numbered. Figures 1–3 are public conceptual illustrations and do not disclose private SGAEIA protocols, algorithms, state machines, policy structures, thresholds, or enforcement mechanisms.
License
Except where otherwise noted, the text and original conceptual illustrations in this article are licensed under the Creative Commons Attribution 4.0 International License (CC BY 4.0).
© 2026 Aridio Silva. You may share and adapt this work for any purpose, provided appropriate attribution is given.
The SGAEIA software and research artifact remain subject to their own Apache License 2.0.
Autonomous AI. Governed by Design. Trusted by Evidence.
Series continuity
This work is SGAEIA Research Series — Article 18. The canonical edition is openly available. The persistent record is identified by DOI 10.5281/zenodo.23284753. Distribution editions are available on Medium and DEV Community.
Suggested citation
Silva, Aridio. (2026). Tracing AI Agent Execution: Foundations, State of the Art, and Implications for Distributed Multi-Agent Systems. SGAEIA Research Series, Article 18. https://doi.org/10.5281/zenodo.23284753. Canonical edition: https://aridiosilva.com/publications/artigo18/. CC BY 4.0.