2026 · Proceedings of the 34th ACM International Conference on the Foundations of Software Engineering (FSE)

Observability and Runtime Governance for Agentic AI Systems

Naser Ezzati-Jivan | Maryam Ekhlasi

Evidence basis: full-text-reviewed · Review status: catalog-reviewed; paper-author approval pending

observability llm-assisted-analysis performance-analysis

agentic AI AgentOps runtime governance AI observability software agents agent tracing silent failures goal drift tool-use failures cross-layer evidence policy checks human escalation FSE 2026

Core contribution: The tutorial presents an end-to-end AgentOps workflow that connects task intent and model decisions to tool calls, memory access, inter-agent communication, external side effects, and runtime-governance decisions.

Abstract

Agentic AI systems increasingly plan, call tools, coordinate with other agents, and act on external software services. Their failures are difficult to diagnose because control flow is generated at run-time, execution is stochastic, and incorrect behavior often appears as goal drift, unsafe tool use, or false success reports rather than as crashes. This tutorial presents a software engineering workflow for AgentOps that connects high-level intent, model-level decisions, tool calls, memory accesses, and low-level system effects. Participants will learn how to construct agent traces, correlate execution evidence, identify silent and drifting behavior, and apply runtime governance patterns such as policy checks, risk tracking, containment, and human escalation.

Source: Exact author abstract from the complete two-page CC BY 4.0 paper matching DOI 10.1145/3803437.3804904, reviewed locally on 2026-08-09.

Problem and motivation

Agentic systems synthesize stochastic control flow at runtime, so failures may appear as goal drift, unsafe tool use, incomplete work, or false success rather than exceptions. Prompt and response logs alone do not explain how decisions propagate into software and system effects.

Method and contribution

The proposed workflow has five stages: characterize the task, risks, constraints, tools, memories, dependencies, and permitted side effects; classify reasoning, planning, tool-use, coordination, memory, and governance failures; construct agent traces using spans, events, correlation identifiers, decisions, retries, memory operations, handoffs, and effects; correlate application-, model-, tool-, and system-level evidence; and apply policy checks, cumulative risk tracking, graded containment, supervisory agents, and human escalation.

Findings and evidence

This is a 90-minute technical tutorial and conceptual workflow rather than an empirical study. Its stated outcomes are a failure vocabulary, an agent-trace model, analysis strategies for loops, divergence, and silent failures, and runtime-governance patterns that support diagnosis and intervention.

Limitations and future directions

Limitations: The paper provides a tutorial architecture and representative forensic examples, but it does not report a controlled evaluation, benchmark, dataset, or quantitative comparison. Effectiveness and overhead therefore remain unmeasured in this publication.

Future work: The paper identifies standardization gaps, scalability limits, evidence-sufficiency questions, and the need to improve cross-layer correlation and risk-aware intervention for long-running and partially successful agent workflows.

Sources and identifiers

When to cite this paper

Cite this paper when designing cross-layer AgentOps observability or runtime governance for tool-using and multi-agent systems.

Citation

BibTeX
@inproceedings{ezzatiJivan2026observabilityand,
  author = {Naser Ezzati-Jivan and Maryam Ekhlasi},
  title = {Observability and Runtime Governance for Agentic AI Systems},
  year = {2026},
  booktitle = {Proceedings of the 34th ACM International Conference on the Foundations of Software Engineering (FSE)},
  pages = {64-65},
  publisher = {ACM},
  doi = {10.1145/3803437.3804904},
  url = {https://doi.org/10.1145/3803437.3804904}
}
Other citation formats for Word and reference managers
APA 7
Ezzati-Jivan, N., & Ekhlasi, M. (2026). Observability and Runtime Governance for Agentic AI Systems. In Proceedings of the 34th ACM International Conference on the Foundations of Software Engineering (FSE) (pp. 64-65). https://doi.org/10.1145/3803437.3804904
IEEE
N. Ezzati-Jivan and M. Ekhlasi, "Observability and Runtime Governance for Agentic AI Systems," in Proceedings of the 34th ACM International Conference on the Foundations of Software Engineering (FSE), pp. 64-65, 2026, doi: 10.1145/3803437.3804904

Readable Markdown record · JSON record · Download RIS