Book an Appointment

Agentic AI Security Testing: Evidence that makes AI agents testable

A shared trace only shows that events belong together. It does not show whether an AI agent respected approval, target, authorization and data provenance along the way. Our paper, accepted at ICSPIS 2026 (IEEE), shows which telemetry that requires. This page maps the results to the OWASP Agent Control Standard (ACS).

Last updated: October 2026 · Valeri Milke, ISO 27001 & ISO 42001 Lead Auditor

Accepted version with IEEE copyright notice

92.9%Security recall of the threat-guided profile P3 across all seven telemetry conditions, at 100% precision and 0% false positives on benign cases
16.7%Recall with purely causal correlation (P2), even though 85.7% of effects were uniquely attributed. Attribution and security assessment are separate requirements
50%Recall of P3 when all decision events are missing (F1). Approval, scope and target checks lose their comparison point
3,024Profile-fault-trace cells: 36 synthetic workflow templates × 3 deterministic repetitions × 4 evidence profiles × 7 telemetry conditions

Agentic applications connect model decisions to tools that read data, delegate work, and modify external systems. A security incident can therefore span content ingestion, planning, approval and tool invocation through to the effect at the target system. In the paper “Threat-Model-Guided Security Testing for Agentic AI: MCP Traces with MAESTRO and STRIDE”, Valeri Milke derives an Evidence Contract for Model Context Protocol workflows from MAESTRO and STRIDE. It was tested on 3,024 profile-fault-trace cells. The central result: causal trace identifiers improve attribution but do not supply the semantics required to judge authorization or provenance. Only security-relevant bindings raise recall from 16.7% to 92.9%. The OWASP Agent Control Standard (ACS) provides an open specification for this: v0.1.0, release 0.1.2, described by its project lead as a “public preview” and part of the OWASP GenAI Security Project since September 2026. It supplies hooks, dispositions, provenance and trace mappings. We show field by field how the two fit together. We also show where ACS v0.1 still has gaps and how you can build robust security tests for your agents from them.

At a glance

The results in five sentences

In short: for the security of an AI agent to be testable, telemetry must evidence every stage, from input through the policy decision and the tool call to the observed effect, with comparable facts. Correlation IDs alone are not enough. The OWASP Agent Control Standard provides hooks, provenance and trace mappings for this. Operators must add Decision Receipts and independent effect evidence.

  1. 92.9%Semantic bindings decideThe threat-guided profile P3 achieved 92.9% recall across all telemetry conditions at 100% precision and 0% false positives. Purely causal correlation reached 16.7%.
  2. 85.7%Attribution is not assessmentP2 and P3 uniquely attributed the same number of effects. Only P3 could judge whether approval, scope, provenance and target were respected.
  3. 50%Decision Receipts are criticalWithout decision events, the recall of P3 dropped by half. Reconstruction and attribution fell to 0%. In ACS, this corresponds in our interpretation to a fail-open Guardian without a decision receipt.
  4. +255%Evidence has its priceP3 produced a median of 2,341 bytes of canonical JSON per workflow versus 659 bytes for fragmented logs. Hence, introduce complete receipts for privileged operations first.
  5. 6Six families, a clear limit of the claimsThe experiment tests the evidentiary value of deterministic, MCP-shaped fixtures. It is not an LLM or prompt-injection benchmark. The next step is a real MCP SDK.

From threat model to control standard

The building blocks on which the paper and ACS rest, and what comes next. Click a milestone for details.

Core concepts from the paper and the standard

Click a card to read the concept in detail. The order follows the paper: from the research question through the Evidence Contract to cost and operations.

Four evidence profiles, and what each one actually demonstrates

The profiles build on one another. Each level adds fields. The rates apply across all seven telemetry conditions (Table II of the paper); the byte values are medians under F0.

Recall 84.5% · False positives 85.7%
  • Only component-local IDs and timestamps. Reconstruction maps an effect to the temporally nearest tool.
  • Unique attribution 0%: concurrent workflows of the same family merge into false chains.
  • The seemingly high recall is not a security achievement. Incorrectly merged chains raise alerts on 85.7% of benign cases, and precision falls to 66.4%.
no ACSTime heuristiconly duplicate counting reliable
Interactive · Trace Replay

Evidence Contract × ACS: check an agent path stage by stage

Choose an attack family from the paper, the evidence profile and, optionally, fault F1. Field names and values come from the experiment; the check rules are the same as in the evaluator. The stage at which the evidence first demonstrates the violation is marked in red.

Scenario
Evidence profile

The approval applied to mail.send with specific arguments. The same tool name is executed with substituted arguments. The argument digest differs.

pInput

Prompt or retrieved datum with origin, requested operation, scope and intended target.

trace_id
"4bf92f3577b34da6a3ce929d0e0e4736"
span_id
"00f067aa0ba902b7"
parent_span_id
null
action_id
"act-9c41e2d07a"
source_trust
"trusted"
requested_operation
"read"
requested_scope
"kb.read"
principal_id
"principal-vamisec"
memory_trust
"trusted"
memory_age_s
30
intended_target
"tenant-a/resource-17"
Counterpart in ACS v0.1
Hook / source
steps/userMessage · steps/knowledgeRetrieval · steps/memoryContextRetrieval
Fields
Provenance { origin, source_id, derived_from }
OpenTelemetry
acs.message.user · acs.knowledge.retrieval · acs.memory.retrieval + acs.provenance.origin
OCSF
6002 Application Lifecycle (ACS: Application Activity) · 6005 Datastore Activity + acs_provenance_origin
dDecision

Explicit policy decision: who approved what, with what, to where and with which capabilities?

trace_id
"4bf92f3577b34da6a3ce929d0e0e4736"
span_id
"a1b2c3d4e5f60718"
parent_span_id
"00f067aa0ba902b7"
action_id
"act-9c41e2d07a"
decision
"allow"
approval_id
"apr-7f3c2a91d0be"
approval_subject
"principal-vamisec"
approved_tool
"mail.send"
approved_args_hash
"a3f9e0…c21e"
approved_target
"recipient:security@example.org"
approved_audience
"mcp://server-a"
approved_capabilities
["mail.send"]
max_delegation_depth
1
idempotency_required
true
Counterpart in ACS v0.1
Hook / source
AcsResult (response envelope) · policy_data
Fields
decision · reasoning · reason_codes · policy_references · cited_provenance_ids · ask_details · chain_hash
OpenTelemetry
span event acs.decision (acs.decision, acs.evaluator)
OCSF
2004 Detection Finding (deny · modify · ask · defer)
tTool

MCP tool invocation that repeats the values actually exercised rather than merely pointing to the decision.

trace_id
"4bf92f3577b34da6a3ce929d0e0e4736"
span_id
"b7ad6b7169203331"
parent_span_id
"a1b2c3d4e5f60718"
action_id
"act-9c41e2d07a"
approval_id
"apr-7f3c2a91d0be"
approval_subject
"principal-vamisec"
tool_name
"mail.send"
args_hash
"9d02b7…e4f8"
target
"recipient:security@example.org"
audience
"mcp://server-a"
effective_capabilities
["mail.send"]
delegation_depth
1
idempotency_key
"idem-5be04d1c"
Counterpart in ACS v0.1
Hook / source
steps/toolCallRequest · steps/subagentStart
Fields
tool · operation · capability · arguments[].provenance · intent
OpenTelemetry
gen_ai.tool.call (ACS span · gen_ai.tool.name, acs.capability)
OCSF
1007 Process Activity
eEffect

Effect observed independently at the target system with target, digest, approval and idempotency key.

trace_id
"4bf92f3577b34da6a3ce929d0e0e4736"
span_id
"5f1e0c2b9d8a7e64"
parent_span_id
"b7ad6b7169203331"
action_id
"act-9c41e2d07a"
tool_name
"mail.send"
args_hash
"9d02b7…e4f8"
target
"recipient:security@example.org"
audience
"mcp://server-a"
approval_id
"apr-7f3c2a91d0be"
idempotency_key
"idem-5be04d1c"
Counterpart in ACS v0.1
Hook / source
steps/toolCallResult + effect sink
Fields
exit_status · request_id_ref (ACS) · observed target, digest, idempotency key (sink)
OpenTelemetry
gen_ai.tool.result (ACS span · gen_ai.tool.name, acs.exit_status)
OCSF
1007 Process Activity
Invariant violatedRule: Approval binding · First stage: t (Tool)P3: all threat-guided bindings. The invariants are checkable.

Shown as a sketch. In ACS, some of the contract fields reside in the hook payload (input: provenance; tool: tool, arguments, capability), the binding fields in policy_data of the response envelope (decision) or in an independent effect sink (effect). The Guardian derives source trust from origin/source_id; it is not a field in the v0.1 schema. Values such as example.org are placeholders.

Interactive · real measurement data

Fault matrix: where detection holds and where it breaks

Four evidence profiles × seven telemetry conditions, 108 workflow instances per cell. Choose a metric and click a cell to see the results per attack family. All values come from the paper's reproducible run.

Metric
Fault matrix: where detection holds and where it breaks — Security recall
ProfilF0unchangedF1no decisionF2trace splitF3no effect parentF4time offsetF5effect duplicateF6action-ID swap
P0Local logs
P1+ Trace ID
P2+ Span/Parent/Action
P3+ Security Bindings
Selected cell

P3 + Security Bindings × F1 no decision

All decision events removed. In ACS, roughly an unreachable Guardian under failure posture proceed, sampled acs.decision events or a failed trace sink (Trace is best-effort and never blocks enforcement).

50%Security recall
0%Unique attribution
0%Exact reconstruction
0%False-positive rate
100%Precision
50%Correct rule

Unsafe cases per attack family

  • Authority Inversion
  • Approval Replay
  • Scope Amplification
  • Memory Substitution
  • Target Redirection
  • Duplicate Effect

detected (12) False positives on benign twins (6)

Interpretation

The dominant residual weakness: without a decision event, source authority, memory and duplicates remain detectable. Approval, scope and target checks lose their comparison point.

Synthetic, deterministic corpus with 36 templates × 3 repetitions. The rates apply only to the six encoded families. Confidence intervals would be misleading because the repetitions vary only IDs and jitter.

Mapping

ACS crosswalk: attack family → predicate → hook → Guardian reaction

For each of the six families from the paper: which evidence demonstrates the violation, where ACS captures it and how a Guardian should react. MAESTRO, STRIDE and OWASP ASI serve as classification.

Attack familyEvidence predicateMAESTRO · STRIDE · OWASPACS hook & fieldsGuardian reactionFirst stage
Authority InversionOrigin untrusted AND requested operation is a writeL2 Data Operations · L3 Agent FrameworksSpoofing · Elevation of PrivilegeASI01steps/knowledgeRetrieval → Provenance.origin; steps/toolCallRequest.capabilitydeny, or ask with a human approver; reason code untrusted_into_consequentialp
Approval Replay / argument substitutionApproval ID, subject, tool or argument digest at invocation ≠ approved valuesL3 Agent Frameworks · L6 Security & ComplianceTampering · RepudiationASI02 · ASI09AcsResult.policy_data ↔ steps/toolCallRequest.argumentsdeny; after modify, rebind the receipt to the rewritten valuest
Delegation Scope AmplificationEffective capabilities ⊄ approved set OR delegation depth > limitL3 Agent Frameworks · L7 Agent EcosystemElevation of PrivilegeASI03steps/subagentStart · steps/toolCallRequest.capabilitymodify (attenuate capabilities) or denyt
Memory Provenance SubstitutionMemory untrusted or expired AND target ≠ intended targetL2 Data OperationsTampering · SpoofingASI06steps/memoryContextRetrieval → Provenance.derived_fromdeny; quarantine the memory entryt
Post-Decision Target RedirectionApproved target ≠ exercised targetL3 Agent Frameworks · L4 Deployment & InfrastructureTampering · Information DisclosureASI02steps/toolCallRequest.arguments (target) ↔ policy_datadeny; normalize the target before comparison (alias resolution)t
Duplicate Effect under RetryMore than one effect per approval (the paper's evaluator); additionally a missing or rotated idempotency keyL4 Deployment & Infrastructure · L5 Evaluation & ObservabilityRepudiation · TamperingASI08steps/toolCallResult + effect sink (idempotency key)Enforce idempotency; effect sink reports the duplicate as a findinge

The MAESTRO, STRIDE and OWASP ASI mapping and the recommended Guardian reactions are VamiSec interpretations based on the paper. Hook and field names correspond to the ACS schemas v0.1.0. The reason code untrusted_into_consequential is named there as an example category.

Interactive · Threat-to-Assertion

Test compiler: turn a threat into a checkable invariant

The paper translates every MAESTRO/STRIDE hypothesis into the tuple H = (boundary, precondition, invariant, evidence, oracle). Choose one of the five invariant groups. The compiler shows the tuple and two sketches: a Guardian policy over ACS inputs and a negative test for CI.

Hypothesis H = (boundary, precondition, invariant, evidence, oracle)
Boundary
Decision → Tool
Precondition
An approval exists (allow, or ask with consent).
Invariant
Approval ID, subject, tool and argument digest at invocation match the approved values exactly.
Evidence
Decision Receipt in policy_data, tool echo in steps/toolCallRequest
Oracle
Alert on any deviation in one of the four fields. First stage: t
package acs.guardian.approval_binding

# Sketch: an approval is valid only for the exact call it approved.
# data.receipts holds the bound values the Guardian wrote into policy_data.
receipt := data.receipts[input.params.metadata.session_id][input.params.payload.tool.name]

# Digest over argument values only (provenance labels excluded).
values := {k: v.value | some k, v in input.params.payload.arguments}

exercised := {
  "args_hash": crypto.sha256(json.marshal(values)),
  "target": input.params.payload.arguments.target.value,
  "principal": input.params.metadata.user_context.user_id,
}

decision := {
  "decision": "deny",
  "reasoning": "Call does not match the bound approval",
  "reason_codes": ["approval_binding_mismatch"],
} if {
  input.method == "steps/toolCallRequest"
  some field in ["args_hash", "target", "principal"]
  exercised[field] != receipt[field]
}

Sketches for illustration, not normative. Paths such as input.params.payload follow the ACS request envelope v0.1.0; fields in policy_data are the bindings of the Evidence Contract.

Calculator

Evidence budget: what the Evidence Contract means for storage

The basis is the paper's median bytes per workflow. The calculator shows the canonical JSON volume per profile and the prioritization from the paper (complete receipts for privileged operations) and, as a VamiSec assumption, causal tracing for the rest. Semantic detection is lost there, however.

Retention
Volume over the retention period
  • P024 GB
  • P131 GB
  • P245 GB
  • P385 GB
Recommended mix53 GBP3 for privileged workflows, P2 for the rest · 146 MB per day

Indicator of serialization volume: canonical JSON without transport envelopes, compression, indexing, replication or privacy controls. Not a production measurement (Paper §V-C).

Deep Dive · 10 chapters

From paper to practice: agent tests with Evidence Contract and ACS

Problem, threat model, Evidence Contract, the OWASP Agent Control Standard, the mapping between the two, experimental setup, results, cost, a practice playbook and the limits of the claims. The figures come from the paper accepted at ICSPIS 2026; the ACS details come from specification v0.1.0.

01Chapter 1

Why agent tests need evidence

Tool-using AI agents cross trust boundaries before a model-selected operation becomes an external effect. That is exactly where evidence is missing in most environments.

In the Model Context Protocol (MCP), a host creates a client for each server. The host is expected to enforce authorization, consent and security boundaries. Tools are discovered and invoked through structured protocol operations. Their descriptions and annotations can influence model behavior and are not inherently trustworthy. An incident therefore spans content ingestion, planning, approval, tool invocation and a remote effect. It does not remain inside a single model call.

Three observable transitions follow from this architecture: host→client, client→server and server→effect. A trace that ends at a successful tool response can still miss whether an external write occurred, was duplicated, or reached a different target. Ordinary component logs or a shared trace identifier rarely record whether the approval matched the principal, arguments, target and delegated authority at execution time.

Three research questions

  • RQ1: Which evidence profiles reconstruct and uniquely attribute an effect under controlled telemetry faults?
  • RQ2: Do threat-derived semantic bindings improve security detection beyond correlation alone?
  • RQ3: What does the evidence cost, and which failure modes remain?

The contribution has three parts: a four-stage Evidence Contract that maps MAESTRO and STRIDE findings to runtime fields; a deterministic fault-injection corpus with six attack families, closely matched benign variants and a ground-truth ledger that the reconstructor never sees; and a comparison of fragmented, correlated, causal and threat-guided profiles across 3,024 cells.

02Chapter 2

MAESTRO × STRIDE: from threat model to test hypothesis

Threat modeling finds dangerous transitions. An operational test additionally needs evidence that connects the transition to what actually happened.

MAESTRO organizes threats across seven interacting layers of an agentic system and emphasizes cross-layer effects. The layers range from foundation models through data operations, agent frameworks, deployment infrastructure, evaluation and observability, and security and compliance to the agent ecosystem. STRIDE adds six property categories: Spoofing, Tampering, Repudiation, Information Disclosure, Denial of Service and Elevation of Privilege.

The paper uses both frameworks to select boundary-relevant scenarios. It does not claim empirical validation of the frameworks. The OWASP Top 10 for Agentic Applications 2026 serves as a cross-check for the vocabulary, including ASI01 Agent Goal Hijack, ASI02 Tool Misuse and Exploitation, ASI03 Identity and Privilege Abuse, ASI06 Memory & Context Poisoning and ASI08 Cascading Failures; they are treated there as system-level concerns.

Attack familyBoundary testedSTRIDE questionOWASP Agentic reference
Authority Inversion (untrusted source)Source trust at inputSpoofing, Elevation of PrivilegeASI01 Agent Goal Hijack
Approval Replay / argument substitutionBinding decision ↔ invocationTampering, RepudiationASI02 Tool Misuse and Exploitation, ASI09 Human-Agent Trust Exploitation
Delegation Scope AmplificationDelegated authorityElevation of PrivilegeASI03 Identity and Privilege Abuse
Memory Provenance SubstitutionProvenance and target consistencyTampering, SpoofingASI06 Memory & Context Poisoning
Post-Decision Target RedirectionApproved targetTampering, Information DisclosureASI02 Tool Misuse and Exploitation
Duplicate Effect under RetryExecution cardinalityRepudiation, TamperingASI08 Cascading Failures

The selection is purposive, organized by evidence predicates at the four boundaries of the contract, not by attack prevalence. The mapping to STRIDE and OWASP ASI is a VamiSec interpretation based on the paper and is not comprehensive coverage of the frameworks.

03Chapter 3

The Evidence Contract: four stages that substantiate one another

The observable unit is the path P = (p, d, t, e). It becomes security-relevant only when adjacent stages repeat the same facts and thereby become comparable.

Each event has an immutable event identifier, component, stage, schema version and timestamp. Correlation fields add trace, span, parent and action IDs. The threat-guided fields bind facts whose equality or ordering implements a test oracle. What matters is that the tool event repeats the exercised values rather than merely pointing to the decision. Only comparing adjacent stages exposes substitutions that remain invisible in correlation-only traces.

StageRecorded fields (Paper §III-A)
p · InputSource trust, requested operation and scope, intended target, memory provenance
d · Decisionallow/deny, principal, approval ID, approved tool, argument digest, target, audience, capability set, delegation-depth limit, idempotency requirement
t · ToolRepetition of the values actually exercised: tool, arguments or digest, target, capabilities, delegation
e · EffectObserved target, operation digest, audience, approval, idempotency key

What the contract deliberately does not record

No profile records hidden chain-of-thought. Decision evidence is restricted to explicit policy outcomes, concise reason codes, approved attributes and hashes suitable for controlled testing. Sensitive prompt content or model reasoning need not be stored. Provenance labels, stable identifiers and keyed digests are sufficient for tests and reduce disclosure.

04Chapter 4

The OWASP Agent Control Standard (ACS) in ten minutes

ACS is an open wire specification. With it, a Guardian Agent can check an AI agent's actions before execution and allow, block, rewrite, escalate for approval or defer them (allow, deny, modify, ask, defer) — over signed envelopes (ACS-Core requires them; the reference implementation, a proof of concept, does not yet implement this) and with an audit trail.

Since September 2026, ACS has been a project of the OWASP GenAI Security Project, at Level 2 (Incubator) according to its own project metadata; specification v0.1.0 dates from 5 June 2026, the current release tag is 0.1.2. It grew out of the Agent Observability Standard (AOS). Under ACS, an agent is considered trustworthy if it is inspectable, traceable and instrumentable. Operators should thus be able to determine after the fact what an agent did and why, and to limit in advance what it may do.

19native steps/* hooks in specification v0.1.0 — from sessionStart to sessionEnd, including the three skill hooks
5dispositions: allow, deny, modify, ask, defer
44JSON schemas under /schema/v0.1.0/
3pillars: Instrument, Trace, Inspect

The three pillars

Instrument

Hooks at consequential execution points: input, knowledge retrieval, memory, tool calls, compaction, sub-agents and skills. Code, shell, file and network actions go through steps/toolCallRequest, which must fire for every action that leaves the reasoning context. The Guardian responds with a disposition. The primary enforcement gate is steps/toolCallRequest. steps/toolCallResult serves as a gate for redacting outputs before they reach the agent.

Trace

Normative vocabulary for OpenTelemetry and OCSF. A tool call becomes the span gen_ai.tool.call with gen_ai.tool.name and acs.capability. Decisions are captured as the span event acs.decision on the span of the checked step, not as a separate span. In OCSF, deny, modify, ask and defer become Detection Findings (class 2004). The mappings have the status “Working draft”; Trace is best-effort and a self-report of the observed environment. The three skill hooks are not mapped, and gen_ai.tool.call is an ACS span name, not an OpenTelemetry GenAI span (there: execute_tool {gen_ai.tool.name}).

Inspect

An Agent Bill of Materials (AgBOM) discloses the agent's tools, models and accessible data, mappable to CycloneDX, SPDX and SWID. The methods agbom/snapshot and agbom/changed map to OCSF class 5001 Device Inventory Info (ACS calls it “Inventory Info”). Serialization to CycloneDX 1.6, SPDX 3.0 or SWID (at least one; status “Working draft”).

Wire, provenance and audit chain

  • Wire format JSON-RPC 2.0. Method namespaces: steps/* (hooks), protocols/* (encapsulated MCP; A2A is reserved for v0.2), agbom/*, system/* and handshake/*. On a version conflict, the handshake ends with UNSUPPORTED_VERSION. The spec prose still assumes MCP mechanisms predating revision 2026-07-28, and whether MCP wrapping belongs to ACS-Core is contradictory in the spec text; MCP tool calls may also be carried via steps/toolCallRequest/-Result.
  • Provenance is a factual origin label on data-carrying fields. It is set by deterministic framework code, never by the LLM: origin (user_input, system, tool_output, retrieved, agent_generated, a2a_inbound, external), source_id and derived_from as lineage edges.
  • Under the acs-provenance profile, field-level provenance is mandatory for every data-bearing field of the session, including every argument of steps/toolCallRequest. The trust classification is made by the Guardian according to local policy; it is not a field in the v0.1 schema.
  • Every response for a content-bearing step carries a chain_hash: the rolling SHA-256 head of the audit chain. policy_references (including the optional policy_version) and cited_provenance_ids make the decision replayable.
  • Signatures are crypto-agile and mandatory in ACS-Core: every request and every response is signed over the canonical envelope; HMAC-SHA256 with an HKDF-derived session key is the baseline that satisfies this. It makes tampering on the network evident, but not a compromised Guardian. Non-repudiation only comes with the acs-crypto profile: ML-DSA-65 (mandatory) and SLH-DSA-128s (recommended as backup).
05Chapter 5

Evidence Contract ↔ ACS: what the standard covers and what you have to add

For input, decision and tool, ACS provides the hooks, envelopes and trace mappings; for the effect, only the tool's own response. The two most important gaps are the decision binding, whose absence explains the dominant residual weakness F1 in the experiment, and independent effect evidence, which the paper justifies as a requirement but does not measure separately.

Contract stageACS equivalent (v0.1.0)What needs adding
p · Input & Provenancesteps/userMessage, steps/knowledgeRetrieval, steps/memoryContextRetrieval; provenance origin/source_id/derived_from; OTel attribute acs.provenance.origin; OCSF enrichment acs_provenance_originMaintain source trust as Guardian policy (the v0.1 schema has no trust field)
d · DecisionAcsResult: decision, reasoning, reason_codes, policy_references (policy_version), cited_provenance_ids, ask_details (approver, timeout_disposition), chain_hash; span event acs.decision; OCSF 2004 Detection FindingBind argument digest, normalized target, capability set, delegation depth, expiry and idempotency requirement in policy_data; persist the Decision Receipt durably
t · Tool & Delegationsteps/toolCallRequest: tool, operation, capability, arguments with provenance, raw_command, intent; steps/subagentStart or agentTrigger (a2a_inbound); turn_id/parent_turn_id; span gen_ai.tool.call; OCSF 1007Compare exercised values against the receipt: equality of argument digest, target and capability as test oracle
e · Effectsteps/toolCallResult (exit_status, request_id_ref); span gen_ai.tool.resultIndependent effect observation at the target system (gateway, DB audit, sink) with observed target, operation digest and idempotency key, correlated via the action ID

The ACS details come from the published JSON schemas v0.1.0 (hooks/*, provenance.json, response-envelope.json, trace/otel-mapping.json, trace/ocsf-mapping.json). policy_data is explicitly provided in the standard as a structured, policy-specific field. According to the specification, ACS Trace events are self-reports of the observed environment and best-effort; they become evidence only when a party outside the emitting runtime attests to them. Signatures (acs-crypto) and content binding (acs-audit) strengthen integrity. The experiment assumes authentic telemetry (Chapter 10).

Why the effect gap is structural

ACS hooks fire inside the agent runtime. steps/toolCallResult reports what the tool returns to the agent. It does not report what actually happened at the target system. The paper shows why this difference matters: a trace that ends at the successful tool response can miss duplicate writes or redirected targets. Effect evidence must therefore come from a second source that is independent of the agent. It is bound to the ACS trace via action and request IDs.

Why fail-open is the F1 problem

If the Guardian is unreachable under failure posture proceed, the action continues. ACS does require logging every fail-open proceed as an audit event, but no decision receipt with approval bindings is created. Operationally, this corresponds to fault F1 from the experiment (VamiSec interpretation). If the decision events are missing, exact reconstruction and attribution fall to 0% and the recall of the threat-guided profile to 50%. An ACS deployment that takes security testing seriously therefore negotiates on_decision_failure: deny (in the reference implementation: ACS_ON_DECISION_FAILURE=deny), sets the startup posture refuse on the Observed Agent and treats a missing receipt as a test failure, not as a gap in the log.

06Chapter 6

Experimental setup: corpus, profiles and faults

36 workflow templates, three deterministic repetitions, four evidence profiles, seven telemetry conditions. Plus a ground-truth ledger that the reconstructor never sees.

The corpus contains six workflow families, each with four unsafe and two closely matched benign variants: 36 templates, of which 24 are unsafe and 12 benign. Three deterministic repetitions yield 108 workflow instances. The repetitions vary identifiers and event jitter. They are a reproducibility check, not independent statistical samples.

A trusted generator emits canonical, MCP-shaped JSON events. Separately, a ground-truth ledger is produced containing expected event identities, parent edges, unsafe labels, the violated rule and the first violation stage. Before reconstruction, private case IDs are stripped from the observable events. Workflows in the same family share an intentionally overlapping time window. This is a difficult case for time-nearest attribution.

ConditionFaultReal-world counterpart in ACS operation
F0Evidence unchangedReference
F1All decision events removedGuardian unreachable under proceed, sampled acs.decision events, discarded Detection Findings, a failed trace sink (Trace is best-effort and never blocks enforcement)
F2Trace IDs replaced at the tool boundaryBroken context propagation between MCP client and server
F3Parent links removed from effect eventsEffect telemetry from target systems without trace context
F4Timestamps shifted, order reversedClock skew, asynchronous exporters
F5Duplicate effect telemetryAt-least-once export, retries in the collector
F6Action IDs swapped between concurrent workflowsConcurrency bugs in instrumentation or collector

The column “Real-world counterpart” is a VamiSec interpretation, not a result of the paper. Faults that target fields not present in a profile are treated as inapplicable for fault-detection scoring.

Reconstruction and metrics

  • P0 groups an effect with the temporally nearest tool, P1 by trace ID. P2 and P3 follow parent edges and fall back to action IDs when an effect parent is missing.
  • Exact reconstruction requires prompt, decision, tool, all effects and, in P2/P3, the expected parent edges.
  • Unique attribution requires all expected effects to map to exactly the correct prompt, decision and tool, even if a parent edge is missing.
  • Recall and precision use the unsafe label, the benign false-positive rate the 12 benign templates. Correct-rule and first-stage rates test the diagnosis.

The implementation uses only the Python standard library. Two clean executions produced byte-identical configuration, case, aggregate and summary hashes. Three unit tests check matrix size and balance, deterministic outputs and the P3-versus-P2 recall comparison.

07Chapter 7

Results: what holds and what breaks

Causal evidence survives trace splits and lost effect parents. A missing decision stage, however, prevents complete reconstruction and removes essential security meaning.

ProfileExactAttributionRecallPrecisionFalse positivesBytes
P0 · local IDs, time0.0%0.0%84.5%66.4%85.7%659
P1 · + Trace ID71.4%71.4%16.7%100%0%843
P2 · + Span/Parent/Action71.4%85.7%16.7%100%0%1,245
P3 · + Security Bindings71.4%85.7%92.9%100%0%2,341

Table II of the paper: rates aggregated across all seven telemetry conditions. Bytes = median canonical JSON per workflow under F0.

P0 never achieved unique attribution because temporally overlapping workflows were merged. Its apparent recall of 84.5% is not useful security performance: incorrectly merged chains raised alerts on 85.7% of benign cases. P1 detected only duplicate-effect cases. P2 raised unique attribution to 85.7%; security recall remained at 16.7%. P3 retained P2's attribution and raised recall, correct-rule rate and first-stage localization to 92.9%.

Fault sensitivity

  1. F0Unchanged evidence

    P1 through P3 reconstruct and attribute every workflow. P3 achieves 100% recall and precision with no false positives.

  2. F2Trace split at the tool boundary

    P1 loses reconstruction and attribution completely. P2 and P3 remain at 100% because parent and action IDs bridge the modified trace boundary.

  3. F3Lost effect parents

    Exact reconstruction of P2/P3 falls to 0%; unique attribution stays at 100% via the action fallback.

  4. F4–F6Time offset, duplicates, action-ID swap

    No loss for P2/P3: the remaining evidence yields an unambiguous path, and corruption checks expose the injected inconsistency.

  5. F1Missing decision events

    The dominant residual weakness: reconstruction and attribution fall to 0% in P1 through P3, and the recall of P3 to 50%. Source authority, memory and duplicates remain detectable. Approval, scope and target binding lose their comparison point.

08Chapter 8

Evidence cost and operations: prioritize instead of sampling

More bindings mean more bytes. The answer is deliberate prioritization, not sampling that cuts away Decision Receipts.

659 BP0 per workflow (median, canonical JSON, F0)
843 BP1: +27.9%
1,245 BP2: +88.9%
2,341 BP3: +255.2% vs. P0, +88.0% vs. P2

P3 adds 1,682 bytes (about 1.64 KiB) over P0 and 1,096 bytes over P2. A collector using the same event representation would retain about 3.55 times the canonical bytes of P0, or 1.88 times those of P2, for a median-sized workflow. These values answer RQ3 as a serialization-volume comparison. They do not measure latency, CPU, transmitted traffic, compression, indexing or monetary cost.

Operating rules from the paper, translated to ACS

  • Complete decision, invocation and effect receipts first for privileged operations: in ACS terms, for capabilities such as filesystem.delete, network.egress or process.execute.
  • Set explicit retention limits. Deduplication, compression and keyed digests can reduce volume but must be measured separately.
  • Anyone who drops an invariant-relevant field or samples away Decision Receipts changes the Evidence Contract and must re-evaluate. F1 shows what is lost then.
  • Model integrity, clocks, retention and access control separately. ACS offers the HMAC baseline (acs-core), asymmetric or PQC signatures (acs-crypto), request_hash (acs-audit) and the chain_hash of the audit chain for this.
09Chapter 9

Practice playbook: ACS-based security tests in the pipeline

This is how you translate paper and standard into a test suite that runs in CI and reliably turns red when evidence is missing.

  1. Step 1Translate threats into hypotheses

    Walk through MAESTRO layers and STRIDE properties for each agent workflow. For every dangerous transition, formulate a tuple H = (boundary, precondition, invariant, evidence, oracle).

  2. Step 2Define ACS hooks and profiles

    ACS-Core as the baseline, plus the acs-trace and acs-provenance profiles. In addition to the six mandatory Core hooks (sessionStart, userMessage or agentTrigger, toolCallRequest, toolCallResult, agentResponse, sessionEnd), instrument steps/knowledgeRetrieval, steps/memoryContextRetrieval and steps/subagentStart — and check that the Guardian lists them in methods_evaluated (otherwise they are ALLOW-by-default).

  3. Step 3Define the Decision Receipt

    Principal, policy version, reason code, operation, normalized target, argument digest, capabilities, delegation depth, expiry and idempotency in policy_data. Persisted and bound to request_id and chain_hash.

  4. Step 4Harden the failure posture

    Set the failure posture to fail-closed: on_decision_failure: deny in the ServerHello (reference implementation: ACS_ON_DECISION_FAILURE=deny) and the startup posture refuse on the Observed Agent in case the handshake itself fails. Sign envelopes. A missing receipt is a test failure.

  5. Step 5Connect an independent effect sink

    Gateway, database or sink audit supplies observed target, operation digest and idempotency key, correlated via the action ID.

  6. Step 6Negative tests with benign twins

    For each family, unsafe and closely matched benign cases, such as permitted alias resolution or repeated reads. Freeze predicates before testing.

  7. Step 7Two CI gates

    Gate 1 reconstruction: required stages and parent edges complete, otherwise quarantine. Gate 2 correctness: invariants satisfied, otherwise abort with the violated rule and first stage.

  8. Step 8Inject telemetry faults

    Run F1 through F6 against your own pipeline regularly. A test that is still green under F1 does not check approvals.

Negative test: Approval Replay

Grant an approval for tool A with argument digest X, then call tool A with digest Y. The expectation is that the approval-binding invariant is violated at stage t. Benign twin: fresh, bound approval with retry.

Negative test: Target Redirection

Approved target tenant-a, exercised target tenant-b. The expectation is that target equality is violated at stage t or e. Benign twin: permitted alias resolution to the same canonical target.

Negative test: Scope Amplification

Sub-agent inherits write although only read was delegated. Expected: effective capabilities ⊄ approved set. Benign twin: attenuated child with read-only.

10Chapter 10

Limits of the claims and the next validation step

The experiment validates evidence sufficiency for deterministic, MCP-shaped fixtures. It is neither an LLM benchmark nor a production robustness benchmark.

The study uses synthetic deterministic fixtures, handcrafted invariants and fixed event schedules. It executes no LLM, no MCP implementation, no OAuth, no policy engine and no external service. It neither measures prompt-injection success nor demonstrates that a control prevents an effect. Precision and recall apply only to the six encoded families. Because the evaluator was designed from the same threat hypotheses, the experiment tests evidence sufficiency and implementation consistency, not the discovery of new threats.

  • The ground-truth ledger is logically separated but generated by the same codebase. Independent implementations, property-based mutation and real protocol captures would reduce common-mode error.
  • Telemetry is assumed to be authentic. Spoofed producers, log truncation beyond the injected faults, compromised collectors and cryptographic verification remain open.
  • New threat classes need their own hypothesis, fields, oracle and independently authored tests. Resource exhaustion, for example, requires budget and consumption evidence; forged telemetry requires producer identity and integrity verification.

Disclosure from the paper: OpenAI Codex was used for literature organization, implementation scaffolding and review of the synthetic evaluation code, initial drafting and language revision, and formatting. The research question, threat model, execution, checking of the results and sources, and all final decisions rest with the author.

Self-check

Evidence Readiness Radar: Are your agents really testable?

Ten questions across five dimensions, derived from the paper's Evidence Contract and the profiles of the OWASP Agent Control Standard. The result shows where your telemetry does not yet evidence approvals, targets, delegation and effects.

  1. ProvenanceACS · Provenance
    Does every piece of content entering the agent (user input, retrieval, tool output, memory) receive a deterministically assigned provenance label with its source?
  2. ProvenancePaper §IV-C
    Does a policy prevent content from untrusted sources from authorizing a consequential operation?
  3. Decision ReceiptsPaper §VI-A
    Is there a persisted decision receipt with principal, policy version and reason code for every consequential tool execution?
  4. Decision ReceiptsPaper §III-A
    Does the decision receipt bind argument digest, normalized target, capability set, delegation depth and expiry?
  5. Tool & DelegationACS · steps/toolCallRequest
    Does the tool event repeat the values actually executed, instead of merely pointing to the decision?
  6. Tool & DelegationACS · steps/subagentStart
    Are delegations to sub-agents recorded with effective capabilities and depth and checked against the approved set?
  7. Effect evidencePaper §II-A
    Is the effect at the target system observed independently of the agent, for example via gateway, database or sink audit?
  8. Effect evidencePaper §III-B
    Do write effects carry an idempotency key, so that duplicate executions under retry are detectable?
  9. Operations & integrityACS · Failure Posture
    Is your Guardian configured fail-closed, so that no action runs without a decision?
  10. Operations & integrityPaper §VI-A
    Do your CI security tests fail when required stages or parent edges are missing from the evidence?
ProvenanceDecisionReceiptsTool &DelegationEffectevidenceOperations& integrity
Provenance0 %Decision Receipts0 %Tool & Delegation0 %Effect evidence0 %Operations & integrity0 %

Answer all ten questions. The result appears here.

Research Edition · free download

Threat-Model-Guided Security Testing for Agentic AI: Paper and ACS Practitioner Brief

The paper accepted at ICSPIS 2026 as the accepted version, supplemented by a practical part. It shows how you implement the Evidence Contract with the OWASP Agent Control Standard and bring it into your pipeline as negative tests.

Cover of the VamiSec Research Edition “Threat-Model-Guided Security Testing for Agentic AI”
Paper (5 pp.) + ACS briefPDF, free of chargeEnglishAs of 10/2026
  • Full paper: method, results across 3,024 cells, fault analysis and validity boundaries
  • Mapping Evidence Contract ↔ ACS v0.1: hooks, provenance, dispositions, OTel and OCSF
  • Decision Receipt minimum schema and recommendations on failure posture (fail-closed instead of proceed)
  • Negative test catalog with six attack families and CI gates for reconstruction and correctness
Free Download

Request Whitepaper

Research Edition: Threat-Model-Guided Security Testing for Agentic AI (ICSPIS 2026)

What the Research Edition contains

The full paper

The accepted version of the ICSPIS 2026 paper with method, Table I/II, fault analysis, evidence cost and validity boundaries. Published with the IEEE copyright notice.

ACS Practitioner Brief

Four pages: the mapping of the Evidence Contract to ACS hooks, provenance, dispositions, OpenTelemetry spans and OCSF classes, including the gaps in ACS v0.1 that you have to close yourself.

Decision Receipt schema

The minimum fields of a privacy-minimized decision receipt: principal, policy version, reason code, operation, target, argument digest, capabilities, delegation depth, expiry and idempotency key.

Negative test catalog

Six attack families with four unsafe and two benign variants each as a template for your own test suite, plus CI gates for reconstruction and correctness.

The standard in depth

OWASP Agent Control Standard: understand runtime control before you test it

This page uses ACS as the vocabulary for test evidence. How the standard itself works – Guardian and Observed Agent, 19 hooks, five dispositions, handshake and failure posture, profiles and the maturity of v0.1.0 – is explained in our deep dive with a Guardian simulator, hook lifecycle and risk matrix.

19native steps/* hooks
5dispositions: allow, deny, modify, ask, defer
7profiles: ACS-Core plus six optional ones
  • Guardian simulator with schema-faithful wire examples
  • Risk matrix OWASP Agentic Top 10 × ACS building blocks
  • ACS readiness assessment and CISO whitepaper (in German)
FAQ

Frequently asked questions about Agentic AI Security Testing and ACS

Short answers to the questions we receive most often about the paper, the Evidence Contract and the OWASP Agent Control Standard.

Threat-Model-Guided Security Testing derives an AI agent's telemetry and test oracles directly from a threat model, in the paper from MAESTRO and STRIDE. Each threat becomes a testable tuple of boundary, precondition, invariant, evidence and oracle. The test then deterministically checks whether the recorded facts satisfy the invariant, for example whether tool, arguments and target at invocation match the approved values.

ACS provides the control points and the vocabulary: hooks such as steps/toolCallRequest, provenance on arguments, dispositions, chain_hash and OpenTelemetry and OCSF mappings. The Evidence Contract from the paper defines which facts must at least be bound at these points so that a test can check approval, target, scope and provenance. Operators add two building blocks themselves: a durable decision receipt with binding fields in policy_data, and effects observed independently at the target system. The standard itself – Guardian, 19 hooks, five dispositions, profiles and maturity – is explained in our deep dive “Agent Control Standard (ACS)” in the Agentic AI Security knowledge area.

No. Trace and parent IDs show which events belong together. They do not show why an action was allowed. In the experiment, the purely causal profile uniquely attributed 85.7% of effects but detected only 16.7% of unsafe cases. Only security-relevant bindings such as approval ID, argument digest, target, capabilities and provenance raised recall to 92.9%.

The paper recommends a durable, privacy-minimized decision receipt that binds principal, policy version, reason code, approved operation, normalized target, relevant argument digest, delegated capabilities and expiry to an action identifier. In ACS, decision, reasoning, reason_codes, policy_references and chain_hash are already part of the response envelope. You add the binding fields in policy_data.

With proceed, an action continues if the Guardian is unreachable. ACS does require logging every fail-open proceed as an audit event, but there is then no decision receipt with approval bindings. In our interpretation, this corresponds to fault F1 in the experiment: without decision events, reconstruction and attribution fell to 0% and recall to 50%. The specification explicitly provides on_decision_failure: deny (fail-closed) for deployments that cannot accept trading enforcement for availability; for failed handshakes, the startup posture refuse, configured outside the protocol on the Observed Agent, is added; otherwise a failed handshake starts the session unprotected (its default is proceed as well). The project README advises setting ACS_ON_DECISION_FAILURE=deny before trusting the Guardian to fail closed. The proceed default itself is contested within the project (issues #32 and #37).

At the median, the threat-guided profile produced 2,341 bytes of canonical JSON per workflow. That is 255.2% more than fragmented logs (659 bytes) and 88.0% more than causal tracing (1,245 bytes). The value is an indicator of serialization volume without compression, transport or indexing. In practice, you introduce complete receipts for privileged operations first.

No. The experiment validates the evidentiary value of the evidence for deterministic, MCP-shaped fixtures. It runs no LLM, no real MCP implementation and no policy engine, and does not measure whether prompt injection succeeds. Benchmarks such as InjecAgent or AgentDojo answer the robustness question. The paper picks up from there and clarifies which evidence an evaluator minimally needs to reconstruct and classify a path.

The paper was accepted for the 9th International Conference on Signal Processing and Information Security (ICSPIS 2026), which takes place in Dubai from 10 to 12 November 2026. The accepted version is available here as the Research Edition for download, with the IEEE copyright notice and a supplementary ACS Practitioner Brief. After publication, we will add the full citation with a link to IEEE Xplore.

Standards & Sources

The content on this page is based on the following publicly available guides and studies.

V. Milke, VamiSec GmbH · ICSPIS 2026 (accepted) · 2026

Threat-Model-Guided Security Testing for Agentic AI: MCP Traces with MAESTRO and STRIDE

Primary source of this page: Evidence Contract, fault-injection corpus and results across 3,024 cells. Accepted version with IEEE copyright notice.

OWASP GenAI Security Project · 2026

Agent Control Standard (ACS)

Listed in the Agentic Security section, dated 1 September 2026

GenAI Security Project (GitHub) · 2026

agent-control-standard: specification, schemas and reference implementation (proof of concept)

README with the known gaps of the reference implementation (wire authentication, failure posture proceed) and roadmap (v0.2.0 target March 2027). Conformance in v0.1.0 is self-declared, without a test suite (issue #19).

OWASP GenAI Security Project · 2026

ACS JSON Schema v0.1.0 (44 Schemas)

Request/response envelope, provenance, hooks and OpenTelemetry and OCSF mappings: basis of the mapping on this page.

Model Context Protocol · 2026

Model Context Protocol: Architecture (Revision 2026-07-28)

Host, clients and servers; the host's responsibility for authorization and consent.

Model Context Protocol · 2026

Model Context Protocol: Tools (Revision 2026-07-28)

Discovery and invocation of tools; recommendation that humans be able to deny sensitive operations.

K. Huang, Cloud Security Alliance · 2025

Agentic AI Threat Modeling Framework: MAESTRO

Seven interacting layers, focus on cross-layer effects.

Microsoft Learn

Threats: Microsoft Threat Modeling Tool (STRIDE)

The six STRIDE categories as property anchors.

W3C · 2021

Trace Context (W3C Recommendation)

Standardized trace and parent IDs for distributed correlation: basis of profiles P1/P2.

OWASP GenAI Security Project · 2025

OWASP Top 10 for Agentic Applications for 2026

ASI01–ASI10 as a cross-check for the scenario vocabulary.

Zhan, Liang, Ying, Kang · Findings of ACL · 2024

InjecAgent: Benchmarking Indirect Prompt Injections in Tool-Integrated LLM Agents

Benchmark for indirect prompt injection via tools.

Debenedetti et al. · NeurIPS Datasets and Benchmarks · 2024

AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents

Extensible environment with realistic tasks, attacks and defenses.

NIST · 2008

NIST SP 800-115: Technical Guide to Information Security Testing and Assessment

Repeatable test procedures and clear evidence.

OWASP Foundation · 2025

OWASP AI Testing Guide

Control-oriented testing of AI systems.

OCSF Project

Open Cybersecurity Schema Framework (OCSF)

Event classes ACS maps to: 3002 Authentication, 6002 (“Application Activity” in ACS, Application Lifecycle in OCSF), 1007 Process Activity, 6005 Datastore Activity, 2004 Detection Finding, 5001 Device Inventory Info.

Are your agents testable, or just logged?

In an initial consultation, we check whether your agent telemetry really evidences approvals, targets, delegation and effects. We also clarify how you can bring ACS hooks, Guardian policies and negative tests into your pipeline.