Disciplines · Audits

Eve SOTA external standards and risk crosswalk — 2026-09-01

Four boundary decisions follow:

5sections16 minread

On this page

Task: EVE_SOTA_GAP_CLOSURE 0.4
As of: 2026-09-01
Exact retrieval: 2026-09-01T17:12:02Z
Machine record: eve-sota-external-crosswalk/2026-09-01.json

This is a point-in-time requirements and risk crosswalk, not a certification. It gives existing implementation its narrowest supportable credit and records what that evidence does not prove. In particular, custom types named MCP, A2A, trace, state, or accessibility do not establish conformance to an external specification with a similar concept.

Normativity and decision rules#

Label Meaning in this record
GUIDANCE Voluntary NIST guidance/taxonomy or OWASP community risk guidance. Eve adopts the risk as planning input; the publisher does not impose a MUST and this record claims no certification.
MUST_IF_ADOPTED A normative protocol requirement only after Eve claims the pinned protocol/version and the applicable capability. Until the named ADR/admission, it is a conditional requirement, not shipped scope.
SHOULD_IF_ADOPTED A normative recommendation under the same conditional protocol boundary. A deviation needs an explicit reason once that protocol/capability is adopted.
NORMATIVE_RECOMMENDATION Normative W3C Recommendation content. A conformance claim is optional; all A+AA criteria bind a Level AA claim once made. Eve has separately adopted WCAG 2.2 AA as an internal release gate.
OPTIONAL_PATTERN AG-UI compatibility is an adopt/adapter/bespoke decision. Its vocabulary is not silently mandatory.
DEVELOPMENT_CONVENTION The OpenTelemetry GenAI material is an untagged development snapshot. It may inform task 13.2 but cannot be described as a stable compatibility target.

Four boundary decisions follow:

  1. NIST AI RMF, the NIST GAI Profile, NIST AML, and OWASP Agentic Top 10 are risk/action guidance. Internal tasks can make a mitigation mandatory without turning the publisher document into a binding standard.
  2. MCP 2025-11-25 and A2A v1.0.0 RFC-style requirements apply only after the protocol and relevant capability are admitted. Current local primitives are precursors, not conformance evidence.
  3. AG-UI remains optional pending task 16.1. Similar bespoke SSE/UI events are useful implementation evidence but are not renamed for compliance theater.
  4. WCAG 2.2 Level AA is an initiative release requirement. Current automation has known carve-outs and cannot support a current conformance or legal claim. OpenTelemetry alignment remains optional and must be pinned by task 13.2.

Dated source ledger#

Source Exact version/date Authority Exact pin Force
NIST AI RMF NIST AI 100-1 · 2023-01-26 NIST publication PDF SHA-256 7576edb531d9848825814ee88e28b1795d3a84b435b4b797d3670eafdc4a89f1 · 1,946,127 bytes Voluntary guidance
NIST GAI Profile NIST AI 600-1 · 2024-07-26 NIST publication PDF SHA-256 6e73620ab6b64e90ef2c04bf0e0d6246185a2f4b1b13cab0df494496cff89b6a · 1,174,643 bytes Voluntary guidance
NIST AML NIST AI 100-2 E2025 · 2025-03-24 NIST CSRC Corrected PDF SHA-256 4811fb6ad73f9c9121843ab77e029b5adc6f2c86d33c2fc5b2099ef133847646 · 1,964,469 bytes Taxonomy/guidance; errata recheck required
OWASP Agentic Top 10 Version 2026 · 2025-12-09 OWASP resource PDF SHA-256 a2db94cd00b08e0b3a5e5b619afe024bdbcd74503111085705e4f3dd886fcb5c · 1,274,186 bytes Community risk guidance
MCP 2025-11-25 · 2025-11-25 Official specification tag 2025-11-25; commit 38c84e9f93ad191d9eb26d92b945d17bd0efcaf3 Normative only if version/capability adopted
A2A v1.0.0 · 2026-03-12 Official specification tag v1.0.0; commit 173695755607e884aa9acf8ce4feed90e32727a1 Normative only if adopted
AG-UI release/2026-07-28 · 2026-07-28 Official documentation tag release/2026-07-28; commit ea3db4630d1e56d289093cbda0fe66929a44b6ed Optional compatibility candidate
WCAG 2.2 W3C Recommendation · 2024-12-12 Dated W3C Recommendation HTML SHA-256 6e3c5fe397257cae509a2fb4752b73062cf8cbeb92c2cec618989b17e4cf7057 · 512,457 bytes Normative content; internally adopted Level AA gate
OpenTelemetry core semantic conventions v1.44.0 · 2026-08-04 Official specification tag v1.44.0; commit e10a930844c6951757a43b849d364f7d056ac32b Optional internal adoption
OpenTelemetry GenAI semantic conventions untagged development snapshot · 2026-09-01 Official repository commit ac46a5d7bfe0b0f47e8ce393e2db3a2c3042f236 Development input, not stable conformance

NIST AML pins the corrected 2025-04-01 PDF. NIST's 2025-06-03 planning note still reports an error on page x and a possible future update, so a byte match does not waive the task 18.3 errata review. The NIST AI RMF is also being revised; 1.0 remains the deliberate dated baseline here.

Repository-backed artifact hashes#

The commit pins above are not substitutes for exact content pins. The verifier locks every entry below against the machine record.

  • MCP: authorization 8182f6a204013b497369c2ad690ff313f8bd1d2ce9ebb68e8f5d0392aa348cb9; cancellation d438bcff38437259bf72e5b38fb3bb86925846fac58b9b166fea9edd7cb2c842; progress 8de8c02945544dde615be30156638c2357125795640d5aef3041898032a0e431; tasks 6afbdf560a2f0740f6f168b40b3f36a9bee522c4d3c6037373fc48b54aa8e462; tools d4b51b2ee107f4b6b28f14479defb7904f373fff7b122aec55f1ef71f4244b73; schema 1ffe4c5577974012f5fa02af14ea88df4b7146679df1abaaad497c8d9230ca8a.
  • A2A: specification 087c6f6272ec08a7fcd90eb16843bf97e25e0dcd155afa5779033d23d5ac377d; protobuf 4b74c0baa923ae0acb55474e548f1d6e5d3f83b80d757b65f8bf3e99a3c2257f.
  • AG-UI: introduction f5e8babd6b5184cba5b7f644279bce3ed0a431628ddf0134e2fb9e1a106c9740; events a67dc47fd3875aa8fa3c1d553c7e7b0d78eeecbfa1efb8751c9b21bcc0a44852; state 40e81e96a57ecffb4ca6e185b66e703d3fdeab4c45ee133d0b45117959b96a7a; interrupts 240fc3e310fb87367f1f6b7dc3558c2a09233ab34612123601ce8cfe9c4d6acc.
  • OpenTelemetry core README: b256742d8553beb3aa9c2264a6eebe219392dad26cc4319caeb708d23c38dc30.
  • OpenTelemetry GenAI: agent spans b6298d75a7dca8dd18511fb206ce3447845537b15bc4d96d2ba35fb491e10a44; spans a5e437839341470fa0110574903ab1198b0240ea144ece045e3f28e57659a287; metrics afe9d852b4122d9300d999aec3979a318b840dbd6410ba63004aa320033cda77; MCP d5eb3b023f553c56052a07aecf3384ac5f64f894825c72eacf8c0eeda3c49a26; spans model 7bc1a3025319821cf8fc3e730e827149a9640ef2e4227bdb5abb02eed7faef69; metrics model e490bcfc2d36696dc563fc642a609ead1bb603cc8a9285a34674f8b48c08395d; registry 8381a0cd47111869d2fbcf847da357504ba7a554ad2ed6c3d4b416d346f9f4fa.

Existing-control evidence register#

Crosswalk rows reference these controls by id. Each link is actual source or a domain test; the machine record carries a separate proves and doesNotProve statement for every item.

Control id What currently exists Direct evidence
initiative-governance Cross-phase admission gates and preregistered outcome owners EVE_SOTA_GAP_CLOSURE_TODOS_2026-09-01.md; 2026-09-01.json
bounded-mcp-tool-surface Six bounded stdio workbench tools, surfaced errors, live smoke server.mjs; eve-codex-agent.mjs
workbench-agent-auth Fail-closed shared/per-agent bearer boundary and impersonation negatives agent-auth.ts; agent-auth.spec.ts
governed-write-confirmation One-shot, user/session/revision-bound, fresh confirmations action-confirm.ts; workbench-kit-write.ts; workbench-agent-tools.integration.spec.ts
attributable-work-ledger Event-sourced lifecycle, leases, attributable actors, machine verification intent-machines.ts; intent-store.ts; work-queue.integration.spec.ts
untrusted-page-bounds Shape and size caps on page context page-context.ts
injection-regression-deck Eight observable injection cases across page, selection, docs, tool, role, and forged-result channels deck-adversarial-cases.ts
operator-memory-boundary Subject-scoped durable memory, refusal, caps, delete, kill switch operator-memory.ts
assistant-safety-and-checkers Safety supersede and evidence-before-claim holds assistant.ts
turn-abort-and-sse Socket-close abort before next provider call and ordered local SSE frames assistant.ts; turn-stream.ts
bespoke-ui-intents Validated highlight/tour/client-tool/confirmation UI model turn-stream.ts; AssistantPanel.tsx
custom-a2a-precursors In-house cards, tasks, streams, gateway, routing, HTTPS/HMAC push agents.ts; integrations.ts
iris-signed-agent-precursor Generated Iris card signing and task/cancel prototype agent-card.ts; task-negotiation.ts; README.md
axe-accessibility-gates Rendered-route WCAG-tagged axe checks with documented carve-outs accessibility-key-workspaces.spec.ts; wcag-aa-signoff-v1-p2-3512.spec.ts; accessibility.ts
wcag22-technique-inventory Complete A/AA criterion inventory and evidence-shape analysis of the six 2.2 additions wcag22-coverage.ts; primitive-accessibility.ts
structural-turn-traces Opt-in content-minimized JSONL turn traces turn-trace.ts
assistant-metrics Durable provider/model outcome, cost, token, latency, routing, checker, and refusal aggregates turn-metrics.ts
in-memory-trace-model In-house trace/span hierarchy and AI monitor API observability.ts
trace-manifest Declared trace and correlation field inventory tracing-manifest.ts

Requirement/risk → control → gap → owner#

The compact rows below are normalized through the control register above. The machine record is authoritative for full requirement ids, evidence boundaries, and exact downstream task arrays.

NIST AI RMF 1.0#

Row Requirement/risk Existing controls Remaining gap Phase/task owner
nist-rmf-govern GOVERN 1–6: policies, accountable roles, workforce, monitoring, engagement, third parties initiative-governance; attributable-work-ledger Evidence manifest, external registry, independent signoff, recurring review Agentic AI PM · 0.5, 0.6, 14.6, 16.5, 18.1, 18.3
nist-rmf-map MAP 1–5: context, categorization, benefits/harms, affected actors, dependencies initiative-governance Threat model, full data flow, modality map, dependency registry Security Lead · 4.1, 14.1, 16.5, 17.1
nist-rmf-measure MEASURE 1–4: metrics, methods, independent assessment, feedback, uncertainty initiative-governance; injection-regression-deck; assistant-metrics Complete families, independent graders, long-horizon and online gates Evaluation Lead · 12.1–12.4, 12.6, 12.7
nist-rmf-manage MANAGE 1–4: prioritize, treat, monitor, communicate, respond initiative-governance; governed-write-confirmation; attributable-work-ledger High-impact policy, SLOs, chaos/restore, ops and game-day evidence Operations Lead · 4.7, 13.1, 13.5–13.7, 18.5

NIST GAI Profile#

Row Requirement/risk Existing controls Remaining gap Phase/task owner
nist-genai-information-integrity Confabulation and information integrity assistant-safety-and-checkers; initiative-governance Grounded/cited conflict-aware answers and long-context provenance floors Sophia PM · 3.5, 12.1, 12.4, 15.3
nist-genai-security-privacy Data privacy and information security workbench-agent-auth; structural-turn-traces; operator-memory-boundary Full data flow, provider posture, exfiltration tests, retention/deletion Privacy Lead · 4.1, 14.1–14.5
nist-genai-human-harm Dangerous/abusive content, bias, human-AI configuration assistant-safety-and-checkers; governed-write-confirmation; initiative-governance Human-trust, oversight, affected-group, blinded-label and acceptance evidence Trust & Safety Lead · 4.1, 4.8, 12.3, 14.6
nist-genai-value-chain Value-chain and component-integration risk bounded-mcp-tool-surface; initiative-governance External registry and model/provider capability, drift, canary, rollback contracts Platform Lead · 4.6, 15.1, 15.2, 15.5, 16.5
nist-genai-ip-media Intellectual-property and provenance risk initiative-governance Eval-data licences and media source/consent/disclosure/redistribution controls Governance Lead · 12.5, 14.7, 17.4
nist-genai-specialized-and-environmental CBRN capability and environmental impact initiative-governance; assistant-metrics Capability-specific admission and verified-outcome resource/cost measurement Agentic AI PM · 12.1, 15.4, 17.1, 17.5

NIST adversarial machine learning#

Row Requirement/risk Existing controls Remaining gap Phase/task owner
nist-aml-evasion NISTAML.022/.025 inference-time evasion injection-regression-deck; assistant-safety-and-checkers Dedicated adaptive, encoded, multimodal, cross-tool statistical gate Security Lead · 4.3, 4.8, 12.1, 12.2
nist-aml-poisoning NISTAML.011/.012/.013/.023/.024/.026/.051 poisoning operator-memory-boundary; untrusted-page-bounds Unified taint/provenance and poisoning tests across memory, retrieval, model/tool/eval supply Security Lead · 3.6, 4.2, 4.6, 9.5, 12.5, 15.5, 16.5
nist-aml-privacy NISTAML.03 privacy compromise structural-turn-traces; operator-memory-boundary Membership/reconstruction/extraction and provider/data-rights coverage Privacy Lead · 4.8, 12.1, 14.2–14.4
nist-aml-generative-prompt-attacks Direct prompting and indirect prompt injection untrusted-page-bounds; injection-regression-deck; assistant-safety-and-checkers Untrusted-data labels, central authority, all-channel red-team evidence Security Lead · 4.2–4.4, 4.8, 17.2

OWASP Agentic Top 10 for 2026#

Row Risk Existing controls Remaining gap Phase/task owner
owasp-asi01 ASI01 Agent Goal Hijack untrusted-page-bounds; injection-regression-deck; assistant-safety-and-checkers End-to-end authority/taint across retrieved, memory, tool, media, and agent content Security Lead · 4.1–4.4, 4.8
owasp-asi02 ASI02 Tool Misuse and Exploitation bounded-mcp-tool-surface; governed-write-confirmation; workbench-agent-auth Central tool policy, dry-run/undo/kill/idempotency, external description attestation Security Lead · 4.4, 4.6, 4.7, 16.5, 16.6
owasp-asi03 ASI03 Identity and Privilege Abuse workbench-agent-auth; attributable-work-ledger Eliminate self-attributed shared mode at admission; tenant/task binding, step-up, rotation, revocation Security Lead · 4.1, 4.4, 4.5, 16.3, 16.4, 16.6
owasp-asi04 ASI04 Agentic Supply Chain Vulnerabilities bounded-mcp-tool-surface; initiative-governance Complete external component registry, drift detection, quarantine and reevaluation Platform Lead · 4.1, 4.6, 15.1, 15.5, 16.5
owasp-asi05 ASI05 Unexpected Code Execution bounded-mcp-tool-surface; governed-write-confirmation Isolation and RCE proof for code, subprocess, browser, DCC, plugins, archives, agents Security Lead · 4.1, 4.5, 4.6, 4.8, 11.3
owasp-asi06 ASI06 Memory & Context Poisoning operator-memory-boundary; untrusted-page-bounds Provenance, confidence, taint, supersession, expiry, conflict, quarantine and harmful-recall gate Iris PM · 4.2, 4.8, 9.2–9.7
owasp-asi07 ASI07 Insecure Inter-Agent Communication workbench-agent-auth; custom-a2a-precursors; iris-signed-agent-precursor Admitted protocol, mutual identity, isolation, replay/downgrade, credential binding, independent test Security Lead · 4.1, 4.4, 4.5, 16.4–16.6
owasp-asi08 ASI08 Cascading Failures attributable-work-ledger; turn-abort-and-sse; assistant-metrics Deadlines, backpressure, breakers, bounded concurrency, fault containment, chaos/soak SRE Lead · 4.1, 4.5, 13.3–13.5, 15.6
owasp-asi09 ASI09 Human-Agent Trust Exploitation governed-write-confirmation; assistant-safety-and-checkers; initiative-governance Confirmation quality, social engineering, trust/fatigue and human outcome gate Trust & Safety Lead · 4.1, 4.7, 4.8, 12.2, 12.3
owasp-asi10 ASI10 Rogue Agents workbench-agent-auth; attributable-work-ledger; turn-abort-and-sse Central authority, durable revoke/kill, drift/containment, no-post-cancel and collusion evals Security Lead · 4.1, 4.4, 4.5, 4.7, 4.8, 13.3

MCP 2025-11-25 — conditional requirements#

Row Requirement/risk Existing controls Remaining gap Phase/task owner
mcp-version-capabilities Initialize, version and capability negotiation bounded-mcp-tool-surface Smoke requests 2024-11-05; no current-version downgrade/conformance matrix Protocol Lead · 16.2, 16.6
mcp-protected-http-auth Optional authorization; protected HTTP discovery/resource/audience/token MUSTs when used workbench-agent-auth; bounded-mcp-tool-surface No protected-resource discovery, OAuth/resource indicator, audience, rotation, isolation or confused-deputy proof Security Lead · 16.2, 16.3, 16.6
mcp-tool-contract-safety Tool capability, schemas, results/errors, untrusted annotations, human-visible control bounded-mcp-tool-surface; governed-write-confirmation Comprehensive schema/result/list-change/malformed/logging and independent tests Protocol Lead · 4.6, 4.7, 16.2, 16.5, 16.6
mcp-progress Unique tokens, monotonic progress, lifetime binding, terminal stop, rate limit turn-abort-and-sse; attributable-work-ledger Local state is not MCP progress and has no token/flood conformance Protocol Lead · 13.1, 13.3, 16.2, 16.6
mcp-cancellation Correct ordinary/task cancel mechanism and race handling turn-abort-and-sse Socket close is not protocol cancellation; no immediate quiescence or post-effect lock SRE Lead · 8.5, 13.3, 16.2, 16.6
mcp-experimental-tasks Experimental capability, lifecycle, get/list/result/cancel, metadata and auth-context isolation attributable-work-ledger; workbench-agent-auth Domain tasks expose no MCP Tasks methods; similarity does not prove semantics or isolation Protocol Lead · 13.3, 16.2, 16.3, 16.6

A2A v1.0.0 — conditional requirements#

Row Requirement/risk Existing controls Remaining gap Phase/task owner
a2a-version-binding-model Version negotiation, canonical model, binding equivalence, errors custom-a2a-precursors; iris-signed-agent-precursor Custom shapes; no admitted binding, A2A-Version, ProtoJSON or independent implementation Protocol Lead · 16.4, 16.6
a2a-agent-card-trust Card availability/declaration, JCS/signatures, verification custom-a2a-precursors; iris-signed-agent-precursor Unwired custom signing; no trust issuer/revocation/cache/change-detection admission Security Lead · 16.4–16.6
a2a-auth-isolation TLS, per-request authentication/authorization, auth-required/HITL, isolation workbench-agent-auth; custom-a2a-precursors Nonempty string is not auth; tenant/task and delegated-credential defenses absent Security Lead · 4.4, 4.5, 16.4, 16.6
a2a-task-stream-artifact-cancel Task/context, ordered streams, artifacts, cancel, subscribe/reconnect custom-a2a-precursors; iris-signed-agent-precursor; attributable-work-ledger No official ordering/context/artifact/durability/cancel/error proof Protocol Lead · 13.3, 16.4, 16.6
a2a-push-security Authenticated push, expected task, idempotency, source verification custom-a2a-precursors Sender-only precursor; receiver, replay, idempotency, isolation, rotation and delivery proof missing Security Lead · 13.3, 16.4, 16.6

AG-UI — optional compatibility decision#

Row Pattern Existing controls Remaining gap Phase/task owner
agui-events Lifecycle/text/tool/error event vocabulary turn-abort-and-sse; bespoke-ui-intents Local frames lack a proven mapping for lifecycle, ordering, ids, version and errors Frontend Platform Lead · 16.1
agui-state State snapshots/deltas and shared-state synchronization bespoke-ui-intents; turn-abort-and-sse No versioned snapshot/delta, conflict, rollback or resume contract Frontend Platform Lead · 8.4, 8.6, 16.1
agui-interrupts-control Human/tool interrupts and cancel/resume patterns governed-write-confirmation; bespoke-ui-intents; turn-abort-and-sse General pause/input/resume/cancel, recovery, mobile parity and ADR remain open Frontend Platform Lead · 8.4–8.6, 16.1

WCAG 2.2 Level A + AA — internally required#

The four principle rows enumerate all 55 current Level A/AA success criteria in the machine record. Criterion 4.1.1, removed in WCAG 2.2, is deliberately not counted. The fifth row keeps the conformance requirements separate from criterion-by-criterion testing.

Row Requirement Existing controls Remaining gap Phase/task owner
wcag-perceivable-aa 20 A/AA Perceivable criteria axe-accessibility-gates; wcag22-technique-inventory Contrast carve-out; qualitative alternatives/media/reflow and affected surfaces unproved Accessibility Lead · 8.7–8.9, 17.3
wcag-operable-aa 20 A/AA Operable criteria axe-accessibility-gates; wcag22-technique-inventory; bespoke-ui-intents Product gaps for 2.4.11/2.5.7/2.5.8 and assistant focus/pointer journeys Accessibility Lead · 8.5, 8.8, 8.9
wcag-understandable-aa 13 A/AA Understandable criteria axe-accessibility-gates; wcag22-technique-inventory; governed-write-confirmation No product instruments/inputs for 3.2.6/3.3.7/3.3.8; generated flow review missing Accessibility Lead · 8.6, 8.8, 8.9
wcag-robust-aa 4.1.2 and 4.1.3 axe-accessibility-gates; bespoke-ui-intents Screen-reader proof for live deltas, tools, errors, tours, confirms and generated UI Accessibility Lead · 8.4, 8.6, 8.8, 8.9
wcag-conformance-aa Level, full-page, complete-process, accessibility-supported, non-interference axe-accessibility-gates; wcag22-technique-inventory Selected scans with disabled rules and no complete assistive-tech/manual process matrix Accessibility Lead · 8.8, 8.9, 18.6

OpenTelemetry — optional, pinned adoption required#

Row Convention/risk Existing controls Remaining gap Phase/task owner
otel-trace-context Core semantic trace/resource/context model structural-turn-traces; in-memory-trace-model; trace-manifest No single runtime context across Eve, models, tools, protocols, ledger and agents SRE Lead · 13.2
otel-genai-agent-spans Development GenAI agent/model spans structural-turn-traces; assistant-metrics; in-memory-trace-model Untagged source; bespoke names; no chosen safe attribute profile SRE Lead · 13.1, 13.2, 15.4
otel-genai-metrics Development token/duration/TTFT metrics assistant-metrics; initiative-governance No convention names/units/histograms, TTFT or cross-plane/cardinality proof SRE Lead · 13.1, 13.2, 13.4, 15.4
otel-mcp-and-sensitive-data Development MCP correlation and sensitive-content posture structural-turn-traces; bounded-mcp-tool-surface; trace-manifest No protocol spans/export; classification, redaction, retention/deletion and pin open Privacy Lead · 13.2, 14.1, 14.3, 14.4

Refresh and claim boundary#

Task 18.3 owns the replacement of this dated record. It must re-fetch primary sources, verify all exact bytes/tags/commits, review NIST errata and protocol normative changes, compare at least two current systems/benchmarks, and assign every newly material gap. A mutable latest page or SDK behavior cannot silently update this record.

This crosswalk closes inventory task 0.4 only. Every implementation or conformance gap above remains owned by its unchecked downstream task. Neither a green verifier nor a similar local abstraction converts those gaps into shipped capability.