Frame
Document mission, system boundaries, business impact, authorized techniques and explicit exclusions.
AI agent security assurance method / v1.0
The six-phase method joins organizational risk, agent-specific attack paths and application security verification in one decision-grade record.
Start with the decision and the owner, map authority and trust boundaries, use the smallest authorized technique that can test the risk, and close only the finding whose changed behavior was actually verified. The output is a bounded record with explicit scope and residual risk — not a generic scanner score.
Document mission, system boundaries, business impact, authorized techniques and explicit exclusions.
Model assets, identities, data, tools, trust boundaries, dependencies and plausible adversaries.
Exercise the highest-risk paths with safe, proportionate and pre-authorized techniques.
Record reproducible proof, affected assets, preconditions, impact and confidence for each finding.
Prioritize structural fixes, compensating controls and ownership against real risk.
Retest the changed behavior and issue a clear record of fixed, accepted and residual risk.
Each control question ends in an observable proof and an accountable decision. This keeps a polished score or framework label from standing in for runtime evidence.
Evidence to inspect: Identity sources, delegated scopes, approval records and revocation evidence.
Evidence to inspect: Tool schemas, API permissions, resource access, egress paths and action logs.
Evidence to inspect: Retrieval provenance, prompt and tool-output boundaries, memory writes and retention rules.
Evidence to inspect: Detection, stop paths, human escalation, rollback, continuity records and recovery tests.
Human, service and agent identities are attributable, authenticated and separately authorized.
Tools, resources, spend and side effects are bounded before execution and revocable during it.
Sources, retention, retrieval and writes are governed so hostile context cannot become durable authority.
Monitoring, human escalation, kill paths and continuity records support containment and recovery.
Use it before an agent gains tools, memory, sensitive data, spend or a material third-party dependency.
Re-run the relevant phases after an architecture, model, workflow, identity or supply-chain change.
Repeat the original proof, check adjacent regressions and record what remains fixed, mitigated, accepted or open.
It is a bounded six-phase method for turning a security question into a decision-grade record: Frame, Map, Challenge, Evidence, Remediate and Verify.
No. The method can include proportionate authorized testing, but it begins with scope, authority and a threat model and uses the smallest technique that can prove or refute a hypothesis. A penetration test is only one possible activity inside an explicitly authorized scope.
It connects identity and authority, tools and side effects, context and memory, and failure and recovery to observable proof: a signed scope and Rules of Engagement, inventories, timestamped logs, reproductions, remediation ownership and a repeatable verification step.
Only after the system owner accepts an auditable scope that names assets, techniques, timing, contacts and stop conditions. Documentation and benign evidence come first, and unexpected sensitive access or instability triggers a stop and incident path.
No. It reports the tested scope, evidence, limitations, findings and residual risk. Framework alignment is a vocabulary for coverage, not a certification or a guarantee about future releases.