AI Security Review Evidence: What You Must Prove in 6 Months
Most AI security reviews produce polished dashboards and weak memories. Six months later, when an auditor, customer, or incident team asks what changed, many teams can show charts but cannot prove control effectiveness, model lineage, or blast-radius reduction. This post explains what an AI security review should leave behind: durable evidence you can test, replay, and defend.
Nesqual Tech AI
A clean dashboard can hide a weak security program. The hard test comes six months later, when a customer asks whether their data touched a fine-tuning job, an auditor wants proof of prompt-injection controls, or your incident team needs to reconstruct which model version generated a harmful action.
If your AI security review ends with screenshots, maturity scores, and a slide that says risk reduced, you bought theater. A useful review leaves evidence: artifacts you can replay, verify, and map to decisions, controls, and outcomes.
Start with the six-month question, not the demo-day dashboard
A strong AI security review begins by asking a blunt question: What will you need to prove six months from now, under pressure? In 2026, that usually means proving five things:
- Which data entered which AI workflow
- Which model, policy, and tool permissions were active at the time
- Which controls blocked, allowed, or escalated a request
- How quickly you could detect and contain misuse
- Whether security posture improved with measurable evidence
That changes the review scope. You stop collecting vanity metrics like total prompts per day and start collecting durable records like signed model manifests, policy decision logs, and red-team replay results.
The difference between observability and evidence
Observability helps you see what is happening now. Evidence helps you prove what happened then.
For example, a dashboard may show that your retrieval-augmented generation stack had 99.3% uptime in Q1 2026. Useful, but not enough. Evidence would show:
- The exact embedding model version used on March 14
- The document access policy in force for finance users
- The tool invocation policy that denied outbound HTTP calls
- The trace ID connecting the user request, retrieval step, model response, and human approval action
That level of proof matters in real scenarios. A Fortune 100 procurement team reviewing your AI assistant will ask whether tenant data can cross retrieval boundaries. If your answer is a dashboard tile saying cross-tenant leakage: low, expect follow-up questions. If you can produce access test results, policy logs, and isolation architecture, the conversation changes.
Define the evidence package your review must produce
Treat the output of an AI security review as an evidence package, not a report. The package should be versioned, queryable, and tied to controls.
1. Model and pipeline lineage you can replay
You should be able to reconstruct a production decision path for a specific time window. At minimum, capture:
- Model provider, model family, and version or snapshot ID
- System prompt and policy prompt hashes
- Retrieval index version and document corpus snapshot
- Tool catalog and permission set
- Safety policy version
- Human approval requirements by action class
A practical pattern is to store these as signed manifests in Git and mirror runtime references into your SIEM or data lake.
apiVersion: ai.nesqual.io/v1
kind: ModelExecutionManifest
metadata:
service: contract-review-assistant
environment: prod
timestamp: "2026-03-14T10:22:31Z"
spec:
model:
provider: openai
name: gpt-4.2-enterprise
snapshot: "2026-03-01"
prompts:
systemHash: "sha256:8fd1..."
policyHash: "sha256:c91a..."
retrieval:
indexVersion: "contracts-index-v118"
corpusSnapshot: "s3://legal-corpus/snapshots/2026-03-10"
tools:
- name: crm_lookup
scope: read_only
- name: send_email
scope: disabled
approvals:
external_send: required
signing:
cosignSignature: "MEUCIQDn..."
If a review does not produce this level of lineage for high-risk workflows, it is incomplete.
2. Policy decision logs, not just final outcomes
Many teams log only the user prompt and model response. That is not enough for an AI security review. You need the policy decisions around the request.
For a support copilot with tool use, a useful record includes:
- User identity and risk tier
- Data classification of retrieved content
- Prompt-injection detector score
- Tool call requested by the model
- Policy engine decision: allow, deny, redact, escalate
- Reason code and policy version
This is where dashboards often fail. They aggregate blocked prompts: 1,284 but cannot explain why specific prompts were blocked or whether the same policy would still fire after a model update.
{
"trace_id": "9d4a7f2e1b",
"timestamp": "2026-03-14T10:22:32Z",
"user": {"id": "u-4471", "role": "support_l2"},
"request": {"classification": "internal", "channel": "web"},
"detectors": {
"prompt_injection_score": 0.91,
"pii_score": 0.08
},
"tool_request": {"name": "refund_api", "action": "approve_refund"},
"policy": {
"engine": "opa-1.4",
"version": "policy-bundle-2026.03.12",
"decision": "deny",
"reason_code": "HUMAN_APPROVAL_REQUIRED"
}
}
3. Red-team replay artifacts with pass/fail thresholds
An AI security review should leave behind a replayable adversarial test suite. Not a PDF summary. A runnable set of prompts, tool-use attempts, retrieval poisoning cases, and expected outcomes.
In 2026, mature teams run these suites on every major model, prompt, or policy change. A realistic benchmark for a medium-sized enterprise assistant is:
200-500adversarial test cases per critical workflow95%+pass rate required for release0tolerance for high-severity failures such as cross-tenant retrieval or unapproved external actions- Under
30 minutesto execute the suite in CI for gating changes
#!/usr/bin/env bash
set -euo pipefail
SUITE=./redteam/customer-support-2026q2.yaml
MODEL_SNAPSHOT=${1:-2026-03-01}
POLICY_BUNDLE=${2:-policy-bundle-2026.03.12}
python run_redteam.py \
--suite "$SUITE" \
--model-snapshot "$MODEL_SNAPSHOT" \
--policy-bundle "$POLICY_BUNDLE" \
--fail-on-severity high \
--min-pass-rate 0.95 \
--export-junit ./artifacts/redteam-results.xml \
--export-json ./artifacts/redteam-results.json
The six-month proof is simple: can you rerun the same suite against the current stack and explain the delta?
Measure controls by blast-radius reduction, not dashboard activity
The point of an AI security review is not to prove you deployed controls. It is to prove those controls reduced risk in measurable ways.
Use outcome metrics that survive executive scrutiny
Good metrics answer: What got harder for an attacker or easier for responders?
Examples that work:
- Cross-tenant retrieval exposure reduced from
1 in 2,000test queries to0 in 50,000 - Median time to revoke a compromised model API key cut from
42 minutesto6 minutes - Tool invocation requiring human approval increased coverage from
61%to98%for money-moving actions - Sensitive data redaction precision/recall improved from
0.89/0.83to0.96/0.92 - Forensic reconstruction time for a harmful response dropped from
9 hoursto35 minutes
Weak metrics, by contrast, include total blocked prompts, number of dashboards, or generic maturity scores detached from incidents and tests.
Example: agentic procurement assistant
Consider a procurement assistant that can read contracts, query ERP data, and draft supplier emails. The original review found three issues:
- The model could request unrestricted web fetches
- Contract retrieval lacked matter-based access controls
- Email drafting logs did not preserve policy decisions
A useful remediation plan produced evidence, not promises:
- Tool broker restricted outbound destinations to an allowlist of
12domains - Retrieval added ABAC rules using matter ID and legal entity
- Email actions required approval for external recipients and attached the policy decision log to each trace
Six months later, the team could prove:
0/18,400replay tests achieved unauthorized external fetch- Unauthorized contract retrieval attempts were denied in
99.98%of test cases; the remaining0.02%were low-severity metadata leaks fixed in one sprint - Incident review time for disputed supplier messages fell from
4.5 hoursto28 minutes
That is what a review should buy you.
Build evidence into the architecture, not the audit folder
If evidence collection is manual, it will decay. The architecture should emit it by default.
Reference architecture for durable AI evidence
[User/App]
|
v
[API Gateway] --> [Identity + Device Context]
|
v
[AI Orchestrator] --> [Policy Engine (OPA/Cedar)]
| | |
| | +--> decision logs -> [SIEM/Data Lake]
| +--> prompt/model metadata -> [Trace Store]
|
+--> [RAG Service] --> [Vector DB + ACL filter logs]
|
+--> [Tool Broker] --> [Scoped tools + approval service]
|
+--> [Model Endpoint]
|
+--> [Immutable artifact store: manifests, evals, signatures]
Three design choices matter here:
- Central policy enforcement: Do not bury access logic inside prompts. Put authorization and tool controls in a policy engine you can version and test.
- Immutable artifacts: Store manifests, evaluation results, and approval records in append-only or object-locked storage for at least your audit retention window.
- Trace correlation: Every request needs a stable trace ID across gateway, orchestrator, retrieval, tool use, and response logging.
A practical policy example
For tool-use governance, a policy should be explicit enough to explain a denial later.
package ai.tools
default allow = false
allow if {
input.user.role == "finance_manager"
input.tool.name == "invoice_export"
input.tool.scope == "read_only"
input.request.data_classification != "restricted"
}
require_human_approval if {
input.tool.name == "payment_release"
}
This is better than a prompt instruction saying, Only release payments when appropriate. One is enforceable evidence. The other is wishful thinking.
Common Pitfalls
Mistaking provider attestations for your own proof
Your model provider may offer SOC 2, ISO 27001, or regional processing guarantees. Useful, but incomplete. They do not prove your retrieval boundaries, your tool permissions, or your tenant isolation.
Avoid this by mapping provider controls to your shared-responsibility model. Then produce your own test evidence for the parts you own.
Logging too little or too much
Some teams log only prompts and responses. Others dump full prompts, retrieved documents, and user data into logs, creating a second security problem.
Log structured metadata by default, and store sensitive payloads selectively with retention and access controls. A common pattern in 2026 is 30 days hot searchable metadata and 180-365 days cold encrypted artifacts for regulated workflows.
Treating red teaming as a one-time event
A single pre-launch exercise is not an AI security review strategy. Model snapshots change, prompts drift, and new tools expand the attack surface.
Tie red-team replay to release gates. For high-risk assistants, rerun on every model snapshot change and every policy bundle change.
Hiding control logic in prompts
If your most important security control lives in a system prompt, you cannot reliably prove enforcement. Prompts are guidance, not hard boundaries.
Move authorization, data filtering, rate limits, and action approvals into services and policies outside the model.
Ignoring responder usability
Evidence that takes eight hours to assemble is not operational evidence. During an incident, your team needs one trace, one timeline, and one set of linked artifacts.
Set an internal SLO: for any high-severity AI event, responders should reconstruct the decision path in under 60 minutes. If you cannot meet that, improve trace correlation before buying another dashboard.
Key Takeaways
- Define your AI security review by what you must prove six months later: lineage, policy decisions, control effectiveness, and response speed.
- Require an evidence package, not a slide deck: signed manifests, decision logs, replayable red-team suites, and immutable artifacts.
- Measure blast-radius reduction with concrete numbers such as retrieval leakage rates, approval coverage, and forensic reconstruction time.
- Put enforcement in policy engines and tool brokers, not prompts, so you can version, test, and explain decisions.
- Make trace correlation mandatory across gateway, orchestrator, RAG, tools, and model endpoints.
- This week, pick one high-risk AI workflow and test whether you can reconstruct a single production decision path from user request to final action in under an hour.
This article was written by an AI system and published pending human review. Verify anything you intend to act on.
Written by
Nesqual Tech AI
Nesqual Tech
Have a project in mind?
Get an instant AI price estimate for it, or talk directly to our team.
One email a month on what we learn building with AI