P01
Requirements-to-evidence traceability
Decision questionWhich requirement applies to which use, component, user, and failure consequence, and what evidence is sufficient for the next gate?
SVAI Assurance & TEV&V
Define test and evaluation, human oversight, traceability, release gates, failure handling, and model-risk controls.
Consequential AI programmes · Legal, security, and assurance teams · Model and operational risk owners

Concept interface · illustrative valuesGovernance evidence
Safety and governance work now has a visual operating model for risk classification, model review, incident response, and control evidence.
Layer 01
Layer 02
Layer 03
Layer 04
01 · Mission context
Policies alone do not show whether an AI-enabled system is suitable for its intended use. Assurance and TEV&V connect requirements, data, models, operators, failure modes, evaluation results, approvals, monitoring, and incidents into a reviewable lifecycle.
Assurance framing
Governance begins with intended use, affected people, operating conditions, prohibited uses, and the consequence of failure. Terms such as reliable, explainable, secure, or human-supervised are rewritten as system-specific claims with owners, evidence methods, and limits, allowing reviewers to distinguish an obligation from a general aspiration.
P01
Decision questionWhich requirement applies to which use, component, user, and failure consequence, and what evidence is sufficient for the next gate?
P02
Decision questionWho is accountable, what can that person observe and override, and under which conditions must the system stop or revert?
P03
Decision questionWhich representative, edge, adversarial, and failure cases reflect the actual operating context and its consequences?
P04
Decision questionWhich changes trigger reevaluation, reapproval, rollback, incident response, or retirement?
The evaluation plan follows those claims into representative tasks, subgroup and edge cases, adversarial behavior, degraded conditions, and human-factors review. Test coverage is chosen because it informs a release question, not because a benchmark is available, and every result stays tied to the exact data, model, prompt, policy, dependency, and environment evaluated.
Requirements-to-evidence traceability. Broad principles such as reliability, explainability, fairness, or human oversight remain ambiguous until translated into system-specific claims and tests.
Human responsibility and governability. A nominal human-in-the-loop control is weak when the operator lacks time, context, authority, or a usable way to challenge the system.
Evaluation realism. Aggregate benchmark scores can hide subgroup failures, distribution shift, adversarial behaviour, tool misuse, degraded environments, and operator misunderstanding.
Lifecycle change. Models, prompts, data, policies, dependencies, environments, and user behaviour change after the initial review and may invalidate earlier evidence.
02 · Delivery system
Inputs, outputs, maturity, and the evidence boundary travel together. Capability is never separated from the condition under which it can be accepted.
Define intended use, prohibited use, affected parties, operating environment, consequences, accountability, and applicable internal or external requirements.
Output · Context record, risk classification, requirement and control crosswalk, accountable owners, prohibited uses, and evidence obligations.
Translate system claims and failure consequences into datasets, scenarios, metrics, qualitative reviews, operator studies, and repeatable evaluation records.
Output · TEV&V plan, test catalogue, evaluation harness, versioned results, coverage gaps, issue register, and release recommendation.
Exercise misuse, prompt injection, data poisoning indicators, evasion, overreliance, automation bias, uncertain output, component failure, and degraded operation.
Output · Scenario traces, findings, mitigations, residual-risk record, operator-control assessment, and retest plan.
Define the decision forum, evidence pack, release conditions, monitoring signals, change triggers, incident roles, rollback, review cadence, and retirement path.
Output · Release checklist, approval record, model and system card, monitoring plan, change policy, incident playbook, and governance calendar.
Define intended use, prohibited use, affected parties, operating environment, consequences, accountability, and applicable internal or external requirements.
Translate system claims and failure consequences into datasets, scenarios, metrics, qualitative reviews, operator studies, and repeatable evaluation records.
Exercise misuse, prompt injection, data poisoning indicators, evasion, overreliance, automation bias, uncertain output, component failure, and degraded operation.
Define the decision forum, evidence pack, release conditions, monitoring signals, change triggers, incident roles, rollback, review cadence, and retirement path.
03 · System boundary
The assurance architecture joins context and requirements to evaluation records, residual issues, operator controls, and a named release authority. Exceptions and operating limits remain visible at the gate, so a favourable aggregate score cannot obscure an unresolved failure path or transfer risk without an explicit decision.
Reference layers support scoping. Interfaces, owners, and target-system constraints remain subject to validation.
Tie intended use, users, consequences, prohibited uses, requirements, accountable owners, and system claims to a controlled record.
Typical elements · AI system inventory, context profile, risk tier, control crosswalk, claim register, RACI, and data-use decision.
Run versioned task, subgroup, robustness, security, human-factors, and failure evaluations with reproducible inputs and review notes.
Typical elements · Golden dataset, scenario library, simulator, adversarial cases, evaluator rubric, trace store, and coverage report.
Combine technical evidence, unresolved issues, operator controls, security review, residual risk, and decision authority before release.
Typical elements · Evidence pack, exception register, approval workflow, signed release manifest, operating limit, and stop condition.
Connect monitoring, user feedback, incidents, changes, drift, complaints, corrective action, reevaluation, and retirement to named owners.
Typical elements · Telemetry and feedback signals, incident record, change trigger, periodic review, corrective action, rollback, and retirement decision.
Handover establishes which telemetry, feedback, incident, supplier, data, or configuration changes trigger review. It also rehearses containment, rollback, corrective action, and retirement responsibilities. The client retains approval and residual-risk authority while the evidence record gives future reviewers a traceable starting point for reevaluation.
04 · Assurance dossier
The primary story remains calm; profiles, scope, handover evidence, and discovery questions stay available as a structured technical annex.
An AI-enabled decision-support or supervised-autonomy component will operate with real users, degraded conditions, and explicit human responsibility.
An internal assistant retrieves sensitive knowledge, drafts content, and may prepare tool actions for employee review.
A model, runtime, sensor, or policy update may change behaviour across a device fleet operating under resource and connectivity constraints.
Included in this service pattern
Not implied by this page
Handover evidence
Each in-scope claim and control links to an owner, test or review method, result, limitation, issue, and release decision for the named system version.
The approved evaluation set, environment, model, prompt, policy, dependency, and scoring versions can be rerun and compared without reconstructing missing context.
Representative scenarios show that authorised users can recognise system state and uncertainty, challenge or override output, escalate, and invoke the declared stop or fallback path.
A simulated issue or change follows the agreed detection, triage, evidence capture, authority, containment, rollback, communication, corrective action, and reevaluation path.
Discovery questions
Evidence register
References shape requirements and review questions. Inclusion does not imply certification, endorsement, partnership, or approval by the publisher.
05 · Engagement record
Inspectable outputs close the engagement; related services point only to the next bounded step.
Deliverables
Engagement artifacts
04 records per engagement
AI Assurance & TEV&V
Build reviewable AI assurance and governance evidence around the requirements that apply to the actual use case.