The short answer
Use AI to support analysis and keep decisions with the responsible people.
A signed validation plan and evidence packet defining intended use, test population, ground truth, metrics, thresholds, operating envelope, failure injections, human workflow, approved version, rollback, and retest triggers.
Fit before tools
Use this workflow only when the operating conditions fit.
Use it when
- The analytic will influence operational attention or investigation.
- Representative scenes and events can be tested lawfully and safely.
- Ground-truth adjudicators and system owners are available.
- The organization can refuse deployment when thresholds are not met.
Do not use it
- To generalize from a showroom demo or vendor benchmark.
- To test biometric or high-impact use without specific legal, privacy, civil-rights, security, and governance review.
- To optimize only for fewer alerts while ignoring misses.
Prepare first
Define the approved input before opening an AI tool.
Inputs to prepare
- Intended event definition and prohibited uses
- Representative positive and negative fixtures
- Scene and condition matrix
- Ground-truth procedure and adjudicator agreement
- Model, firmware, camera, and configuration versions
- Expected operator workflow
- Metric thresholds and consequence weighting
- Rollback and retest triggers
Keep out of the workflow
- Unconsented or unlawfully collected test data
- Hidden demographic or environmental exclusions
- Changing thresholds after results without documenting the change
- Using alert count alone as accuracy
- Deploying a version different from the tested version
Default rule: If the information boundary is not explicit, do not paste, upload, connect, or transmit the material. Practice with made-up information until the responsible owner approves the tool and information you can use.
Implementation workflow
Complete the work in six reviewable steps.
- 01
Define the event operationally
Write inclusion and exclusion rules that an adjudicator can apply. Avoid labels such as suspicious behavior. Define observable conditions and the exact output the analytic produces.
- 02
Map the operating envelope
List day, night, glare, weather, seasonal change, density, occlusion, camera angle, compression, motion, background, device type, network state, and other conditions that materially affect the intended use.
- 03
Create ground truth
Use qualified reviewers and a documented adjudication process. Measure agreement and resolve ambiguous fixtures separately instead of forcing them into positive or negative labels.
- 04
Select decision-relevant metrics
Measure precision, recall, false alerts per camera-hour, missed valid events, detection latency, availability, and operator correction. Report each condition, not only the overall average.
- 05
Test failure and change
Inject network loss, delayed streams, stale configuration, unavailable inference, time drift, camera movement, partial occlusion, and recovery. Confirm status, logging, fallback, and no silent success.
- 06
Approve a versioned release
Bind acceptance to model, firmware, camera, configuration, rule, and interface versions. Define monitoring, drift indicators, change control, rollback, and the events that require partial or full retest.
Copyable working aid
Use this template, then adapt it to the approved workflow.
The template deliberately exposes missing evidence and preserves human approval. Replace bracketed fields; do not paste prohibited information.
VALIDATION PLAN
Intended use: [observable event and analyst action]
Prohibited uses: [identity, intent, autonomous consequence, etc.]
Tested versions: model | firmware | camera | configuration | interface
CONDITION MATRIX
lighting | weather | density | occlusion | angle | compression | network | seasonal/environmental
GROUND TRUTH
Positive definition: [rule]
Negative definition: [rule]
Ambiguous handling: [rule]
Adjudicators and agreement method: [method]
METRICS AND THRESHOLDS
precision | recall | false alerts/camera-hour | missed events | latency p50/p95 | availability | operator correction
FAILURE TESTS
network loss | stale config | unavailable model | time drift | camera movement | recovery
RELEASE
owner | approval | monitoring | rollback | retest triggersWorked example
A finished example you can check.
These fictional examples and corrections illustrate the review process. They are not records of real incidents or measured model performance.
A perimeter analytic performs well in clear daytime tests but produces repeated false alerts in rain and misses partially occluded events at night.
Condition-level confusion matrices, false alerts per camera-hour, detection latency, and versioned configuration.
ANALYTIC ACCEPTANCE — FULL ENVELOPE NOT ACCEPTED Clear daytime: Favorable performance reported. Rain: Repeated false alerts. Night with partial occlusion: Relevant events missed. Decision: Do not accept the full intended operating envelope. An average cannot resolve condition-specific failures. Evidence gaps: Numeric thresholds and metric values are not reproduced. Preserve condition-level confusion matrices, false alerts per camera-hour, latency, and configuration. Next step: Remediate rain and occlusion failures; repeat affected tests and check for daytime regressions. Daytime-only use requires separate approval, proof against thresholds, and an enforceable restriction. Release record: Name permitted conditions, owner, monitoring measures, rollback, and retest triggers. Failed conditions remain excluded until supported by new evidence.
- Favorable daytime performance is not approval; verify thresholds and enforceable restrictions before any constrained release.
Quality control
Review the artifact and measure whether it improved the work.
Release checklist
- The event definition is observable and adjudicable.
- Positive, negative, and ambiguous fixtures represent the actual operating envelope.
- Metrics include misses, false alerts, latency, availability, and condition-level results.
- Ground-truth creation and reviewer agreement are documented.
- The deployed version exactly matches the accepted evidence.
- Monitoring, rollback, change control, and retest triggers are owned and tested.
Measures worth tracking
- Precision and recall by operating condition
- False alerts per camera-hour
- Missed valid events by consequence class
- Detection latency median and 95th percentile
- Availability and silent-failure rate
- Operator correction rate
- Performance change after environment or version updates
Stop conditions
Treat these outcomes as failures, not minor editing issues.
A strong overall average hides a failed critical condition.
Ground truth is created by one uncalibrated reviewer.
Threshold tuning leaks test outcomes into the final evaluation without a holdout set.
The deployed configuration changes after acceptance without retest.
Escalate instead of improvising: Site-specific risk assessment, emergency action, legal interpretation, employment action, identity determination, biometric use, and consequential access or dispatch decisions require the approved professional and organizational process.
Practical questions
Questions to resolve before operational use.
What is a good accuracy target?
There is no universal target. Thresholds must reflect the intended use, event prevalence, operating conditions, human workflow, and consequence of false positives and false negatives.
Why use false alerts per camera-hour?
It expresses operator burden against exposure time and is often more actionable than a raw false-positive count. It should be reported with recall and missed-event evidence.
When is retesting required?
At minimum when the model, firmware, camera, scene, configuration, interface, environment, or intended use changes materially, or monitoring shows drift outside the accepted range.
Sources and scope
Use authoritative guidance, then apply the organization’s own requirements.
- NIST AI Risk Management Framework Voluntary framework for governing, mapping, measuring, and managing AI risk across the lifecycle.
- NIST AI Resource Center Operational resources for testing, evaluation, verification, and validation of AI systems.
- CISA Guidelines for Secure AI System Development Secure-by-design guidance covering development, deployment, and operation of AI systems.
This guide is vendor-neutral practitioner planning guidance, updated 2026-09-07. It is not a compliance determination, site risk assessment, emergency procedure, or substitute for qualified legal, privacy, cybersecurity, safety, engineering, or security review. Product capabilities and applicable requirements change; verify them with current primary documentation.
Next step
Finished reading? Turn the pattern into practice.
Video search
Read guideAI Readiness and Governance Kit
Open toolProgress is saved only in this browser. Nothing is sent to physicalsecurity.AI.