Expert · Assurance

Validate AI video analytics under the conditions you actually operate.

A vendor benchmark or successful demonstration does not establish site performance. Acceptance testing must represent the intended scene, event, consequence, lighting, weather, density, occlusion, device, network, operator, and degraded conditions.

Published by physicalsecurity.AI · Methodology · Report a correction

The short answer

Use AI to support analysis and keep decisions with the responsible people.

A signed validation plan and evidence packet defining intended use, test population, ground truth, metrics, thresholds, operating envelope, failure injections, human workflow, approved version, rollback, and retest triggers.

Fit before tools

Use this workflow only when the operating conditions fit.

Use it when

  • The analytic will influence operational attention or investigation.
  • Representative scenes and events can be tested lawfully and safely.
  • Ground-truth adjudicators and system owners are available.
  • The organization can refuse deployment when thresholds are not met.

Do not use it

  • To generalize from a showroom demo or vendor benchmark.
  • To test biometric or high-impact use without specific legal, privacy, civil-rights, security, and governance review.
  • To optimize only for fewer alerts while ignoring misses.

Prepare first

Define the approved input before opening an AI tool.

Inputs to prepare

  • Intended event definition and prohibited uses
  • Representative positive and negative fixtures
  • Scene and condition matrix
  • Ground-truth procedure and adjudicator agreement
  • Model, firmware, camera, and configuration versions
  • Expected operator workflow
  • Metric thresholds and consequence weighting
  • Rollback and retest triggers

Keep out of the workflow

  • Unconsented or unlawfully collected test data
  • Hidden demographic or environmental exclusions
  • Changing thresholds after results without documenting the change
  • Using alert count alone as accuracy
  • Deploying a version different from the tested version

Default rule: If the information boundary is not explicit, do not paste, upload, connect, or transmit the material. Practice with made-up information until the responsible owner approves the tool and information you can use.

Implementation workflow

Complete the work in six reviewable steps.

  1. 01

    Define the event operationally

    Write inclusion and exclusion rules that an adjudicator can apply. Avoid labels such as suspicious behavior. Define observable conditions and the exact output the analytic produces.

  2. 02

    Map the operating envelope

    List day, night, glare, weather, seasonal change, density, occlusion, camera angle, compression, motion, background, device type, network state, and other conditions that materially affect the intended use.

  3. 03

    Create ground truth

    Use qualified reviewers and a documented adjudication process. Measure agreement and resolve ambiguous fixtures separately instead of forcing them into positive or negative labels.

  4. 04

    Select decision-relevant metrics

    Measure precision, recall, false alerts per camera-hour, missed valid events, detection latency, availability, and operator correction. Report each condition, not only the overall average.

  5. 05

    Test failure and change

    Inject network loss, delayed streams, stale configuration, unavailable inference, time drift, camera movement, partial occlusion, and recovery. Confirm status, logging, fallback, and no silent success.

  6. 06

    Approve a versioned release

    Bind acceptance to model, firmware, camera, configuration, rule, and interface versions. Define monitoring, drift indicators, change control, rollback, and the events that require partial or full retest.

Copyable working aid

Use this template, then adapt it to the approved workflow.

The template deliberately exposes missing evidence and preserves human approval. Replace bracketed fields; do not paste prohibited information.

VALIDATION PLAN
Intended use: [observable event and analyst action]
Prohibited uses: [identity, intent, autonomous consequence, etc.]
Tested versions: model | firmware | camera | configuration | interface

CONDITION MATRIX
lighting | weather | density | occlusion | angle | compression | network | seasonal/environmental

GROUND TRUTH
Positive definition: [rule]
Negative definition: [rule]
Ambiguous handling: [rule]
Adjudicators and agreement method: [method]

METRICS AND THRESHOLDS
precision | recall | false alerts/camera-hour | missed events | latency p50/p95 | availability | operator correction

FAILURE TESTS
network loss | stale config | unavailable model | time drift | camera movement | recovery

RELEASE
owner | approval | monitoring | rollback | retest triggers

Worked example

A finished example you can check.

These fictional examples and corrections illustrate the review process. They are not records of real incidents or measured model performance.

Situation

A perimeter analytic performs well in clear daytime tests but produces repeated false alerts in rain and misses partially occluded events at night.

Fictional practice input
Condition-level confusion matrices, false alerts per camera-hour, detection latency, and versioned configuration.
Finished example · For practice
ANALYTIC ACCEPTANCE — FULL ENVELOPE NOT ACCEPTED

Clear daytime: Favorable performance reported.
Rain: Repeated false alerts.
Night with partial occlusion: Relevant events missed.

Decision: Do not accept the full intended operating envelope. An average cannot resolve condition-specific failures.
Evidence gaps: Numeric thresholds and metric values are not reproduced. Preserve condition-level confusion matrices, false alerts per camera-hour, latency, and configuration.
Next step: Remediate rain and occlusion failures; repeat affected tests and check for daytime regressions. Daytime-only use requires separate approval, proof against thresholds, and an enforceable restriction.
Release record: Name permitted conditions, owner, monitoring measures, rollback, and retest triggers. Failed conditions remain excluded until supported by new evidence.
Reviewer corrections and checks
  • Favorable daytime performance is not approval; verify thresholds and enforceable restrictions before any constrained release.

Quality control

Review the artifact and measure whether it improved the work.

Release checklist

  • The event definition is observable and adjudicable.
  • Positive, negative, and ambiguous fixtures represent the actual operating envelope.
  • Metrics include misses, false alerts, latency, availability, and condition-level results.
  • Ground-truth creation and reviewer agreement are documented.
  • The deployed version exactly matches the accepted evidence.
  • Monitoring, rollback, change control, and retest triggers are owned and tested.

Measures worth tracking

  • Precision and recall by operating condition
  • False alerts per camera-hour
  • Missed valid events by consequence class
  • Detection latency median and 95th percentile
  • Availability and silent-failure rate
  • Operator correction rate
  • Performance change after environment or version updates

Stop conditions

Treat these outcomes as failures, not minor editing issues.

01

A strong overall average hides a failed critical condition.

02

Ground truth is created by one uncalibrated reviewer.

03

Threshold tuning leaks test outcomes into the final evaluation without a holdout set.

04

The deployed configuration changes after acceptance without retest.

Escalate instead of improvising: Site-specific risk assessment, emergency action, legal interpretation, employment action, identity determination, biometric use, and consequential access or dispatch decisions require the approved professional and organizational process.

Practical questions

Questions to resolve before operational use.

What is a good accuracy target?

There is no universal target. Thresholds must reflect the intended use, event prevalence, operating conditions, human workflow, and consequence of false positives and false negatives.

Why use false alerts per camera-hour?

It expresses operator burden against exposure time and is often more actionable than a raw false-positive count. It should be reported with recall and missed-event evidence.

When is retesting required?

At minimum when the model, firmware, camera, scene, configuration, interface, environment, or intended use changes materially, or monitoring shows drift outside the accepted range.

Sources and scope

Use authoritative guidance, then apply the organization’s own requirements.

This guide is vendor-neutral practitioner planning guidance, updated 2026-09-07. It is not a compliance determination, site risk assessment, emergency procedure, or substitute for qualified legal, privacy, cybersecurity, safety, engineering, or security review. Product capabilities and applicable requirements change; verify them with current primary documentation.

Next step

Finished reading? Turn the pattern into practice.

Put it into practice

AI Readiness and Governance Kit

Open tool

Progress is saved only in this browser. Nothing is sent to physicalsecurity.AI.