Skip to Content
DEEPCRAFT™ Studio 5.14 has arrived. Read more →

Stage 0: Plan your test

Complete this stage before you set up any hardware.

Confirm the model fits your product

Compare the report with your product requirements:

  • The operating range and sensor configuration match what your product can provide.
  • The published sensitivity and false positive rate are close to your minimum requirements. Use the figures from the tests closest to real use, such as tests on people or sources not used in training, and on-device or field tests.

Your deployment environment can differ from the conditions tested in the report, including the confusables. The model may still work well in your conditions. Performance can differ from the published figures, so note each difference and add it to your test scenarios.

If you are unsure whether the model fits your product, contact support before you set up hardware. Refer to Support.

Write acceptance criteria

Write down your acceptance criteria before you start testing. Base them on your product requirements. For example:

  • Detects at least 80% of target events in typical conditions.
  • Triggers falsely no more than 5 times per hour in typical background.
  • Responds within 500 ms of the event.

Define how you count detections

Decide how you count detections and false positives before you start testing. The same test can give very different results with different counting rules. Use the following approach to design your rule:

  1. Start from the product question. For example, “Is the parent alerted when the baby starts crying?”
  2. Define a detection. Count detections after the post-processing recommended in the report. Set a time tolerance from how late a detection can be and still be useful to your product, for example a detection within a set time after the event starts. Record the value in your counting rule.
  3. Define a false positive the same way. For example, count one false positive per episode of triggers, and state the rate per hour of relevant background, or per non-gesture movement in the zone for Gesture.
  4. Report more than one view where it helps. A per-episode result shows whether the product does its job. A per-event result shows how many individual events the model catches.

Write the rule down with your acceptance criteria. The report for each model shows an example of this approach. You can start from the report’s method and adapt it to your product.

Design your scenarios

Include all four scenario types in your test plan. Together, they give a complete picture of performance.

Scenario typePurposeSource
TargetThe event the model should detect.Product use cases.
Non-targetEveryday background the model should ignore.Deployment environment.
ConfusableNon-target events that resemble the target.Negative data and known false positives in the report.
BoundaryEvents at the stated operating limits, and real situations where the model behavior is unknown.Operating range in the report, product use cases, and product risk analysis.

Refer to the audio and radar test matrices for examples.

Control variables and ground truth

  • Change one variable at a time. This way, you know which change caused each result.
  • Choose an independent ground truth method, such as a human observer, a second recorder, or a reference device. Document the method in your test plan.

Sample size

A test event is one occurrence of what you test, such as one cough, one gesture, or one confusable sound. For models that detect periods, it is one episode, such as a snoring period or a crying session. Count test events the same way you count detections.

We recommend at least 50 test events for each detection scenario with a pass/fail criterion. At 50 test events, the measured sensitivity is within about ±10 percentage points of the true value. Size false positive scenarios by duration instead. Refer to False positive test duration. Scenarios that only record how the model behaves do not need a fixed number of test events. Expand testing beyond the minimum where you can. More test events, more conditions, and more test units give a clearer picture of true performance.

Decide the number of test events before you start. Use the performance figures in the report to judge how many test events you need for an accurate result.

Test with the range of people or sources your product will meet. The number depends on how certain you need to be. About 10 people or sources is a good start, around half the number of people the report tested who were not part of the training data. Add more if your product needs a stronger generalization claim. Record the count and the range in your test plan.

If you find a problem with the test setup, discard the affected data, fix the problem, and record the scenario again. Note the cause in your run log. To save time, run the tests in smaller batches and review the results after each batch. This way you find setup problems early and apply fixes before the next batch.

False positive test duration

Relevant background is normal activity in your deployment environment, with the product on and no target event. Cover all the common scenarios in which your product will operate.

We recommend at least 3 hours of relevant background per environment. Size the rest of the test from your false positive criterion. If you allow one false positive per period, test for at least three times that period with no false positives.

Your criterionTest duration with no false positives
At most 1 per hour3 hours
At most 1 per night (8 active hours)3 nights (24 hours)
At most 1 per day, always on3 days (72 hours)
At most 1 per week, always on3 weeks

If false positives occur, record each one with the conditions at the time. Divide the count by the test hours to get the measured rate. If the rate is close to your criterion, extend the test.

For models with a small detection zone, such as Gesture, false positives depend on how often something moves inside the zone. Size the test by the number of non-gesture movements in the zone, using the same numbers as in Sample size. Add a long session only if your product sees regular movement near the sensor.

Run longer tests where you can. Full-day or overnight sessions in the deployment environment show how often false positives occur in real use. When we validate models for customer production, we sometimes test with several typical users for one week each. This covers a wide range of real scenarios and activities.

Safety-critical use

If the model output feeds a safety-critical function, test with more rigor:

  • Run more test events to increase statistical confidence.
  • Add fail-safe behavior at the application level.
  • Contact support early to confirm the model suits the intended use.

Plan for several evaluation rounds before a safety-critical deployment.

Last updated on