Skip to Content
DEEPCRAFT™ Studio 5.14 has arrived. Read more →

DEEPCRAFT™ Ready Model field testing guide

A field test measures how a DEEPCRAFT™ Ready Model performs on your hardware, in your product environment, and against your product acceptance criteria. This guide describes what to test, how to set up the test environment, and how to judge the results. The steps apply whether you evaluate in DEEPCRAFT™ Studio Graph UX or in your own framework.

ℹ️

All recommendations, performance information, and other details refer exclusively to the Ready Model included in the project. Any modification to the model — such as retraining, altering its architecture, or changing its parameters — is likely to invalidate these details, including the stated recommendations and performance metrics.

Overview

The evaluation runs in four stages. Complete each stage before you move to the next.

StagePurpose
Stage 0: Plan your testConfirm the model fits your product and design the test scenarios.
Stage 1: Validate your setupConfirm the sensor, physical layout, and signal are correct so your test results are reliable.
Stage 2: Controlled testingMeasure core performance in controlled conditions.
Stage 3: Field conditionsConfirm performance in the deployment environment and in boundary cases.

After testing, refer to Interpret results. For sensor-specific checks and example test matrices, refer to Audio models and Radar models.

Before you start

Read the Ready Models page for the recommended confidence threshold and post-processing for each model. Then read the report for your model, for example the Baby Cry Detection Ready Model Report. Copy the values you need into your test plan, such as the operating range, confidence threshold, post-processing, sensor configuration, performance metrics, minimum SNR, known confusables, and on-device test setup.

The Ready Model is delivered as a Studio project. Refer to Getting started to download the project and to Code generation for Ready Models to generate C code. Ready-to-use deployment code is available in ModusToolbox™:

Key terms

  • Ground truth: an independent record of what actually happened, captured without using the model output.
  • True positive (TP): the target event occurred and the model detected it.
  • False negative (FN): the target event occurred and the model did not detect it.
  • True negative (TN): no target event occurred and the model did not trigger.
  • False positive (FP): the model triggered when no target event occurred.
  • Sensitivity (recall): TP / (TP + FN). The share of real events the model detects.
  • False positive rate: how often the model triggers incorrectly. Choose the unit from how your product meets non-target activity. For models that listen all the time, such as the audio models, use false positives per hour of relevant background. For models with a small detection zone, such as Gesture, false positives only occur when something moves inside the zone. Count them per non-gesture movement in the zone, or as the share of detections that are wrong.
  • Confidence threshold: the confidence level required before the model triggers. Set in Studio post-processing.
  • Model variant: the deployed numeric format, either float32 or int8x8 quantized, where available.
  • Confusable: a non-target event that resembles the target and can cause false positives. For example, laughing or sneezing can trigger the Cough Detection model. Each report lists the confusables tested.
  • Boundary case: a condition at the edge of what the model is designed for. This includes the stated operating limits, such as maximum range or an off-axis angle. It also includes real situations where the model behavior is unknown. For example, a baby crying into a blanket or pillow, or someone rubbing, tapping, or dropping the device.
Last updated on