Skip to Content
DEEPCRAFT™ Studio 5.14 has arrived. Read more →

Run the tests

This page covers Stage 1 to Stage 3, how to record each run, and how to interpret the results. Complete Stage 0: Plan your test first.

Stage 1: Validate your setup

A correct sensor configuration and physical layout are the basis for reliable results. Check both before you start testing.

Sensor setup
  • Set the sensor configuration exactly as described in the report, including sample rate, channels, bit depth, and any preset.
  • Pass the raw sensor data straight to the model. The Ready Model includes its own preprocessing. Extra filtering, automatic gain control (AGC), resampling, or feature extraction changes the signal the model expects.
Physical layout
  • Fit the intended enclosure and mount.
  • Check that no port cover, bracket, or cable blocks the sensor. Any enclosure wall in front of the sensor can have an impact on the signal.
Compare your signal with the project dataset

The project includes the labeled dataset used to train and test the model. Use it as a reference for a correct signal.

  1. Record raw sensor data on your device. Refer to Real-time Data Collection to know how to stream data into Studio.
  2. Compare your recording with samples from the dataset. Check the level, clipping, noise floor, and, for audio, the spectrogram shape.
  3. If you see a large difference, find and fix the cause before you run any scenarios.

Use the dataset only to check signal quality. To measure detection performance, test with live events on your device.

Log the model output

Print each prediction with a millisecond timestamp so you can align it with your ground truth. For example:

printf("%lu,%s,%.2f\r\n", timestamp_ms, class_label, confidence);

Record the microphone or radar gain and any software gain in your run log.

First live check

Trigger one clear, real target event and check that the model detects it. When it does, move on to your test scenarios. If it does not, check the setup first.

You can use speaker playback as a quick sanity check of the setup. Keep in mind that the model is designed for real sounds, so it may behave differently with played-back or synthetic sounds.

NOTE: Do not use speaker playback to measure performance. Compression, gain settings, and the speaker frequency response change the sound. Models often detect real events better than played-back recordings, so playback results can understate performance.

Optional integration check

Record a short clip on your device and run it through the Ready Model in Studio or in your own framework. Compare the offline output with the device output for the same clip. If the outputs differ, check the integration on the device.

Stage 2: Controlled testing

This stage gives you the core performance results. Run the tests in order, from simple to complex. Move on when the current test meets your acceptance criteria.

  • 2a Baseline: test at close range in realistic conditions with no background noise. This confirms the model detects the target on your hardware. To check the integration, you can also count this test the same way as the report, under conditions similar to the on-device test in the report, and compare the result with the published figures.
  • 2b Placement: test at the intended distance, angle, and mount, with the enclosure fitted.
  • 2c Confusables and false positives: test with the confusables listed in the report and with typical background. Record at least the duration recommended in False positive test duration.

Stage 3: Field conditions

  • 3a Deployment environment: test in the environment where the product will run, with typical background noise and activity.
  • 3b Boundary: test the boundary cases in your test plan. Cover both the stated operating limits and the real situations that may affect performance or cause unknown behavior.

Expect performance to drop beyond operating limits. Record where the drop begins and how the model behaves in each boundary case.

Run one full session

Run at least one full session that matches real product use, such as an overnight sleep session, a full work shift, or a drive. Long sessions confirm that the system stays stable over time, including buffering, timing, and drift.

Repeat and record each run

Repeat key scenarios until you are satisfied that performance meets your product requirements. Run the repeats on different days and, where possible, on more than one device. Consistent results across repeats give you confidence in the measured performance. If results vary between identical runs, check for a change in setup or conditions.

Record the following for each run. These details let you reproduce and compare results.

  • Scenario ID from your test matrix. Refer to the audio and radar test matrices for examples.
  • Model version from the model file name, for example babycry-rm-v2.0.h5, and variant (float32 or int8x8).
  • Board, sensor, and gain settings.
  • Environmental conditions and any changes during the run.
  • Ground truth and model output.
  • Observations, especially anything unexpected.

NOTE: Field tests often capture data that includes people. Inform participants and get consent where recordings may identify them. Store recordings securely and follow your organization’s data protection rules.

Run log

Copy this table and use one row per test run.

DateRun IDScenario IDModel versionVariantBoardSensorGainPlacementBackgroundExpectedObservedGround truthNotes

Interpret results

Count detections with the rule you defined in Define how you count detections.

Convert the false positive rate to daily impact

A rate per hour is easier to judge when you convert it to daily use. Multiply the rate by the number of hours the product is active each day.

For example, 6 false positives per hour × 8 active hours = 48 false positives per day. Compare this number with what your product and users can accept.

Check results below your criteria

If results fall below your acceptance criteria, check the following in order. The first three steps solve most early issues.

  1. Setup: sensor configuration, gain, physical layout, and signal chain.
  2. Product fit: check that your use case and test conditions match what the model is designed for, such as the operating range, sensor placement, and type of event. Refer to Confirm the model fits your product. Conditions outside the model design can explain lower results.
  3. Application: how your application uses the model output. The recommended post-processing is included in the code generated from Studio. If you modified the post-processing or did not generate the code from Studio, apply the post-processing recommended on the Ready Models page. Acting on every raw prediction lowers the measured performance.
  4. Model variant and threshold:
    • Check that the deployed variant (float32 or int8x8) matches the on-device test in the report. The quantized int8x8 variant runs on the PSOC™ Edge NPU and can give slightly different results from float32.
    • Check that the confidence threshold matches the recommended value on the Ready Models page. The threshold is set in Studio post-processing and applied when you regenerate the code. You can adjust it within the acceptable range, but the published figures only apply at the recommended value.
  5. Model: if results are still below your criteria, the model may need data from your conditions. Refer to Customize Ready Model in DEEPCRAFT™ Studio, or contact support.
Decide on the outcome

Base your decision on the results and your acceptance criteria.

  • Accept: performance meets your acceptance criteria across the tested scenarios. Move on to wider product validation.
  • Accept with conditions: performance meets your criteria under specific conditions. Record those conditions.
  • Investigate further: results are below your criteria and the cause is not yet clear. Define the next steps, such as more tests in specific conditions.
  • Adapt the model: performance is below your criteria in some conditions after you check the setup and application. Collect data in those conditions and retrain the model in Studio. After retraining, the published figures no longer apply, so test the new model again.
  • Not suitable: performance is below your criteria and adapting the model is not an option. Contact support before you decide.
Report template

Use the following structure to write up your test.

  • Model: name, version from the model file name, and variant (float32 or int8x8).
  • Setup: board, sensor, gain settings, and code generation path (Graph UX Option 1 or Option 3).
  • Environments: quiet baseline, typical deployment, and noisiest realistic environment.
  • Scenarios: target, non-target, confusable, and boundary.
  • Acceptance criteria and counting rule: as defined before testing.
  • Results: per scenario, counted with your counting rule.
  • Performance by condition: where performance was strong and where it dropped.
  • Daily impact: false positive rate multiplied by daily active hours.
  • Decision: Accept, Accept with conditions, Investigate further, Adapt the model, or Not suitable.
Support

Contact support if a result is unclear, if performance stays below your criteria after you check the setup and product fit, or if you are unsure whether the model fits your product.

Have the following ready:

  • Model name, version, and variant.
  • Board and sensor configuration, including gain.
  • Timestamped output logs for the runs in question.
  • A short recording of the issue.
  • Your ground truth for those runs.

Submit a ticket on the Infineon community forum  DEEPCRAFT™ Studio page.

Checklist

Before testing

  • Report read, and values copied into the test plan.
  • Acceptance criteria and counting rule written down.
  • Sensor configuration matches the model report, with no extra preprocessing.
  • Timestamped model output checked.
  • Ground truth method chosen.
  • Test plan covers target, non-target, confusable, and boundary scenarios.

During testing

  • One variable changed at a time.
  • Planned number of test events completed for each scenario.
  • Key scenarios repeated on different days or devices.
  • Model version, variant, board, and gain recorded for each run.
  • Unexpected outputs recorded with the conditions at the time.

After testing

  • Detections counted with the rule defined in Stage 0.
  • False positive rate converted to daily impact.
  • Results below your criteria checked in order: setup, product fit, application, variant and threshold, then model.
  • Decision written with the setup, results, performance by condition, and reasoning.

Tips for reliable results

  • Validate the setup before you start testing.
  • Test beyond ideal conditions, including distance, background, and boundary cases.
  • Run enough test events to support your conclusions.
  • Set the acceptance criteria and counting rule before you see any results.
  • Test on more than one device where possible.
  • Include confusables in your test plan.
  • Apply the recommended post-processing before you count detections.
  • Use an independent ground truth.
  • Include a time tolerance in your counting rule so late detections count correctly.
  • Change one variable at a time.
  • Use live events for performance results, and keep playback for setup checks.
Last updated on