Skip to Content
DEEPCRAFT™ Studio 5.14 has arrived. Read more →

DEEPCRAFT™ Gesture Classification Ready Model Report

In this document, we describe the DEEPCRAFT™ Ready Model for Gesture Classification, a radar-based AI model developed by Imagimob, an Infineon Technologies company. This model detects and classifies when a person is performing the following five hand gestures in front of a radar sensor:

  • Push - Open palm moving horizontally in parallel to the ground from the user’s body towards the radar sensor

  • Orthogonal Swipes

    • Swipe Left - open palm moving horizontally in parallel to the ground from right side to left side of the radar sensor
    • Swipe Right - open palm moving horizontally in parallel to the ground from left side to right side of the radar sensor
    • Swipe Up - open palm moving vertically, perpendicular to the ground, from the bottom side to the top side of the radar sensor
    • Swipe Down - open palm moving vertically, perpendicular to the ground from the top side to the bottom side of the radar sensor

We provide details about the technical specifications of this machine learning model, its performance in common scenarios, and various test results for the model including the real-time testing on an Infineon PSOC™ 6 board connected to an Infineon XENSIV™ 60 GHz radar sensor.

All recommendations, performance information, and other details mentioned in this report refer exclusively to the Ready Model included in the project. Any modification to the model — such as retraining, altering its architecture, or changing its parameters — is likely to invalidate these details, including the stated recommendations and performance metrics.

Model Specification

Model Overview

The DEEPCRAFT™ Ready Model for Gesture Classification is designed to detect in real-time the above-mentioned five gestures with the purpose of enhancing and/or assisting human-machine interaction and experience. Such a model is developed with the intent to run, for instance, on a monitor or on any device connected to a monitor in order to provide contactless control. The model’s predictions will be the input of an application used to interact with any device.

The model will not distinguish between fast or slow gestures, diagonal swipes, right-handed versus left-handed gestures.

Expected Performance

The aim of this model is to detect the user’s hand gestures, defined as push and swipe movements, performed with one hand. The hand movement is expected to:

  • last between one and two seconds
  • occur within 10 cm and 70 cm from the radar sensor
  • be within an angle of 10 degrees in the field of view of the radar sensor

The model is expected to detect 92% of gesture events for a standing or sitting user. A gesture is expected to be misclassified or to trigger one of the other gestures fewer than five times out of every 100 gesture events. This performance can significantly increase with the user’s practice and experience in using the model. Besides different distances, angles and velocities, the model is expected to be robust against different hand shapes, arm lengths and user heights, as well as hand and body movements different from the actual gestures.

Detailed results are provided in Appendix II - Detailed On-Device Test Results.

Operations

The user is expected to stand or sit in front of the radar sensor without any obstacle in between. Missed detections or unwanted triggers may occur when the five gestures are performed in the following ways:

  • faster or slower than 1-2 seconds
  • outside the 10-70 cm range from the radar sensor
  • more than 10 degrees from the radar sensor’s field of view
  • with obstacles between the user’s hand and the radar sensor
  • when holding an object, especially one with moving parts
  • when wearing clothes with long, hanging sleeves
  • when wearing a hanging bracelet or a gadget close to the hand

In addition to the relative positioning of the radar sensor and the user, the accuracy of the gesture model strongly depends on the way the user moves their hand and arm. Any movement which resembles a push or a swipe will be detected as such. Diagonal hand movements, as well as stopping the gesture halfway, may result in unexpected model output.

This model detects any of the five gestures in its vicinity without distinguishing the person who performed it; this means that the model may detect the gestures of someone else close by. Also, if one or more people are moving in between the radar sensor and the user, the model’s output is expected to be less accurate. On the other hand, people moving or walking behind the user might have no impact on model performance.

Performance degradation may occur because of complementary movements that may resemble one of the five gestures. For instance, a swipe down can be triggered after a swipe up or a push, because the user will naturally bring their hand down. Similarly, a swipe left or right can be preceded by a push because of the hand movement towards the radar sensor that the user has to make to engage with the radar sensor.

The model is in general robust against body movements as long as they are not the same speed as a gesture, although one of the gestures might be detected when a user is sitting, standing, bending down quickly, jumping, or running.

It is recommended that this model should be used to enhance the user’s overall experience in controlling a device/system rather than focusing on each single gesture classification.

Model Tech Specs and Deployment

The DEEPCRAFT™ Ready Model for Gesture Classification is able to detect a gesture from radar sensor data, parsed on-device, with the following characteristics:

  • Features: range, velocity, azimuth, elevation, magnitude
  • Features sample rate: 33 Hz

For more details about the radar configuration see Appendix I - Radar Configuration.

Target board                 Model versionlibrary - Flash (KiB)library - RAM (KiB)ml-middleware - Flash (KiB)ml-middleware - RAM (KiB)tflite-micro - Flash (KiB)tflite-micro - RAM (KiB)
AnyANSI C9944.0267.30N/AN/AN/AN/A
PSOC™ 6float3253.56132.423.144.52259.690.09
PSOC™ 6int8x825.29104.003.144.52259.690.09
PSOC™ Edgefloat32136.87132.434.274.62264.320.12
PSOC™ Edgeint8x8106.80102.204.274.62264.320.12

Memory-footprint calculation details are provided in Appendix III - Memory Footprint Calculation.

For on-device testing and inference time calculation, the model uses C code generated for PSOC™ 6. It performs preprocessing in C and uses the ModusToolbox ML runtime to execute an int8x8 quantized TensorFlow Lite for Microcontrollers model. The inference time, excluding the radar data parsing time, is about 8 ms when running on a XENSIV™ KIT CSK BGT60TR13C with a PSOC™ 6 clocked at 150 MHz and a 60 GHz radar sensor. The model outputs a prediction every 91 ms.

Data Set

This DEEPCRAFT™ Ready Model has been built using internally collected positive and negative radar sensor data. The positive data consists of gestures from several individuals with different heights performing gestures at different distances and angles. Most of the participants were right-handed. The negative data consists of different random movements in front of the radar as well as sessions without any movement.

For the collection of all the data, the radar sensor was mounted about 1 m from the ground. The person performing the gestures was standing mostly in front of the radar sensor (azimuth angle zero degrees) and within the range 0.5-1 m from the radar sensor.

Model Evaluation

In this section we present the evaluation of the model. The model was evaluated in two ways, each described in one of the following subsections:

  • Validation set results: performance on the validation set used during model development.
  • On-device test results: real-time performance with the model deployed on a XENSIV™ Connected Sensor Kit board with 20 people not part of the model training.

We evaluate model performance at two levels: file level and sliding-window level. At file level, each data file counts as one result: a positive file is classified as detected if the model triggers at least once anywhere in the file, and a negative file is classified as a false positive if the model triggers at least once anywhere in the file. At sliding-window level, each model input window and resulting prediction counts as one result.

Validation Set Results

The overall accuracy of the model on the Validation set is 96.6% when the output of the model with the highest confidence level or probability is counted as a trigger. The table below is the so-called confusion matrix of the model which shows its performance on the Validation set in more detail. The percentages are normalised within each actual class column, so the diagonal values are recall values for the corresponding class:

  • Top Left Value (Negative Recall / True-Negative Rate) : actual non-gesture data predicted as non-gesture
  • Other Values in “Unlabelled” Column (False-Positive Rates): actual non-gesture data predicted as a gesture
  • Other Values in “Unlabelled” Row (False-Negative Rates): actual gesture data predicted as non-gesture
  • Other Values in Diagonal (Per-Gesture Recall / True-Positive Rates): actual gesture data predicted as the correct gesture
  • All Other Off-Diagonal Values (Misclassification Rates): actual gesture data predicted as another gesture

This report uses the sliding-window-level evaluation rather than a file-level matrix. Because it evaluates short windows, it can overestimate False Negatives, so the True Positives percentages below set a roughly lower limit for the model’s real True Positives.

Sliding-Window-Based Confusion MatrixActual non-gestureActual pushActual swipe downActual swipe leftActual swipe rightActual swipe up
Predicted non-gesture97.24 %6.07 %6.59 %4.96 %8.41 %6.13 %
Predicted push0.76 %91.44 %0.39 %0.25 %0.35 %0.10 %
Predicted swipe down0.78 %1.41 %91.22 %0.35 %0.56 %1.72 %
Predicted swipe left0.57 %0.65 %0.36 %93.59 %1.77 %1.32 %
Predicted swipe right0.30 %0.22 %0.47 %0.50 %88.61 %0.96 %
Predicted swipe up0.35 %0.22 %0.97 %0.35 %0.30 %89.77 %

In this case, another way to interpret the high value for the False Negatives is that the model may miss gestures which are performed further from and/or not aligned with the radar sensor as well as gestures which are executed with incorrect movements. Typically, this happens more when the user interacts with the radar sensor for the first time.

When looking at the True Positives (diagonal values), we see that the model has the highest detection recall with Swipe Left (93.6%) and lowest with Swipe Right (88.6%). Overall, this means that it is easier to trigger a swipe when moving the hand from right to left than the opposite. This 5% difference or asymmetry can be a consequence of the fact that most of the people who contributed to model building with their data are right-handed. Using the right hand to perform a Swipe Left can indeed be easier, more consistent and more accurate than a Swipe Right, which, has more variation and requires more training data.

On the other hand, we notice that most of the gesture-to-gesture misclassification rates are below 1% except for four cases: Push detected as Swipe Down (1.41%), Swipe Up detected as Swipe Down (1.72%) or Swipe Left (1.32%), and Swipe Right detected as Swipe Left (1.77%). In all these cases, the misclassifications may be triggered by complementary movements that the users do in order to perform a gesture. For instance, the user will put his hand down after a Push or a Swipe Up, or execute a Swipe Right after moving the hand to the left side of the radar sensor.

Overall, by summing up the misclassification rates per gesture, Swipe Up and Swipe Right have the highest misclassification rates with a total of 3.0% and 4.1%, respectively. On the other hand, Swipe Left is expected to have the highest precision since it has in total a misclassification rate equal to 1.5%.

On-Device Test Results

We performed on-device testing with 20 people that were not part of the model training. For this test, we deployed the model on a XENSIV™ Connected Sensor Kit board with PSOC™ 6 mounting a BGT60TR13C radar sensor. Post-processing combines a confidence threshold with multiple consecutive positive predictions to confirm a gesture and applies debouncing to prevent a single gesture from producing multiple confirmed events. We summarise the results in the table below. The detailed results of this testing are reported in the table in [Appendix II - Detailed On-Device Test Results](Appendix II - Detailed On-Device Test Results).

GestureRecallPrecision
Push88.1 %93.3 %
Swipe Down91.0 %96.1 %
Swipe Left93.8 %97.5 %
Swipe Right91.7 %95.0 %
Swipe Up94.4 %94.4 %
TOTAL91.8 %95.3 %

The participants had different heights ranging from 160 cm to 190 cm, and they were all right-handed. Most importantly, for the majority of the participants it was the first time they interacted with a radar gesture model.

To test the model on-device, we equally divided the participants into groups depending on the positioning of the radar sensor and the positioning of the person’s body and height. In particular, the test has been run with the following conditions:

  • Person standing or sitting
  • Radar mounted at about 1 m from the ground
  • For sitting, the radar sensor placed on a table
  • Two radar sensor-to-person’s body distances: about 1 m and 0.5 m
  • Gestures performed in front of the radar (zero degrees azimuth angle)

Each participant performed each gesture several times in a row, with and without a person moving close by (see details in tables in Appendix II).

The total recall value for this test was 91.8%, and the total precision was 95.3%, meaning that 4.7% of predicted gesture events were incorrect. By looking at the performance per gesture, this test shows that the model is more likely to detect Swipe Up and Swipe Left events than Swipe Down and Push. The overall detection increases if the person is sitting close to the radar (see detailed results in Appendix II). In general, the further the user stands or sits, the higher the number of missed gestures.

The performance of the model running on-device also shows that Swipe Left (93.8%) is more accurate than Swipe Right (91.7%), in agreement with the results on the validation set.

When analysing the misclassifications, only a few gestures were actually detected as another gesture, most of the recorded False Positives in the test occurred because of complementary movements. Similarly to what is discussed in the Validation Set Results section, this happened most often in Swipe Up events which were followed by a Swipe Down gesture because the participants were naturally lowering their hand and arm down. Also, we observed that in some cases a Swipe Left or Right could be preceded by a Push, as some participants were moving their hand towards the radar to engage with it in a way similar to a Push gesture. Using additional post-processing that filters out unwanted triggers may result in a more precise detection and a lower learning effort for the user. An example is applying a long enough delay after each detected gesture to block the additional triggers.

Most importantly, during this test, we clearly observed that the more each participant performed the gestures, the fewer False Negatives and False Positives occurred. In particular, the Push gesture was the most difficult to trigger and to learn. This is one of the main reasons why the recall for Push is only 88.1%.

Besides performing the five gestures, we also tested the robustness of the model against non-gestures and other body movements. No False Positives were triggered for instance for the following movements:

  • Hand waving
  • Hand/finger rotation
  • Hands moving while talking
  • Clapping
  • Moving around the radar
  • Walking to/away from the radar

In some cases, however, misclassifications could occur when a hand or body movement was fast enough to resemble a gesture. Here are some examples:

  • Sitting on a chair or bending to the ground in a relatively fast way in front of the radar sensor could trigger Push
  • Jumping from the right to left (or vice versa) side of the radar sensor could be detected as Swipe Left (or Swipe Right)
  • Quickly putting both hands on the table, where the radar sensor was placed, when standing up from a chair could trigger one of the gestures
Summary

The evaluation shows that the model reliably detects the five supported gestures, with performance depending on the user’s movement, position, and experience:

  • On the validation set, the model achieved 96.6% overall sliding-window accuracy. Per-gesture recall ranged from 88.6% for Swipe Right to 93.6% for Swipe Left, and most gesture-to-gesture misclassification rates were below 1%.
  • In on-device testing with 20 people who were not part of the model training, the model achieved 91.8% overall recall and 95.3% overall precision. Swipe Left had the highest recall at 93.8%, while Push had the lowest at 88.1%.
  • The model produced no false positives for several ordinary movements, including hand waving, hand rotation, clapping, and walking to or away from the radar. Fast movements that resemble a gesture may still trigger a detection or misclassification.

Detection generally improves when the user is closer to the radar and becomes more familiar with the gestures. Missed detections and unwanted triggers are more likely when gestures are performed too quickly or slowly, outside the recommended distance or angle, or with complementary movements such as lowering the hand after a Push or Swipe Up.

Appendix I - Radar Configuration

The radar needs to be configured with the following parameters:

  • Start Frequency: 58.5 GHz
  • End Frequency: 62.5 GHz
  • Samples per chirp: 64
  • Chirps per frame: 32
  • Number of Receivers: 3
  • Number of Transmitters: 1
  • Sample rate: 2 MHz
  • Chirp repetition time: 0.299787 ms
  • Frame repetition time: 30.0446 ms

Appendix II - Detailed On-Device Test Results

In the tables below, we report true positives (TP) and false positives (FP) for the on-device testing with 20 people as well as values for the Recall and Precision of the model in different scenarios. For the TPs we also write the total number of expected gestures per each gesture, which is 20 for most of the participants.

Person-Radar Distance: about 0.5 m (close case), about 1 m (far case).

Standing Close

PersSwipe Right TPSwipe Right FPSwipe Left TPSwipe Left FPSwipe Up TPSwipe Up FPSwipe Down TPSwipe Down FPPush TPPush FP
120/20020/20019/20420/20020/200
210/20017/20219/20116/20014/205
320/20020/20018/20320/20016/201
417/20319/20019/20520/20318/206

Standing Far

PersSwipe Right TPSwipe Right FPSwipe Left TPSwipe Left FPSwipe Up TPSwipe Up FPSwipe Down TPSwipe Down FPPush TPPush FP
120/20320/20319/20020/20018/200
219/20015/20018/20416/20020/200
320/20320/20020/20120/20017/200
414/20014/20014/20011/20215/203

Sitting Close

PersSwipe Right TPSwipe Right FPSwipe Left TPSwipe Left FPSwipe Up TPSwipe Up FPSwipe Down TPSwipe Down FPPush TPPush FP
119/20020/20020/20019/20020/200
218/20120/20019/20019/20019/202
317/20219/20020/20020/20020/200
420/20020/20020/20018/20016/200
515/15013/13016/16016/17117/170
610/1006/10110/1028/1008/100

Sitting Far

PersSwipe Right TPSwipe Right FPSwipe Left TPSwipe Left FPSwipe Up TPSwipe Up FPSwipe Down TPSwipe Down FPPush TPPush FP
119/20019/20119/20019/20518/204
219/20320/20118/20120/20216/201
318/20019/20119/20016/20015/200
419/20019/20018/20019/20016/201
520/20220/20020/20018/20119/200
610/10110/10010/1008/10010/101

Performance

OverallSwipe RightSwipe LeftSwipe UpSwipe DownPushTOTAL
Recall91.7 %93.8 %94.4 %91.0 %88.1 %91.8 %
Precision95.0 %97.5 %94.4 %96.1 %93.3 %95.3 %
StandingSwipe RightSwipe LeftSwipe UpSwipe DownPushTOTAL
Recall87.5 %90.6 %91.3 %89.4 %86.3 %89.0 %
Precision94.0 %96.7 %89.0 %96.6 %90.2 %93.2 %
SittingSwipe RightSwipe LeftSwipe UpSwipe DownPushTOTAL
Recall94.9 %96.2 %96.8 %92.2 %89.4 %93.9 %
Precision95.8 %98.1 %98.6 %95.7 %95.6 %96.8 %
CloseSwipe RightSwipe LeftSwipe UpSwipe DownPushTOTAL
Recall89.7 %95.1 %96.8 %94.1 %89.8 %93.1 %
Precision96.5 %98.3 %92.3 %97.8 %92.3 %95.4 %
FarSwipe RightSwipe LeftSwipe UpSwipe DownPushTOTAL
Recall93.7 %92.6 %92.1 %87.9 %86.3 %90.5 %
Precision93.7 %96.7 %96.7 %94.4 %94.3 %95.1 %

Appendix III - Memory Footprint Calculation

All memory values in the deployment table are KiB, where 1 KiB = 1024 bytes. Memory usage was measured by building the radar deployment examples with the GCC_ARM 14.2.1 compiler and its Debug configuration flags, then analyzing the generated linker map, as the examples existed on September 4, 2026:

The calculation parses the Linker script and memory map portion of the GNU linker map. Each input section is attributed to the generated model object, TensorFlow Lite for Microcontrollers (libtensorflow-microlite.a), or ML Middleware (libraries_shared/ml-middleware). The size reported for each input section is then added to the Flash and/or RAM total for its component according to its enclosing output section.

For PSOC™ 6, Flash is the sum of .text + .rodata + .data, and RAM is the sum of .data + .bss + .noinit. For PSOC™ Edge, Flash is the sum of .app_code_main + .app_code_itcm + .app_code_socmem + .ram_vectors + .cy_socmem_data + .data, and RAM is the sum of .app_code_itcm + .app_code_socmem + .ram_vectors + .cy_socmem_data + .data + .bss + .noinit + .cy_sharedmem. Sections that are stored in Flash and executed or initialized in RAM are counted in both totals. Consequently, these results describe the linked deployment example rather than a platform-independent measurement of the model alone.

The PSOC™ Edge Flash figures are currently overestimated because of a known bug in the code examples. A fix is work in progress, so the PSOC™ Edge Flash values should be treated as provisional until the corrected examples are available.

Last updated on