Start Of Movement Thresholds Methods For Force Plate Testing

DECIDING THE START OF MOVEMENT ONSET USING FORCE PLATES

IS THERE A BEST METHOD FOR CALCULATING SOM?

Recently I was asked a question regarding the start of movement detection (SoM) for force plate analysis. This is a topic that some applied practitioners care about pretty highly, likely because they want to ensure the most valid and reliable data collection possible. So by correctly identifying the SoM this will allow for the most valid and reliable assessment to be performed. All this is well intentioned, but is it that important?

I think due to increased visibility of researchers and the need to suggest a best method in search of the “sport science holy grail” of best practices, this has now led to increased confusion surrounding “what is best?” for the SoM.

As there are a number of potential SoM detection method thresholds that exist and are being used in both research and practice. It would therefore seem pertinent to review what is actually recommended in the literature. There is quite a bit of noise out there – so how do we make sense of it all (and understand the signal – literally)? (hopefully this article can help).

START OF MOVEMENT (SoM) METHODS

Here are some of the key SoM thresholds that are typically employed in the literature:

The beauty of having these defined criteria thresholds means analysis can be performed in different ways. This also means that software can be used to identify the SoM algorithmically, based on the defined criterias of the F-T signal (outlined above with the exception of manual selection of course!). Meaning that time can be saved in interpreting the F-T signal and outcomes, which is why a number of commercial force plate softwares exist (as no one has the time to manually inspect traces anymore). 

However this also means we have a bunch of different methods to choose from, so which is right?

At least for now as we have a sense of what the different thresholds are and how they are calculated, this article will focus on a review of the current literature that discusses these methods, to determine whether or not we should lose sleep over our choice of SoM within our force plate assessments.

THE PRIMARY DISCUSSION... "WE NEED PRECISION"

In a world of measurement and evaluation in sports and the rise of applied sport science utilizing technology, it has become more important than ever to ensure the correct decisions are being made with data to influence the performance process effectively. This means collecting valid and reliable measures on athletes to ensure the most accurate representation outcome is being determined. 

Therefore we need technology (and assessments) that provide a high amount of precision and where possible the lowest amount of error to enhance this decision making process (afterall noisy data = noisy outcomes).

As the SoM is one of the first steps in determining calculations related to force plate movement assessments, it would seem common sense that we want to have the most precise estimate of the SoM to ensure the most accurate level of data is being collected. 

The thought process here is that if the SoM is calculated incorrectly then this will have a downstream effect on the variables derived from the F-T signal later in the movement (we explore this in more detail further down in the next sections). As such this has led a number of researchers to investigate SoM with the following questions…

  1. What is the most precise calculation for the SoM?
  2. Does the SoM calculation create “significant” differences in calculation of metrics? (i.e. does it have a downstream effect)

What is the most precise calculation to use for SoM?

From my understanding of the recent landscape related to the SoM, discussions have focussed primarily on the use of the 3x and 5x SD threshold methods (which uses the SD of the weighing phase capturing system weight). The reason for this choice is related to our knowledge of statistics and distributions, in which ~99.7% and  ~99.9% of the variation in the distribution of the weighing phase will be represented by 3 and 5 SD respectively. 

Using these set thresholds means that any deviation beyond this threshold is likely to be a real and meaningful change in weight as detected in the F-T Signal and related ground reaction force (GRF). Therefore this change beyond this threshold would represent a true change in movement and indicate the SoM. 

During assessments like countermovement jumps performed on force plates, it’s extremely rare that an individual can remain perfectly still prior to initiation of the jump (despite the best efforts to remain still). This is thought of as one such reason for using the SD threshold, as this uses the captured variance of the individual’s system weight as a means to assess the change due to the change in resultant GRF related to the movement being performed. 

At first glance this seems favorable compared to other fixed thresholds (e.g.10-50N) or percentage of body weight (e.g. 2.5-10%) which may be seen as insensitive, due to variation between movements and individuals (i.e. those who weigh more or less will likely contribute closer or further from these thresholds) [1]. As such some researchers have made a case for the 5SD to be the most valid criteria for performance assessments [2]. However, while the 5SD threshold may reduce some of these potential sensitivity “issues”, it doesn’t necessarily mean that this is the best representation of movement onset as this may be too conservative of an estimate especially in dynamic tasks where movement is increased. It should also be noted that the SD threshold originated from isometric assessments in which a fixed body position is used – meaning it is likely best suited when variability in weighing is lowered (as greater variability will mean a larger threshold). 

There are however other thresholds related to SoM (reverse scanning and the first derivative method aka yank) that have also been suggested to be more accurate for the detection of SoM. For example Sahrom et al [3] investigated yank as a means of SoM in relation to the CMJ and compared across other thresholds, showing that the first derivative method was the closest relating to motion capture (Vicon – which offers the gold standard in vertical jump and center of mass displacement capture) and of those other threshold methods the SoM was significantly delayed and had greater bias (comparisons shown in the table below). Therefore the authors suggested that this means the other methods miss vital information that can be subsequently captured by using the first derivative (yank) method. Pinto & Callaghan [1] show a similar finding, suggesting a 10Hz filtered first derivative shows the highest level of agreement to a manually selected method (that had high intra-rater reliability). While this “gold standard” criterion isn’t as robust as the one in the previously mentioned study, it still represents what we would say is our best guess of identifying the SoM. Combined both these studies suggest there are limitations to the 5SD and other methods relying on the weighing period. 

Given these outcomes it would seem important to critically assess the reliability and validity of these different thresholds (a quick overview of some studies are presented in the table below).

Per the table above what we see is our first main point of interest. In a lot of cases all the calculated metrics associated with any given threshold offer a good level of reliability (based on the statistics used). Therefore this suggests each SoM offers a reliable means to estimate data within its own realm. 

This would make a lot of sense, given that the SoM onset is only likely to influence those variables that are associated with a time component (i.e impulse, RFD, force at a given time point, time to take off and phase duration mainly eccentric) [2] to which this will be a few milliseconds in almost all cases. This is why in the table above I selected metrics that related as closely to time in the given studies (there are more measures in each study most of which show the same results in regards to reliability of the measures for the most part – except time to peak force in some instances which is a poor indicator regardless of SoM due to the variation in waveforms).

From these results we can likely conclude that each SoM threshold method offers a reliable means (i.e. it hits the target consistently enough), however this doesn’t mean that it is valid (does it measure what it’s supposed to measure – the true start of movement?).

From the current available research there are two studies [2] [3] that have assessed validity in relation to a criterion “gold standard”. Both suggest that the first derivative method to be the most valid in comparison to the other methods (as we outlined in the table above)

However what should be noted from these studies is that the mean bias range for all SoM threshold methods always includes 0 for each presented comparison and is generally pretty close to 0. Though the limits of agreement are suggested to be large in some cases, with narrower limits preferable, one could arguably make the case that these are still clinically acceptable ranges (as there is no cut off threshold of good and bad when it comes to LOA and this will be specific to the measurement being taken).  Meaning it’s generally up to us as practitioners to decide how much error we are comfortable with in relation to the gold standard (after all each method can still be collected reliably once implemented).

Another thing that may be important is understanding how different each method is. This means we would then need to turn to how relevant the differences are between each of the SoM calculations… That is to say how much the metrics are affected overall by each SoM method…

Does the SoM calculation create “significant” differences in calculation of metrics? (i.e. does it have a downstream effect).

As mentioned previously, the downstream effects of identifying the correct SoM threshold are likely more consequential to those variables that have a time component associated with them (e.g. contraction time, phase durations, impulse and rate of force development). For example Sahrom et al [3] showed that the first derivative method on average had an earlier start of onset time (75 +/- 88ms) compared to the fixed and SD methods (however note the large SD!). 

In such a case this would mean that contraction time (the SoM to toe off in a CMJ for example) would be longer as the movement onset occurred earlier on the F-T trace. The authors reported on average the other methods had a detection “error” being in the region of ~163ms compared to the criterion (this is around 20% of the total CMJ time – and likely meaningful if we want to get the closest SoM). That being said, the authors did not present any calculated metrics. So the effects on forward dynamics cannot be stated between these methods and the impact of this. Though its likely that things like RSImodified and potentially phase durations may be over / under estimated based on the different methods. Taking this study at face value, one could imply that the first derivative and the 2x SD method likely offer the best options of the SoM thresholds observed (as they have the lowest mean bias compared to motion capture).

This is also likely why we see that force plate derived jump height measures are generally lower estimates than the derived measures from motion capture. We are using a derivative to estimate the center of mass displacement and given that this can be influenced by the starting position, accurately detecting initial movement and the height the CoM started at, it is unsurprising we see differences. As relying on impulse (the area under the curve) iis impacted by an accurate weighing period. 

The table below details some comparisons between different SoM thresholds and variables that have been noted in the literature:

Now I quickly want to mention here that when we make comparisons this way we want to assess the magnitude of the difference in relation to the smallest meaningful effect (or the smallest effect size of interest). That is to say that just because a study shows a “significant difference” or an effect size, we have to ensure that this isn’t due to sampling variance, the studies power and potential artifacts (i.e. we should state the effect size of interest apriori). 

Without doing a complex meta analysis here, we are going to use some known reliability data for both CMJ and SJ to compare the changes based on what we know about measurement error for the given metric. The data set can be accessed here: Collings et al 2024 – Reliability and Validity

We use the SEM for the selected metrics and then interpret them as trivial if observed difference < SEM and “possibly real” > SEM 

Comparing the differences above we see that almost all differences reported in the selected variables fall below the expected SEM, with the exception of 4 comparisons (which are lower than the minimal detectable change MDC but higher than the SEM). We can interpret this as any differences between thresholds are likely negligible. Now I didn’t have time to do every variable and comparison (feel free to do this yourself at this time). 

What we should note here is that not all papers assess the right things (i.e. variables reported). That is to say it’s most likely that a change in the SoM will likely impact eccentric phase durations and potentially subsequent concentric phase durations [4] . However, it’s unlikely that these differences will yield any significance compared to the expected systematic and random errors we would see in any given assessment trial to trial using different SoM thresholds. Suggesting that the different SoM thresholds relying on system weight largely have very minimal (millisecond) impacts on the derived metrics from the F-T signal (if we used yank then we would likely see greater differences if the right filter is applied as this closer reflects the outcomes of motion capture).

Given this and the previously mentioned reliability data it’s likely that any given SoM threshold will still be useful within its own domain. Which would mean we are best served to stick to a method rather than comparing results across methods. 

OTHER THINGS WE SHOULD CONSIDER...

One of the most important things with any force plate analysis is capturing an accurate system weight [5]. This has been the focus of research in terms of what is an ideal weighing period to identify…

As such research has demonstrated that a 1 second weighing period should be performed with the individual as still as possible, and likely no more than 2 seconds as this demonstrates poor repeatability [5] and can create further errors in calculations [6]. This also means when assessing someone you cue them to minimize movement (such as hand gestures and head movement) to minimize any additional sway that will influence ground reaction forces (this is likely why hands on hips has been suggested to be a more reliable means than its counterpart arm swing). 

Another important aspect to consider here is the movement at initiation. For example some individuals may elevate their center of mass (COM) by performing a heel raise at the SoM. As such practitioners have been advised to always assess their F-T trace and discard any jumps that deviate >50N above the weigh-in threshold (as advised as a general rule of thumb) or recognize this change in technique can have an impact on the forward dynamic calculations.

This is also important when comparing different individuals, when performance profiling, as this would signify different types of jump signatures despite colloquially being considered CMJs. As such the manual inspection method though tenuous should likely always be performed in conjunction with assessments to ensure the highest level of reliable data is collected. So when you suspect something looks off from a metric standpoint – it likely relates to the initiation of movement following the weighing period.

SO WHAT THRESHOLD SHOULD YOU USE?

Ultimately the choice of SoM method threshold is entirely up to you the practitioner, who after reading this you should have an idea of the limitations of each method. We should also understand that each SoM has its own “within method reliability” so you will be okay in the long run. However some SoM thresholds may serve a better purpose when a less noisy weighing periods are apparent (e.g. a 2 – 5 SD threshold) and others may be better suited for this situation (fixed thresholds).

Simply, choose an SoM method and stick to it – as this will be the key over time, as long as your protocol is standardized. Given that the more important consideration is protocol standardization and ensuring the start of movement is initiated after being as still as possible, I would spend more time making sure athletes are well familiarized with the testing protocol and ensuring that movement execution is standardized each time you test (as there is already a fair amount of variability in the waveform trial to trial). 

Given that deviations in weight and the “quiet stance” prior to movement onset will impact some of the SoM thresholds more than others this is important to consider. But also this system weight is more likely to impact the calculated metrics irrespective of the SoM threshold you use. So much to say that the SoM method will have minimal impact on most of the important outcome measures, but an inappropriate weighing period can introduce greater systematic and random errors to the calculations and give you poor data (e.g. a 1N difference in system weight can impact jump height ~1cm or ~0.25″). 

Ultimately the differences between SoM methods are likely trivial and can most likely be accounted for by the typical variation (systematic and random error) we would see in the testing measure, with the weighing period, “quiet stance” and technical execution having a greater impact on the calculated metrics. Although one may make a case for the first derivative as being the closest to valid form to detect SoM (though this likely needs a filter to be applied as interpreting the raw signal may come with its own set of challenges).

So don’t lose sleep over your SoM. If you are going to get super technical the SoM likely should be one related to yank (the first derivative – potentially filtered at 10Hz) though we need more studies on a larger population to say this with any confidence. However theoretically this makes sense as this is a threshold that does not rely on the weighing period so should be the most consistent test to test and between individuals (at least in theory). 

REFERENCES
  1. Pinto, B. L., & Callaghan, J. P. (2023). Movement onset detection methods: a comparison using force plate recordings. Journal of Applied Biomechanics39(2), 118-123.
  2. Owen, N. J., Watkins, J., Kilduff, L. P., Bevan, H. R., & Bennett, M. A. (2014). Development of a criterion method to determine peak mechanical power output in a countermovement jump. The Journal of Strength & Conditioning Research28(6), 1552-1558.
  3. Sahrom, S. B., Wilkie, J. C., Nosaka, K., & Blazevich, A. J. (2020). The use of yank-time signal as an alternative to identify kinematic events and define phases in human countermovement jumping. Royal Society open science7(8), 192093.
  4. Eagles, A. N., Sayers, M. G. L., Bousson, M., & Lovell, D. I. (2015). Current methodologies and implications of phase identification of the vertical jump: A systematic review and meta-analysis. Sports Medicine45(9), 1311-1323.
  5. Pinto, B. L., & Callaghan, J. P. (2024). Effects of weighing phase duration on vertical force-time analyses and repeatability. Sports biomechanics23(12), 2862-2872.
  6. Street, G., McMillan, S., Board, W., Rasmussen, M., & Heneghan, J. M. (2001). Sources of error in determining countermovement jump height with the impulse method. Journal of Applied Biomechanics17(1), 43-54.
  7. Barefoot, M., Lamont, H., & Smith, J. C. (2022). Comparison of the reliability of four different movement thresholds when evaluating vertical jump performance. Sports10(12), 193.
  8. Smith, J. C., Lamont, H. S., & Barefoot, M. (2024). Comparison of Different Take-off Thresholds When Assessing Vertical Jump Performance. International Journal of Exercise Science17(1), 660.
  9. Pérez-Castilla, A., Rojas, F. J., & García-Ramos, A. (2019). Assessment of unloaded and loaded squat jump performance with a force platform: Which jump starting threshold provides more reliable outcomes?. Journal of Biomechanics92, 19-28.
  10. Donahue, P. T., Hill, C. M., Wilson, S. J., Williams, C. C., & Garner, J. C. (2021). Squat jump movement onset thresholds influence on kinetics and kinematics. International Journal of Kinesiology and Sports Science9(3), 1-7.
  11. Meylan, C. M., Nosaka, K., Green, J., & Cronin, J. B. (2011). The effect of three different start thresholds on the kinematics and kinetics of a countermovement jump. The Journal of Strength & Conditioning Research25(4), 1164-1167.
  12. Dos’ Santos, T., Jones, P. A., Comfort, P., & Thomas, C. (2017). Effect of different onset thresholds on isometric midthigh pull force-time variables. The Journal of Strength & Conditioning Research31(12), 3463-3473.

Leave a Reply

Shopping Cart

Discover more from The Science-Based Lifter

Subscribe now to keep reading and get access to the full archive.

Continue reading