Why These Five Tests
The evidence behind each instrument in the battery, and where it stops
This is written for practitioners (physiotherapists, clinicians, coaches) rather than clients. It sets out why each test is in the battery, what the evidence supports, and, for each one, the limitation I think matters most.
I'd rather state those limitations myself than have someone find them. If you think I've overstated anything here, I'd genuinely like to know which.
The architecture first
The most useful finding in this literature isn't about any individual test. It's that multifactorial models outperform single-metric assessment: combining movement quality, strength data, workload, injury history and psychological factors predicts and describes better than any one measure alone. Every systematic review cited below that addresses the question reaches the same conclusion.
So the battery isn't five tests chosen because each is strong. It's five tests chosen because together they cover domains that don't substitute for one another, repeated at weeks 8 and 16 so the result is a record rather than a snapshot.
None of the individual tests is unusual. The combination, applied to recreational and clinical populations rather than professional athletes, is.
1. Movement screen
Why it's included. Structured observation of loading tolerance before progressive exercise is prescribed. Seven patterns, 0 to 3 scoring.
What the evidence supports. Movement screening of this type demonstrates good reliability when performed consistently: pooled inter-rater ICC 0.843 and intra-rater 0.869 in Cuchna et al.'s meta-analysis of the FMS composite. Bonazza et al. (2017) reached similar reliability conclusions and additionally found an association between composite scores of 14 or below and subsequent injury.
The limitation. Dorrel et al. (2015) found sensitivity too low for the composite to function as a standalone injury prediction tool, and I take that position over Bonazza's. The disagreement in the literature isn't resolved, and I'd rather name it than cite only the half that suits me.
Two further points of honesty. My screen is FMS-informed but it is not the FMS: a different set of patterns, with different scoring conventions. The reliability figures above describe the FMS composite and I don't claim they transfer to my screen. And movement screening has a structural ceiling: it describes movement under unloaded, standardised conditions, which is not the condition under which people typically get hurt.
I use it to decide what not to load and where to look next. Not to generate a risk score.
2. Isometric strength and limb symmetry
Why it's included. Objective force production at standardised joint angles, with a symmetry index. The strongest evidence base in the battery.
What the evidence supports. Liao et al.'s COSMIN-methodology review of 123 studies (n ≈ 5,000) found moderate-to-high certainty evidence of sufficient reliability (ICC ≥ 0.70 across most muscle groups, positions and fixation methods), with high-certainty evidence specifically supporting external fixation for knee extensors and flexors, and sufficient criterion validity against isokinetic reference for seated knee testing (r = 0.82 to 0.90). Smaller validation studies against isokinetic dynamometry reach the same conclusion.
The limitations, and there are three.
Liao et al. conclude explicitly that portable dynamometers should not yet be treated as direct substitutes for isokinetic dynamometry. They're appropriate for longitudinal monitoring and clinical screening, which is how I use them, not for one-off clearance decisions.
The LSI overestimates function when the uninvolved limb has also weakened through compensatory loading or prior injury (Wellsandt et al., 2017). Bilateral injury history mitigates this. It doesn't eliminate it.
The ≥90% threshold derives overwhelmingly from ACL return-to-sport research. It does not transfer cleanly to the lumbar spine, the shoulder, or to recreational adults not returning to competitive sport. I treat it as a reference point and say so.
3. Force-velocity profile
Why it's included. It answers a question no other test here answers: not where someone sits on absolute strength or symmetry, but which neuromuscular quality is under-developed, and therefore what type of stimulus will move them.
What the evidence supports. The force-velocity relationship is among the most established concepts in exercise science. Field-based profiling using velocity-based devices shows good intra-session reliability and strong agreement with laboratory reference systems. Meta-analytic evidence supports individualised profiling to guide training design over generic programming.
The limitation. Extrapolating theoretical maximum force and velocity from field data introduces non-trivial estimation error relative to force plates. I use the profile directionally (force-dominant, velocity-dominant, balanced) and not for mathematical extrapolation or for prescribing loads from calculated values. Between-session reliability is also less well established than within-session, so profiles are retested under identical conditions or not compared.
4. Heart rate recovery (HRR-60)
Why it's included. An objective, standardised marker of autonomic recovery, repeatable weekly, requiring no wearable.
What the evidence supports. HRR-60 is well established as a non-invasive marker of parasympathetic reactivation, with a substantial clinical literature behind it including large cohort data.
The limitation, and it's the most important one here. The great majority of that literature is in cardiovascular disease populations. I apply the marker in a fitness monitoring context for training load management, not cardiac assessment of any kind. The physiological mechanism is the same across populations, but the prognostic thresholds published in cardiology do not transfer to healthy active adults and are not used. HRR-60 is interpreted only as within-person change against a standardised protocol. Any cardiac concern is a GP referral, without exception.
5. Reactive strength index
Why it's included. Among the most reliably measured performance metrics available without laboratory equipment, and it captures a quality that declines earlier than maximal strength and is almost never tracked.
What the evidence supports. Drop jump RSI shows excellent within-day reliability (ICC 0.95 to 0.99, CV 2 to 3% in force plate studies), and acceptable between-day reliability (ICC ~0.87 to 0.88, CV 4 to 6%). The most recent reliability data using IMU-based measurement in elite athletes reports CV ≤ 10%, ICC ≥ 0.8.
The limitation. All of that reliability data comes from elite and professional athletic populations, as do the published benchmarks. They are not applicable to recreational adults aged 35 to 65 and I don't grade anyone against them. Between-day reliability being lower than within-day also sets a floor on how small a change can honestly be called real, which matters directly when interpreting an 8-week retest.
6. Validated questionnaires
PSQI for sleep, DASS-21 for psychological load, TSK-11 for fear of movement, PRS for session readiness. All psychometrically validated, all used as screening instruments rather than diagnostic ones.
Two notes on correct use, because both are commonly got wrong. DASS-21 raw subscale scores are doubled before comparison against published severity norms: the cut-offs apply to the doubled score. And the TSK-11 has a range of 11 to 44; the frequently quoted threshold of 37 belongs to the original 17-item scale (range 17 to 68) and is routinely misapplied to the short form.
Elevated depression or anxiety scores go to a GP. They don't go into a training plan.
What this battery does not do
Stated plainly, so no referring clinician has to infer it from what's absent.
It does not diagnose. It does not predict injury: no test here, alone or combined, has demonstrated adequate sensitivity for that in this population. It does not clear anyone for return to sport; that decision belongs to the treating clinician, and I supply data into it rather than making it. It does not replace isokinetic testing. It does not assess cardiac function. And it does not grade anyone against athlete norms.
Why I've written it this way
Assessment tools get oversold routinely, and practitioners are right to be sceptical of a five-test battery presented with confidence. The limitations above aren't caveats added at the end. They're the reason I can tell you what the tests are good for with any credibility.
If a patient of yours has finished treatment and you'd like to know what they can currently produce before they return to load, that's exactly the gap this is built for. I send the full report back regardless of whether anything else comes of it.
And if you think I've overstated something above, tell me which. I'd rather adjust it than defend it.