The Journal Training

How Performance Testing Changes Your Training

What changes once there are numbers on the table

By Giacomo Di Mauro, MSc · 5 min read · October 2026

Here are two runners.

Both 41. Both running four times a week, both have been for years. Both have the same complaint: the calf goes every few months. It settles with rest, they come back, and some months later it goes again. Both have been told to strengthen their calves. Both have.

Same sport, same age, same training history, same symptom. Indistinguishable from the outside.

They need opposite interventions, and giving either one the other's would waste twelve weeks.

Where they differ

The first produces good force through the calf but poorly off a short ground contact. He's strong, slow to express it. His tissue can generate what's needed; it can't absorb and return load quickly, so every ground contact costs more than it should and he accumulates more total work per kilometre than his training volume suggests.

The second is the reverse. Reasonable reactive quality, force production well down. His tissue is being asked for more than it can currently make, several thousand times per run.

One is a velocity problem. One is a capacity problem. They look identical in a training log and feel identical to the person inside them.

What each needs

The first one needs the work that teaches a system he already has to express itself faster. Plyometric progressions. Contrast work. Short-contact drills with strict quality criteria rather than accumulated reps. The load is low; the intent is everything.

The second one needs the boring version. Heavy, slow calf loading. Progressive and patient, measured in months. Nothing exciting and nothing that looks like performance training, because the base has to exist before anything can be built on it.

Now swap them.

The first spends twelve weeks on heavy loading. He gets stronger, measurably. His actual limiter doesn't move, because it was never force. Here's the part that makes this difficult to catch: his numbers go up. The intervention looks like it's working right up until the calf goes again.

The second spends twelve weeks on plyometrics. You're now asking a system to produce force rapidly that can't yet produce it slowly. That one isn't just ineffective.

The honest version of what testing does

It doesn't tell you what to prescribe. Judgement still applies, and some of these measures are directional rather than precise.

What it does is narrow four plausible explanations to one. That's the entire value. Without it, exercise selection for these qualities is an educated guess, and educated guesses are wrong often enough to cost someone a training year.

Choosing the right exercise isn't expertise. It's information. A coach who has measured isn't smarter than one who hasn't. He just knows something the other one is guessing at.

How this looks across the five domains

Each test changes a different kind of decision.

Movement quality decides what you don't load yet. A hip hinge scoring 1 out of 3 means the pattern isn't organised. Load it and the adaptation goes somewhere other than intended. So the first weeks are spent establishing the pattern (unloaded, high frequency, short sessions), and the loading starts once there's something to load. This is the least glamorous output of an assessment and the one that most often changes the first month.

Asymmetry decides volume distribution, not exercise choice. A meaningful side-to-side difference rarely means a different exercise. It means the weaker side gets more of it: often double the volume for a block, sometimes unilateral work substituted where bilateral work would let the stronger side cover. And it means retesting, because the only thing that matters is whether the gap is closing. The exercises are usually unremarkable; the dosing isn't.

The force-velocity profile decides training emphasis. Force-dominant means the next block is weighted toward velocity: jumps, contrast methods, lighter loads moved with intent. Velocity-dominant means the opposite: compound loading, progressive, unexciting. Balanced means the limiter is somewhere else and this test has told you where not to look, which is still useful information.

Recovery capacity decides dose. This one overrides the others. If heart rate recovery is suppressed against your own baseline, the answer is almost never more training. It usually means reducing volume for a period and rebuilding as the number returns. This is the measure that stops a good intervention being applied at the wrong time, and it's the reason two people with identical test results can need different week-one plans.

Reactive strength decides whether the plyometric work happens at all. Low RSI alongside low force production is not an argument for jumping. It's an argument for building capacity first and introducing reactive work later, once there's something to react with. Low RSI with good force production is the opposite: that's the person for whom plyometrics are the whole answer.

What a block actually looks like

Three phases, in an order the numbers set.

Phase one. Restore patterns that scored poorly, and create recovery space. If HRR is depressed, training volume comes down here rather than up. Nothing heavy, nothing fast. This phase exists because you can't load a pattern that isn't organised, and you can't build on a system that isn't recovering.

Phase two. Load the corrected pattern progressively. The weaker side carries more volume. This is where most of the strength change happens and where most of the time goes.

Phase three. Build the qualities the profile identified as missing: velocity, reactive strength, whichever the test named. This phase comes last because it requires the base established in phase two.

Then retest. Same protocol, same conditions, same equipment.

What the retest is actually for

This is the part most people underestimate.

Eight weeks in, you want to know whether the training worked. Without a baseline, you're asking whether it felt like it worked, and feeling is a bad instrument, heavily shaped by sleep, mood, and what you expected to feel. People reliably report improvement in interventions that produced none, and reliably report nothing from interventions that produced a lot.

With a baseline, the question gets answered. Hip abduction asymmetry 71% to 84% is an answer. HRR from 11 to 19 bpm is an answer. And when the number hasn't moved, that's an answer too, and a more valuable one, because it means something in the plan was wrong and you now know within eight weeks rather than within a year.

That's the actual argument for measuring. Not the first set of numbers. The second set.

The thing this doesn't solve

Testing narrows the options. It doesn't remove judgement, and it isn't a formula where numbers go in and an intervention comes out.

Two people with identical profiles can still need different work, because one plays tennis and the other runs, one has three hours a week and the other has six, one has a twenty-year training history and the other started in March. Those things matter and no dynamometer captures them.

What testing does is stop the guessing happening at the one point where guessing is most expensive: which quality is actually limiting this person. Get that wrong and everything downstream is well-executed work aimed at the wrong target.

That's not a small thing to get right. It's most of the job.