Exercise physiology has a sample size problem. The classic study recruits perhaps twelve to twenty trained runners, controls what it can, and reports an effect that may or may not survive contact with the messiness of actual racing. A review published this month in the Journal of Applied Physiology, by Daniel Muniz-Pumares with Ed Maunder, Barry Smyth and Ben Hunter, makes the case for a different approach: treat the mass-participation marathon as a natural experiment, and read the resulting datasets of hundreds of thousands of runners for patterns no laboratory could afford to generate.
The logic is that marathons already vary the things a researcher would want to manipulate. Different races are run in different temperatures, at different altitudes and on different profiles. Runners arrive with widely different training histories, adopt different pacing strategies, and make different fuelling decisions. Nobody randomised any of it, but the variation exists, it is documented in GPS files and split data, and the numbers involved are large enough that the noise starts to cancel out. That is the definition of a natural experiment, and it is how epidemiology has worked for a century.
The findings the authors draw together are more interesting than the method alone would suggest. Marathons, it turns out, are run close to but below critical speed, the intensity threshold above which fatigue accumulates in a fundamentally different way. Faster runners complete the distance at a higher fraction of critical speed than slower ones. That is not a trivial observation: it means the marathon is not a fixed physiological challenge scaled by ability, but a different sort of event depending on where in the field you are running, and it partly explains why pacing advice that works for a 2:30 runner translates poorly to a 4:30 one.
The review also gives durability a central role. Durability describes how well a runner preserves physiological function deep into prolonged exercise, and it is the trait that separates athletes who look identical in a twenty-minute laboratory test but diverge sharply after 30km. A durable runner loses less of their economy, their mechanics and their capacity as fatigue accumulates. Mass race data is unusually good at exposing this, because the second-half slowdown of a large field is effectively a durability measurement taken on tens of thousands of people at once, stratified by training volume, by race-day temperature and by starting pace.
The obvious caution is that natural experiments do not establish causation. Runners who train more are different from runners who train less in ways that go well beyond mileage, and a correlation between a training characteristic and a finishing time carries all of that baggage with it. The authors are explicit that the approach generates hypotheses rather than settling them, and that the useful workflow runs in both directions: find the pattern at scale, then take it into the laboratory and test it properly. For anyone reading training studies as a runner rather than a scientist, that is the practical takeaway. A finding drawn from half a million race files is a good reason to pay attention, and not yet a reason to change what you do on Tuesday.
