A study tracking 643 adult runners through a full marathon training cycle for the New York City Marathon has produced one of the more sobering results in recent sports science: a well-built predictive model, trained on weekly surveys and Strava-sourced GPS logs, that performs impressively overall and considerably worse at the one job most runners actually want it to do — telling a healthy runner whether next week is the week they get hurt.
Researchers collected baseline surveys and sixteen weekly check-ins from each of the 643 participants, cross-referenced against their own training logs, yielding 9,002 runner-week observations in total. Just under half of the group, 48 percent, reported at least one injury serious enough to force a change to their training during the sixteen-week block, while the remaining three-quarters of individual weekly observations were injury-free. A generalised additive model, a form of machine learning suited to messy, non-linear relationships like the one between training load and soft-tissue injury, was trained on this data to predict which weeks would end in an injury requiring modified training.
Run across the whole dataset, the model performed well: an area under the ROC curve of 87 percent, a standard measure of how well a binary classifier separates two outcomes, with 50 percent being no better than chance. But that headline figure was doing a lot of work that the authors were careful to unpick. Restricting the same model to only the weeks where a runner had reported no pain and no prior injury — precisely the population a prevention tool would need to serve — collapsed its performance to 67 percent, with a precision score of just 8 percent. In plain terms, a model that looks accurate across a whole training block becomes close to useless at the exact moment a runner would want a warning: before anything hurts.
The predictors that did carry weight are, on reflection, less surprising than the headline result. A runner's pain and injury status the previous week was the single strongest signal, essentially confirming that injuries tend to announce themselves before they fully arrive, if anyone is paying attention. An acute:chronic workload ratio — the now-familiar comparison of a runner's most recent week of training against their rolling four-week average — clustered around 1.3 as a meaningful threshold, broadly consistent with injury-risk research published on much larger recreational cohorts using the same metric. Training volume, weekly peak distance, and a prior history of running injury all featured too, alongside a curious but intuitive finding: risk climbed in weeks seven through nine of training, the point where cumulative fatigue has built but the taper has not yet begun to offer relief.
None of this amounts to a workable early-warning system, and the study's authors say so plainly: predictive power was insufficient to offer personalised prevention guidance to runners who are currently healthy. For the growing industry of wearables and training apps promising to flag injury risk before it happens, the finding is a useful corrective. The data that best predicts an injury is the runner's own body starting to complain about it — which means the humbler advice, to take a niggle seriously in week seven of a marathon block rather than push through it, may still be doing more good than any algorithm built on top of it.
