Research published in PM&R has tested one of the more persistent assumptions in recreational running: that combining survey data with wearable training logs should be enough to predict who is about to get injured. The study followed 643 adult runners, 53 per cent of them female with a mean age of 43, through their preparation for the New York City Marathon. Participants completed a baseline survey and 16 weekly interval surveys, and linked their Strava accounts so that GPS watch and smartphone training logs could be aggregated alongside the self-reported data.
The headline finding is a distinction that matters more than it might appear. The models performed reasonably well at classification, meaning they could identify which runners were currently reporting an injury from the combination of their training history and their survey responses. They performed considerably less well at prediction, meaning forecasting the injury status of the following week. That gap is the central problem in this field. A model that tells a runner they are injured is describing something the runner already knows. A model that tells them next Tuesday's long run is the one to shorten would be genuinely useful, and remains out of reach.
The result sits alongside the Garmin-RunSafe findings from earlier this year, which followed more than 5,200 runners and concluded that overuse injuries tend to arrive suddenly during a single session rather than accumulating gradually. If that is correct, it explains a good deal about why week-ahead prediction is so hard. A model trained on weekly aggregates is looking for a slow trend in data generated by an event that may be decided in a single run, and the strongest single predictor identified in that study was a sharp distance spike relative to the previous 30 days, which is a session-level variable rather than a weekly one.
There is also a measurement problem running underneath both studies. Injury status here is self-reported, and self-reported means different things to different runners. A participant who describes themselves as injured may have stopped running entirely or may have run 60 miles that week with a sore Achilles. Training logs from Strava capture distance, pace and elevation but not surface, not sleep, not the two hours spent standing at work, and not strength training done off the platform. The models are working with a partial picture and are, to their credit, honest about it.
None of this makes the work useless. Establishing that classification is achievable and prediction is not, on a cohort of this size with linked objective training data, narrows the problem considerably for the next study. It also has an immediate practical implication for runners: the injury-risk figures displayed by consumer watches are built on similar logic and should be read as descriptions of what has already happened rather than warnings about what is coming. The load-management advice that survives all of this is unglamorous and unchanged. Avoid large single-session spikes, treat a new niggle as information rather than noise, and do not expect an algorithm to make the decision for you.
