A study published in BMC Sports Science, Medicine and Rehabilitation has applied machine-learning models to one of the larger running-injury datasets assembled anywhere: 1,744 runners drawn from the University of Calgary Running Injury Clinic database, of whom 1,088 were injured and 656 were not. The stated aim was to compare models capable of distinguishing the two groups using multidimensional data, and to work out which variables the models were actually leaning on.
The candidate predictors were unusually broad. Alongside the demographic and anthropometric measures that appear in most such work, the researchers included training-related variables, running speed, and three-dimensional biomechanical data captured during treadmill running. That last category is what separates this dataset from the questionnaire-based studies that dominate the field, and it is expensive enough to collect that few groups outside a dedicated clinic could attempt it at this scale.
Rather than reporting accuracy and stopping there, the authors used Shapley additive explanations to quantify how much each feature contributed to the models' output. This matters more than it sounds. A classifier that performs well while remaining opaque tells a clinician nothing actionable; one that can say which measurements carried the decision at least points towards the variables worth arguing about. It is a modest form of interpretability, but it is the direction the field has needed to move in for some years.
The design imposes a hard limit on what any of it means. This is a cross-sectional study: the runners were injured at the moment they were measured. Any biomechanical difference the models detect may be a consequence of the injury rather than a cause of it, and there is no way within this dataset to tell the two apart. An antalgic gait pattern is exactly the kind of signal a classifier would seize on, and exactly the kind that would be useless as a screening tool.
That limitation is thrown into relief by the prospective work running alongside it. A separate study following runners through New York City Marathon training — weekly surveys, GPS watch data, smartphone logs — found machine learning had low discriminatory power when asked to predict injury week by week going forward. Retrospective classification is tractable; forward prediction, so far, is not. Runners and coaches should treat both results as a map of where the research is, not as the basis for a screening protocol that does not yet exist.
