From Continuous Predictors to Clinical Thresholds: Early Evidence on Performance Trade-offs of Guideline-Based Categorisation for Ischaemic Stroke Outcome Prediction
06:00 · August 7, 2026 · arXiv cs.AI RSS

Machine learning models achieve strong predictive accuracy for 90-day outcome prediction in acute ischaemic stroke, yet clinical adoption is limited by the misalignment of model explanations with clinicians' reasoning. Motivated by a clinician user study calling for clinical guideline-aligned cut-offs, we ask whether continuous predictors can be replaced by clinically informed categorical encodings without sacrificing performance. On a multi-centre European registry stratified into three treatment cohorts, we compare standard and fully categorised gradient-boosted models, the latter using stroke guideline-aligned, treatment-specific thresholds. The fully categorised models are statistically indistinguishable from their continuous counterparts in two of the treatment cohorts, with a significant drop in predictive accuracy in one cohort. Global feature importance rankings remain consistent, suggesting that discretising continuous predictors into guideline-based categories preserves the core hierarchy of prognostic factors across all treatment groups. Guideline-based categorisation is thus a viable design choice for stroke-outcome models.
Summary
Machine learning models for predicting 90-day functional outcomes after acute ischaemic stroke routinely reach high accuracy with gradient-boosted trees and SHAP explanations, yet adoption stalls when the resulting attributions diverge from the threshold-based reasoning clinicians apply in practice. A prior user study with twelve stroke specialists revealed that continuous variables such as blood pressure, glucose and onset-to-door times introduced clinically irrelevant granularity, prompting calls for encodings that mirror guideline ranges rather than raw values.
To test whether such re-encoding could be performed without performance loss, researchers trained paired gradient-boosted models on a multi-centre European stroke registry of 3,017 patients. The data were stratified into three treatment cohorts—no recanalisation, thrombolysis alone, and thrombectomy with or without thrombolysis—and continuous predictors were replaced in one model variant by categorical features whose cut-offs followed AHA/ASA and stroke-specific guidelines. Missing values were imputed within each train-test split to prevent leakage, and model performance was assessed with standard discrimination metrics.
In two of the three cohorts the fully categorised models produced results statistically indistinguishable from their continuous counterparts, while global feature-importance rankings derived from SHAP values remained consistent across all groups. A measurable drop occurred in the remaining cohort, indicating that the effect of categorisation is pathway-dependent. The authors conclude that guideline-aligned discretisation offers a practical route to explanations that clinicians already recognise, thereby reducing the cognitive friction that currently limits deployment of otherwise accurate stroke-outcome models.
Why it matters
The article is highly relevant for researchers focusing on Explainable AI (XAI) and clinical decision support systems. It provides empirical evidence on how to bridge the gap between technical model explanations and clinical reasoning, aligning well with the Dutch and EU focus on transparent, trustworthy AI in healthcare.






