Page 9 - Read Online
P. 9

Hsu et al. Vessel Plus 2021;5:2  I  http://dx.doi.org/10.20517/2574-1209.2020.45                                                     Page 3 of 12

               hemorrhage, with follow-ups up to 1 year. These patients went through clinical examinations, including
               computed tomography (CT) and/or magnetic resonance imaging (MRI) for the indexed event, following
               the international clinical guidelines for stroke. The demographic data, stroke type, National Institutes of
               Health Stroke Scale (NIHSS) scores, Barthel Index, blood pressure upon admission, medical history, pre-
               existing comorbidities, and treatment data as well as modified Rankin scale (mRS), medications and some
               discharge and follow-up data were recorded. Using TSR data as a human study protocol was approved
               by the institutional review board of all participating hospitals. The details of diagnosis, inclusion criteria,
                                                                                    [10]
               and collection of clinical variables of this registry have been presented elsewhere . The full list of Taiwan
               Stroke Registry participating investigators is listed in Supplemental Appendix I.


               Data preprocessing
               The TSR included the following four categories of datasets derived at different time points from admission
               to discharge and follow-ups: (1) demographic data; (2) measurement/diagnosis; (3) inpatient treatments
               and medications; and (4) discharge information plus follow-ups for up to one year. To ensure data quality,
               we performed data cleaning, validation, and resampling to remove missing data, outliers, and miss-coded
                                                                                 [11]
               data, as well as inconsistent clinical measurements before building the models .
               The primary goal of this study was to develop a multiparametric tool to estimate the probability of
               predicting the best functional improvement after stroke. The mRS is a clinician-reported and -quantified
               measure of disability and has been widely used to evaluate stroke outcomes [12-14] . Several studies attested to
               the validity and reliability of the mRS at different time points [15-17] , and we followed the model of functional
               mRS outcomes measured at 90 days divided into good outcome (mRS 0-2) and poor outcome [18-20]  (mRS
               3-6) to determine what clinical variables and treatments showed significant predictive value for future
               disability status in studied stroke patients, and how accurate we can predict disability with this set of stroke
               big data-selected information.

               Statistical models
               We used the independent t-test and chi-square test to compare the clinical variables between patient
               groups and utilized univariate logistic regression to calculate the univariate odds ratio of variables. Also,
               multivariable logistic regression (LR) was employed for feature selection and 90-day mRS (mRS_3m)
               outcome prediction. We built two supervised learning models to compare the prediction performances
               of functional assessments and clinical data in the registry. In the first model, all the clinical data in TSR at
               admission and discharge were included in LR. In the second model, the variables selected 100/100 times
               were then used as the input features to predict mRS_3m. Furthermore, we evaluated the two models
               in four different subgroups of patients, including male, female, ischemia, and hemorrhage for a better
               understanding of how these subgroups may affect the performance and prediction models in different
               populations of stroke patients. The flowchart of model construction with different subgroup datasets is
               shown in Supplemental Figure S1.

               There were two clinical time points selected for our prediction models, i.e., admission and discharge.
               In models evaluated at admission, a total of 140 variables of information registered during admission
               were used. By adding to the data registered during admission, such as treatments, medications, and
               complications, a total of 262 variables were used in models at discharge. The 10-fold cross-validation
               method with 70% of the data for training and 30% for testing was used to select variables using stepwise
                                               [21]
               Akaike information criterion (AIC) . Variables were then included in the final models based on the
               criteria that they were selected ten times in each of the ten rounds of the bootstrap. The counts of selection
               and the coefficient of each selected variable were recorded and evaluated. Accuracy and area under the
               curve (AUC) were assessed by comparing model predictions to the actual mRS of patients in holdout test
               sets [Supplemental Figure S1]. The performances of each model were then evaluated and compared by
               statistical analysis using SPSS Statistics version 22 and RStudio version 1.2.1335 software.
   4   5   6   7   8   9   10   11   12   13   14