<p><strong>Mixed Effects Modelling of Brain Morphology for Alzheimer's Disease Progression</strong></p><p>This workflow shows a code-free pipeline for modeling longitudinal structural MRI data and turning each subject's change over time into features for machine learning. It takes regional brain measurements across multiple visits, normalises them, selects the most informative regions, fits linear mixed effects models to describe each subject's trajectory, and passes the resulting trajectory features to downstream classifiers or regressors for trajectory prediction.</p><p><strong>Input data</strong></p><p>The workflow ships with a sample dataset of features extracted with FreeSurfer 7.0 (regional volumes and related morphometric measures) from a longitudinal cohort, where each subject has scans from two or more visits. The sample is included only so the workflow runs out of the box. Users can replace it with their own FreeSurfer output by pointing the reader node to their file. The input should be in long format (one row per subject per visit) and contain a subject identifier, a time variable (for example age at scan, months since baseline, or visit number), the intracranial volume (ICV), the regional features, and the target label used for prediction.</p><p><strong>ICV normalisation (Developed in Python)</strong></p><p>Head size varies considerably between individuals, and raw regional volumes partly reflect that difference rather than disease-related change. The ICV normalisation nodes adjust each regional measure for intracranial volume so that subjects with different head sizes can be compared fairly and the downstream models focus on biologically meaningful variation.</p><p><strong>Feature selection with Mutual Information (Developed in Python)</strong></p><p>FreeSurfer produces a large number of regional measures, many of them redundant or only weakly related to the outcome. The Mutual Information (MI) filter ranks features by how much information they share with the target variable and keeps the most informative ones. This reduces dimensionality, lowers the number of LME models that need to be fitted, and helps keep the final model interpretable.</p><p><strong>Linear mixed effects trajectory modeling (Developed in Python)</strong></p><p>The core of the workflow is the pair of LME nodes, which follow the familiar Learner/Apply pattern used elsewhere in KNIME.</p><p>The LME Modelling (Learner) node is trained on the training partition. For each selected brain region it fits a linear mixed effects model with the chosen time variable as the fixed effect and subject-specific random intercepts and slopes. This captures both the population-level trend and how each individual deviates from it. The node outputs a model port object and a table of per-subject trajectory features.</p><p>The LME Modelling (Apply) node takes the learned model and applies it to unseen subjects in the test partition. Random intercepts and slopes for new subjects are estimated, so test subjects get the same feature set as training subjects without refitting the model. The node warns when a test subject's time values fall outside the range seen during training.</p><p>Both nodes return a wide table with one row per subject. For each region, the following trajectory features are produced, prefixed with the region name: the random intercept (the subject's individual baseline level relative to the population), the random slope (the subject's individual rate of change), the deviation at baseline, the deviation at the last visit, and the maximum change observed over the follow-up period. Subject-level summaries of the time variable (mean and maximum) are added once per subject.</p><p>Data are split into training and test sets at the subject level before LME fitting, so that no subject contributes visits to both partitions and information does not leak from test to training.</p><p><strong>Trajectory prediction with machine learning</strong></p><p>The trajectory features from the LME nodes summarise how each brain region is changing for each subject, not just its value at a single time point. These features are passed to standard KNIME machine learning nodes (for example logistic regression, random forest or gradient boosting) to predict the subject's disease trajectory or progression group, followed by scorer nodes to evaluate performance on the held-out test subjects.</p><p><strong>Requirements</strong></p><p>The LME nodes are part of the KnimeVis Python extension. Please install the extension before running the workflow.</p><p><strong>Citation</strong></p><p>Please refer to the following repositories for the missing nodes.</p><p>Reference Repository:<br>https://github.com/AliBhatti21/Data-Science-with-KNIME.git</p><p>Reference Papers:</p><p>https://ieeexplore.ieee.org/document/11607592</p>
Mixed Effects Modelling of Brain Morphology for Alzheimer's Disease Progression
This workflow shows a code-free pipeline for modeling longitudinal structural MRI data and turning each subject's change over time into features for machine learning. It takes regional brain measurements across multiple visits, normalises them, selects the most informative regions, fits linear mixed effects models to describe each subject's trajectory, and passes the resulting trajectory features to downstream classifiers or regressors for trajectory prediction.
Input data
The workflow ships with a sample dataset of features extracted with FreeSurfer 7.0 (regional volumes and related morphometric measures) from a longitudinal cohort, where each subject has scans from two or more visits. The sample is included only so the workflow runs out of the box. Users can replace it with their own FreeSurfer output by pointing the reader node to their file. The input should be in long format (one row per subject per visit) and contain a subject identifier, a time variable (for example age at scan, months since baseline, or visit number), the intracranial volume (ICV), the regional features, and the target label used for prediction.
ICV normalisation (Developed in Python)
Head size varies considerably between individuals, and raw regional volumes partly reflect that difference rather than disease-related change. The ICV normalisation nodes adjust each regional measure for intracranial volume so that subjects with different head sizes can be compared fairly and the downstream models focus on biologically meaningful variation.
Feature selection with Mutual Information (Developed in Python)
FreeSurfer produces a large number of regional measures, many of them redundant or only weakly related to the outcome. The Mutual Information (MI) filter ranks features by how much information they share with the target variable and keeps the most informative ones. This reduces dimensionality, lowers the number of LME models that need to be fitted, and helps keep the final model interpretable.
Linear mixed effects trajectory modeling (Developed in Python)
The core of the workflow is the pair of LME nodes, which follow the familiar Learner/Apply pattern used elsewhere in KNIME.
The LME Modelling (Learner) node is trained on the training partition. For each selected brain region it fits a linear mixed effects model with the chosen time variable as the fixed effect and subject-specific random intercepts and slopes. This captures both the population-level trend and how each individual deviates from it. The node outputs a model port object and a table of per-subject trajectory features.
The LME Modelling (Apply) node takes the learned model and applies it to unseen subjects in the test partition. Random intercepts and slopes for new subjects are estimated, so test subjects get the same feature set as training subjects without refitting the model. The node warns when a test subject's time values fall outside the range seen during training.
Both nodes return a wide table with one row per subject. For each region, the following trajectory features are produced, prefixed with the region name: the random intercept (the subject's individual baseline level relative to the population), the random slope (the subject's individual rate of change), the deviation at baseline, the deviation at the last visit, and the maximum change observed over the follow-up period. Subject-level summaries of the time variable (mean and maximum) are added once per subject.
Data are split into training and test sets at the subject level before LME fitting, so that no subject contributes visits to both partitions and information does not leak from test to training.
Trajectory prediction with machine learning
The trajectory features from the LME nodes summarise how each brain region is changing for each subject, not just its value at a single time point. These features are passed to standard KNIME machine learning nodes (for example logistic regression, random forest or gradient boosting) to predict the subject's disease trajectory or progression group, followed by scorer nodes to evaluate performance on the held-out test subjects.
Requirements
The LME nodes are part of the KnimeVis Python extension. Please install the extension before running the workflow.
Citation
Please refer to the following repositories for the missing nodes.