Download molegro virtual docker
Transcript
13 Data Analyzer page 195/321 created indicating the accuracy of the current model setup. Notice: For Percentage split validation, the prediction is only made for the test set (training set entries are set to 'NaN'). It is possible to create general regression models by first training a model using the N-CV or LOO procedure in order to identify promising descriptors and model training parameter settings. Therefore, the Regression Wizard must be invoked more than once. To aid in the selection of descriptors and parameter settings, the wizard remembers the previously used settings making it easier to adjust the parameters. When a model of high generality has been identified (using the correlation coefficient as a measure of generality), a regression model can be created using the Create new model and prediction option. A way to check whether a regression model is over-fitted or not is to compare the correlation coefficient of the trained model (R train) with the correlation coefficient obtained from N-CV or LOO validation (R cv). If Rtrain is much higher than Rcv, the model is probably over-fitted. Feature Selection The built-in feature selection algorithms can be used to identify relevant descriptors (see Figure 136). Reducing the number of descriptors makes it easier to interpret the model, and makes overfitting less likely. molegro virtual docker – user manual