Download molegro virtual docker

Transcript
13 Data Analyzer
page 195/321
created indicating the accuracy of the current model setup. Notice: For
Percentage split validation, the prediction is only made for the test set (training
set entries are set to 'NaN').
It is possible to create general regression models by first training a model
using the N-CV or LOO procedure in order to identify promising descriptors and
model training parameter settings. Therefore, the Regression Wizard must
be invoked more than once. To aid in the selection of descriptors and
parameter settings, the wizard remembers the previously used settings making
it easier to adjust the parameters. When a model of high generality has been
identified (using the correlation coefficient as a measure of generality), a
regression model can be created using the Create new model and
prediction option.
A way to check whether a regression model is over-fitted or not is to compare
the correlation coefficient of the trained model (R train) with the correlation
coefficient obtained from N-CV or LOO validation (R cv). If Rtrain is much higher
than Rcv, the model is probably over-fitted.
Feature Selection
The built-in feature selection algorithms can be used to identify relevant
descriptors (see Figure 136). Reducing the number of descriptors makes it
easier to interpret the model, and makes overfitting less likely.
molegro virtual docker – user manual