fitcnb
statistics: Mdl = fitcnb (X, Y)
statistics: Mdl = fitcnb (Tbl, ResponseVarName)
statistics: Mdl = fitcnb (Tbl, formula)
statistics: Mdl = fitcnb (Tbl, Y)
statistics: Mdl = fitcnb (…, name, value)
Fit a naive Bayes classification model.
Mdl = fitcnb (X, Y) returns a naive Bayes
classification model, Mdl, with X being the predictor data and
Y the class labels of the observations in X.
A naive Bayes model fits one univariate density to each predictor within each class, and treats the predictors as conditionally independent given the class. An observation’s likelihood under a class is therefore the product of its per-predictor densities, and its posterior follows by Bayes’ rule from the class prior.
Mdl = fitcnb (…, name, value) returns a naive
Bayes model with additional options specified by Name-Value pair
arguments listed below.
| Name | Value |
|---|---|
'PredictorNames' | A cell array of character vectors specifying the names of the predictors. The length of this array must match the number of columns in X. |
'ResponseName' | A character vector specifying the name of the response variable. |
'ClassNames' | Names of the classes in the class labels,
Y, used for fitting the model. ClassNames are of the same
type as the class labels in Y. Naming a subset of the classes keeps
only the observations belonging to them. The model keeps the classes in
this order; by default they are sorted. |
'Prior' | A numeric vector specifying the prior probability
of each class, in the order of ClassNames, or the character vector
'empirical' (default) to take the class frequencies, or
'uniform' to give every class the same probability. |
'Cost' | A square numeric matrix of misclassification
costs, where Cost(i,j) is the cost of classifying an observation of
class i into class j. The default is one off the diagonal and
zero on it. |
'ScoreTransform' | A character vector naming a transform
applied to the posterior returned by predict, or a function handle
taking and returning a matrix of the same size. The default is
'none'. |
'DistributionNames' | A character vector naming the
distribution fitted to every predictor, or a cell array of character vectors
naming one per predictor. Supported are 'normal' (default),
'kernel', 'mvmn' for a categorical predictor, and
'mn' for token counts. 'mn' describes the whole predictor
vector at once and so cannot be named for only some predictors. |
'Kernel' | The smoothing kernel of the predictors fitted
with a kernel density, one of 'normal' (default), 'box',
'epanechnikov' or 'triangle', given once for every predictor
or once per predictor. |
'Support' | The support of the kernel densities, either
'unbounded' (default), 'positive', or a two element numeric
vector giving finite bounds. |
'Width' | The bandwidth of the kernel densities, given as a scalar, as one value per predictor, as one per class, or as a matrix of one per class and predictor. By default each density chooses its own. |
'Weights' | A nonnegative single or double vector of
observation weights, one per row of X. An empirical prior sums them
per class; a normal density takes weighted means and standard deviations, a
kernel density weighs its observations but chooses its bandwidth from them
alone, and the multinomials count each observation by its weight. A row of
zero or missing weight is left out. The model’s W keeps the class of
the weights, while every computation runs in double, so Prior is
double where MATLAB returns single. The default is uniform. |
A predictor that takes one value throughout a class has no normal density
to fit, and that combination of class and predictor is refused rather than
answered. Only the combination is refused, not the model: giving that
predictor a 'kernel' or a 'mvmn' distribution fits the same
data, and leaves the other predictors normal.
See also: ClassificationNaiveBayes
Source Code: fitcnb
Fit a naive Bayes classifier to Fisher's iris data and see how often it classifies a training observation into its own species.
load fisheriris Mdl = fitcnb (meas, species)
Mdl =
ClassificationNaiveBayes
ResponseName: Y
CategoricalPredictors: []
ClassNames: {'setosa' 'versicolor' 'virginica'}
ScoreTransform: none
NumObservations: 150
DistributionNames: {'normal' 'normal' 'normal' 'normal'}
DistributionParameters: {3x4 cell}
printf ("resubstitution loss: %g\n", resubLoss (Mdl));
resubstitution loss: 0.04
The petal measurements separate the species far better than the sepal ones, and a kernel density follows a skewed predictor where a normal one cannot.
load fisheriris
normalMdl = fitcnb (meas, species);
kernelMdl = fitcnb (meas, species, 'DistributionNames', 'kernel');
printf ("normal : %g\n", resubLoss (normalMdl));
normal : 0.04
printf ("kernel : %g\n", resubLoss (kernelMdl));
kernel : 0.0333333
Fit from a table, and predict on one
load fisheriris
T = table (meas(:,1), meas(:,2), meas(:,3), meas(:,4), ...
'VariableNames', {'SL', 'SW', 'PL', 'PW'});
T.Species = categorical (species);
A column holding levels is a categorical predictor without being named one
T.Wide = categorical (meas(:,2) > 3, [false true], {'narrow', 'wide'});
The response is named by its column, and everything else is a predictor
Mdl = fitcnb (T, 'Species'); Mdl.PredictorNames
ans =
1x5 cell array
{'SL'} {'SW'} {'PL'} {'PW'} {'Wide'}
Mdl.CategoricalPredictors
ans = 5
A model formula names them instead, holding main effects only
Mdl2 = fitcnb (T, 'Species ~ PL + PW'); Mdl2.PredictorNames
ans =
1x2 cell array
{'PL'} {'PW'}
predict reads a table by the names the model was fitted on, so the columns may come in any order
label = predict (Mdl, T(1:5, [6, 5, 4, 3, 2, 1])); label'
ans =
1x5 categorical array
setosa setosa setosa setosa setosa