stats.drift.DriftDiagnostics
statistics: stats.drift.DriftDiagnostics
The drift found between two data sets, as detectdrift returns it.
A stats.drift.DriftDiagnostics object holds, for each variable
compared, the metric measuring how far the target data has moved from the
baseline data, the p-value of the permutation test of that change with its
confidence interval, and the drift status the interval gives, together with
the drift status of the data as a whole. Every property is read-only.
detectdrift is the documented way to create one.
summary tabulates the results, ecdf and histcounts
return the distributions compared, and plotDriftStatus,
plotEmpiricalCDF, plotHistogram and
plotPermutationResults draw them.
See also: detectdrift
Source Code: stats.drift.DriftDiagnostics
The stats.drift.DriftDiagnostics class contains the following properties:
The stats.drift.DriftDiagnostics class offers the following public methods:
stats.drift.DriftDiagnostics: summary (obj)
stats.drift.DriftDiagnostics: tbl = summary (obj)
Where p-values were estimated, summary prints the drift status
of the data as a whole and a table with one row per variable holding
DriftStatus, PValue and ConfidenceInterval.
tbl is that table with a last row, MultipleTest, holding
the status of the data as a whole and NaN for its p-value and
interval. Where they were not, the table holds MetricValue and
Metric.
See also: detectdrift
stats.drift.DriftDiagnostics: tbl = ecdf (obj)
stats.drift.DriftDiagnostics: tbl = ecdf (obj, 'Variable', var)
tbl holds one row per variable, named by it, and three variables
of cells: x, the pooled values sorted with the first repeated,
and F_Baseline and F_Target, the two distribution
functions there, starting from 0. Where values tie, only the last copy
takes the full value, so that the points draw the two functions as
steps. A variable holding levels has NaN in each cell.
var names the variables, by name or index, all of them by
default.
stats.drift.DriftDiagnostics: tbl = histcounts (obj)
stats.drift.DriftDiagnostics: tbl = histcounts (obj, 'Variable', var)
tbl holds one row per variable, named by it, and three variables
of cells: Bins, and Counts_Baseline and
Counts_Target, each the percentage of its sample in each bin.
For a continuous variable Bins holds the bin edges
histcounts chooses for the two samples pooled. For a variable
holding levels it holds the levels as a categorical array, and
each count is increased by 0.5 before the percentages are taken, as the
metrics take them. var names the variables, by name or index,
all of them by default.
stats.drift.DriftDiagnostics: plotDriftStatus (obj)
stats.drift.DriftDiagnostics: h = plotDriftStatus (obj)
Each variable is drawn as its p-value with its confidence interval as
horizontal error bars, one set of bars for each drift status, the
variables in alphabetical order up the vertical axis, with a line at
each threshold. h holds the three sets of bars, for
"Stable", "Warning" and "Drift" in turn, each
spanning every variable with NaN where a variable has another
status. MATLAB draws on a categorical axis, which Octave does not have,
so the variables sit at 1, 2, … with their names as tick labels.
See also: stats.drift.DriftDiagnostics.summary
stats.drift.DriftDiagnostics: plotEmpiricalCDF (obj)
stats.drift.DriftDiagnostics: plotEmpiricalCDF (obj, 'Variable', var)
stats.drift.DriftDiagnostics: h = plotEmpiricalCDF (…)
The baseline and the target are drawn as steps, from the points
ecdf returns. var names the variable, by name or index;
by default it is the one with the smallest p-value, or the first where
p-values were not estimated. h holds the two lines, baseline
first.
See also: stats.drift.DriftDiagnostics.ecdf
stats.drift.DriftDiagnostics: plotHistogram (obj)
stats.drift.DriftDiagnostics: plotHistogram (obj, 'Variable', var)
stats.drift.DriftDiagnostics: h = plotHistogram (…)
The percentages histcounts returns are drawn as bars side by
side, over the bins of a continuous variable or the levels of one
holding levels. var names the variable, by name or index; by
default it is the one with the smallest p-value, or the first where
p-values were not estimated. h holds the two sets of bars,
baseline first.
See also: stats.drift.DriftDiagnostics.histcounts
stats.drift.DriftDiagnostics: plotPermutationResults (obj)
stats.drift.DriftDiagnostics: plotPermutationResults (obj, 'Variable', var)
stats.drift.DriftDiagnostics: h = plotPermutationResults (…)
The percentage of permutations in each bin is drawn as bars, split at
the observed metric, which is marked by a line: the bars below it and
those at or above it, whose share is the p-value. var names the
variable, by name or index; by default it is the one with the smallest
p-value. h holds the two sets of bars, those below first.
MATLAB draws them as histogram objects, which Octave does not
have.
See also: detectdrift