Categories &

Functions List

Class Definition: stats.drift.DriftDiagnostics

statistics: stats.drift.DriftDiagnostics

The drift found between two data sets, as detectdrift returns it.

A stats.drift.DriftDiagnostics object holds, for each variable compared, the metric measuring how far the target data has moved from the baseline data, the p-value of the permutation test of that change with its confidence interval, and the drift status the interval gives, together with the drift status of the data as a whole. Every property is read-only. detectdrift is the documented way to create one.

summary tabulates the results, ecdf and histcounts return the distributions compared, and plotDriftStatus, plotEmpiricalCDF, plotHistogram and plotPermutationResults draw them.

See also: detectdrift

Source Code: stats.drift.DriftDiagnostics

The stats.drift.DriftDiagnostics class contains the following properties:

The stats.drift.DriftDiagnostics class offers the following public methods:

stats.drift.DriftDiagnostics: summary (obj)

stats.drift.DriftDiagnostics: tbl = summary (obj)

Where p-values were estimated, summary prints the drift status of the data as a whole and a table with one row per variable holding DriftStatus, PValue and ConfidenceInterval. tbl is that table with a last row, MultipleTest, holding the status of the data as a whole and NaN for its p-value and interval. Where they were not, the table holds MetricValue and Metric.

See also: detectdrift

stats.drift.DriftDiagnostics: tbl = ecdf (obj)

stats.drift.DriftDiagnostics: tbl = ecdf (obj, 'Variable', var)

tbl holds one row per variable, named by it, and three variables of cells: x, the pooled values sorted with the first repeated, and F_Baseline and F_Target, the two distribution functions there, starting from 0. Where values tie, only the last copy takes the full value, so that the points draw the two functions as steps. A variable holding levels has NaN in each cell. var names the variables, by name or index, all of them by default.

See also: stats.drift.DriftDiagnostics.plotEmpiricalCDF

stats.drift.DriftDiagnostics: tbl = histcounts (obj)

stats.drift.DriftDiagnostics: tbl = histcounts (obj, 'Variable', var)

tbl holds one row per variable, named by it, and three variables of cells: Bins, and Counts_Baseline and Counts_Target, each the percentage of its sample in each bin. For a continuous variable Bins holds the bin edges histcounts chooses for the two samples pooled. For a variable holding levels it holds the levels as a categorical array, and each count is increased by 0.5 before the percentages are taken, as the metrics take them. var names the variables, by name or index, all of them by default.

See also: stats.drift.DriftDiagnostics.plotHistogram

stats.drift.DriftDiagnostics: plotDriftStatus (obj)

stats.drift.DriftDiagnostics: h = plotDriftStatus (obj)

Each variable is drawn as its p-value with its confidence interval as horizontal error bars, one set of bars for each drift status, the variables in alphabetical order up the vertical axis, with a line at each threshold. h holds the three sets of bars, for "Stable", "Warning" and "Drift" in turn, each spanning every variable with NaN where a variable has another status. MATLAB draws on a categorical axis, which Octave does not have, so the variables sit at 1, 2, … with their names as tick labels.

See also: stats.drift.DriftDiagnostics.summary

stats.drift.DriftDiagnostics: plotEmpiricalCDF (obj)

stats.drift.DriftDiagnostics: plotEmpiricalCDF (obj, 'Variable', var)

stats.drift.DriftDiagnostics: h = plotEmpiricalCDF (…)

The baseline and the target are drawn as steps, from the points ecdf returns. var names the variable, by name or index; by default it is the one with the smallest p-value, or the first where p-values were not estimated. h holds the two lines, baseline first.

See also: stats.drift.DriftDiagnostics.ecdf

stats.drift.DriftDiagnostics: plotHistogram (obj)

stats.drift.DriftDiagnostics: plotHistogram (obj, 'Variable', var)

stats.drift.DriftDiagnostics: h = plotHistogram (…)

The percentages histcounts returns are drawn as bars side by side, over the bins of a continuous variable or the levels of one holding levels. var names the variable, by name or index; by default it is the one with the smallest p-value, or the first where p-values were not estimated. h holds the two sets of bars, baseline first.

See also: stats.drift.DriftDiagnostics.histcounts

stats.drift.DriftDiagnostics: plotPermutationResults (obj)

stats.drift.DriftDiagnostics: plotPermutationResults (obj, 'Variable', var)

stats.drift.DriftDiagnostics: h = plotPermutationResults (…)

The percentage of permutations in each bin is drawn as bars, split at the observed metric, which is marked by a line: the bars below it and those at or above it, whose share is the p-value. var names the variable, by name or index; by default it is the one with the smallest p-value. h holds the two sets of bars, those below first. MATLAB draws them as histogram objects, which Octave does not have.

See also: detectdrift