Categories &

Functions List

Function Reference: detectdrift

statistics: DDiagnostics = detectdrift (Baseline, Target)

statistics: DDiagnostics = detectdrift (Baseline, Target, Name, Value)

Detect drift between two data sets, one variable at a time.

DDiagnostics = detectdrift (Baseline, Target) measures, for each variable, how far the distribution in Target has moved from the one in Baseline, tests the change by permutation, and returns the results as a stats.drift.DriftDiagnostics object. Baseline and Target are both numeric matrices with the same number of columns, both categorical arrays, or both tables, read over the variables they share, which must be all the variables of one of them. An observation holding a missing value in a variable used is left out.

Each variable is measured by a metric: a continuous one by 'ContinuousMetric', one holding levels by 'CategoricalMetric'. With F and G the empirical distribution functions of the two samples and H that of the two pooled, the metrics are:

MetricDefinition
'wasserstein'The area between F and G, the default for a continuous variable.
'ks'The largest difference between F and G.
'ad'The Anderson-Darling statistic, the sum over the distinct pooled values but the largest of (F - G)^2 / (H (1 - H)), divided by the number of pooled observations.
'energy'The energy distance, sqrt (2 E|x - y| - E|x - x'| - E|y - y'|).
'hellinger'sqrt (1 - sum (sqrt (p q))), the default for a variable holding levels.
'bhattacharyya'-log (sum (sqrt (p q))).
'tv'The total variation distance, sum (abs (p - q)) / 2.
'psi'The population stability index, sum ((p - q) log (p / q)).
'chi2'sum ((p - q)^2 / p).

For a variable holding levels, p and q are the shares of each level in Baseline and Target, over the levels either holds, each count increased by 0.5 first so that a level one sample lacks leaves every metric finite.

The p-value of a variable is the share of permutations whose metric is at least the observed one, the observed arrangement counted as the first permutation, so it is never below one over their number. Its 95% confidence interval is the Clopper-Pearson interval. The drift status is "Drift" where the interval lies below 'DriftThreshold', "Stable" where it lies above 'WarningThreshold', and "Warning" otherwise. Permutations are drawn in stages, 1000 or 'MaxNumPermutations' if fewer, then four times as many at each stage, until the interval lies within a single one of the three regions or the next stage would pass 'MaxNumPermutations'. They are drawn with randperm, so the state of rand decides them.

The drift status of the data as a whole comes from the p-values corrected for testing every variable: by Bonferroni, the smallest p-value times the number of variables, or by the false discovery rate, the smallest Benjamini-Hochberg adjusted p-value. It is "Drift" below 'DriftThreshold', "Warning" below 'WarningThreshold', and "Stable" otherwise.

DDiagnostics = detectdrift (…, Name, Value) takes the following options.

NameValue
'VariableNames'The variables to use, among those the two tables share. All of them by default.
'CategoricalVariables'The variables holding levels: 'all', their indices, a logical vector over the variables, or, for tables, their names. A table variable holding logical values, an unordered categorical array, a string array or a cell array of character vectors, and every column of a categorical array, holds levels whether named here or not.
'ContinuousMetric''wasserstein', the default, 'ks', 'ad' or 'energy'.
'CategoricalMetric''hellinger', the default, 'bhattacharyya', 'tv', 'psi' or 'chi2'.
'DriftThreshold'0.05 by default.
'WarningThreshold'0.1 by default, and above 'DriftThreshold'.
'MultipleTestCorrection''bonferroni', the default, or 'fdr'.
'MaxNumPermutations'The most permutations drawn for a variable, 1000 by default.
'EstimatePValues'true by default; with false only the metrics are computed.
'Options'A structure as statset returns. Parallel computing and random streams are not implemented, and are refused where asked for.

MATLAB draws at least 1000 permutations whatever 'MaxNumPermutations' says, and finishes a stage that passes it, so that a maximum of 3000 can end at 4001; here no variable is given more permutations than 'MaxNumPermutations'. MATLAB’s stages are otherwise its own, so the number of permutations and the p-values differ from MATLAB’s, as they would between any two random draws.

See also: stats.drift.DriftDiagnostics, knntest, mmdtest

Source Code: detectdrift