detectdrift
statistics: DDiagnostics = detectdrift (Baseline, Target)
statistics: DDiagnostics = detectdrift (Baseline, Target, Name, Value)
Detect drift between two data sets, one variable at a time.
DDiagnostics = detectdrift (Baseline, Target)
measures, for each variable, how far the distribution in Target has
moved from the one in Baseline, tests the change by permutation, and
returns the results as a stats.drift.DriftDiagnostics object.
Baseline and Target are both numeric matrices with the same
number of columns, both categorical arrays, or both tables, read
over the variables they share, which must be all the variables of one of
them. An observation holding a missing value in a variable used is left
out.
Each variable is measured by a metric: a continuous one by
'ContinuousMetric', one holding levels by
'CategoricalMetric'. With F and G the empirical
distribution functions of the two samples and H that of the two
pooled, the metrics are:
| Metric | Definition | |
|---|---|---|
'wasserstein' | The area between F and G, the default for a continuous variable. | |
'ks' | The largest difference between F and G. | |
'ad' | The Anderson-Darling statistic, the sum over the distinct pooled values but the largest of (F - G)^2 / (H (1 - H)), divided by the number of pooled observations. | |
'energy' | The energy distance, sqrt (2 E|x - y| - E|x - x'| - E|y - y'|). | |
'hellinger' | sqrt (1 - sum (sqrt (p q))), the default for a variable holding levels. | |
'bhattacharyya' | -log (sum (sqrt (p q))). | |
'tv' | The total variation distance, sum (abs (p - q)) / 2. | |
'psi' | The population stability index, sum ((p - q) log (p / q)). | |
'chi2' | sum ((p - q)^2 / p). |
For a variable holding levels, p and q are the shares of each level in Baseline and Target, over the levels either holds, each count increased by 0.5 first so that a level one sample lacks leaves every metric finite.
The p-value of a variable is the share of permutations whose metric is at
least the observed one, the observed arrangement counted as the first
permutation, so it is never below one over their number. Its 95%
confidence interval is the Clopper-Pearson interval. The drift status is
"Drift" where the interval lies below 'DriftThreshold',
"Stable" where it lies above 'WarningThreshold', and
"Warning" otherwise. Permutations are drawn in stages, 1000 or
'MaxNumPermutations' if fewer, then four times as many at each
stage, until the interval lies within a single one of the three regions or
the next stage would pass 'MaxNumPermutations'. They are drawn with
randperm, so the state of rand decides them.
The drift status of the data as a whole comes from the p-values corrected
for testing every variable: by Bonferroni, the smallest p-value times the
number of variables, or by the false discovery rate, the smallest
Benjamini-Hochberg adjusted p-value. It is "Drift" below
'DriftThreshold', "Warning" below
'WarningThreshold', and "Stable" otherwise.
DDiagnostics = detectdrift (…, Name, Value)
takes the following options.
| Name | Value | |
|---|---|---|
'VariableNames' | The variables to use, among those the two tables share. All of them by default. | |
'CategoricalVariables' | The variables holding
levels: 'all', their indices, a logical vector over the variables,
or, for tables, their names. A table variable holding logical values, an
unordered categorical array, a string array or a cell array of
character vectors, and every column of a categorical array, holds
levels whether named here or not. | |
'ContinuousMetric' | 'wasserstein', the
default, 'ks', 'ad' or 'energy'. | |
'CategoricalMetric' | 'hellinger', the
default, 'bhattacharyya', 'tv', 'psi' or
'chi2'. | |
'DriftThreshold' | 0.05 by default. | |
'WarningThreshold' | 0.1 by default, and above
'DriftThreshold'. | |
'MultipleTestCorrection' | 'bonferroni', the
default, or 'fdr'. | |
'MaxNumPermutations' | The most permutations drawn for a variable, 1000 by default. | |
'EstimatePValues' | true by default; with
false only the metrics are computed. | |
'Options' | A structure as statset returns.
Parallel computing and random streams are not implemented, and are refused
where asked for. |
MATLAB draws at least 1000 permutations whatever
'MaxNumPermutations' says, and finishes a stage that passes it, so
that a maximum of 3000 can end at 4001; here no variable is given more
permutations than 'MaxNumPermutations'. MATLAB’s stages are
otherwise its own, so the number of permutations and the p-values differ
from MATLAB’s, as they would between any two random draws.
See also: stats.drift.DriftDiagnostics, knntest, mmdtest
Source Code: detectdrift