nomdist
statistics: D = nomdist (X)
statistics: D = nomdist (X, measure)
statistics: D = nomdist (…, Name, Value)
Dissimilarity between every pair of rows of nominal data.
D = nomdist (X) returns the Goodall 3 dissimilarity
between each pair of rows of X, whose columns are nominal variables
and whose values are their levels. D is a row vector in the order
pdist returns, pairs (1,2), (1,3), …,
(2,3), …, so squareform turns it into a matrix and
linkage takes it as it stands. An X of one row gives a
1×0 D.
X is a numeric, logical, categorical, string or cellstr matrix, each
distinct value being a level, or a table whose every variable is one column
of these. A missing value, NaN, an empty character vector, a
missing string or an undefined category, is refused.
The measures weigh a match or a mismatch by how often the levels involved
occur, counted over the rows of X unless 'Reference' names
another sample.
D = nomdist (X, measure) uses the measure named,
in any case, from those of Boriah, Chandola and Kumar (2008) and of Sulc and
Rezankova (2019), as the R package nomclust 2.8.1 implements them.
In the table, N is the number of rows the frequencies are counted
over, f the count of a level, p = f/N its frequency,
p2 = f(f-1)/(N(N-1)) the chance of drawing it twice, and n
the number of levels the variable holds. A subscript names the level of
either row, and a variable scores 0 wherever its entry names no score.
| measure | Score of one variable | |
|---|---|---|
'anderberg' | The similarity is \(\frac{\sum_{x = y} c/p_x^2} {\sum_{x = y} c/p_x^2 + \sum_{x \ne y} c/(2 p_x p_y)},\) with \(c = 2/(n(n+1)).\) | |
'burnaby' | 1 on a match; on a mismatch \(\frac{B}{\log \frac{p_x p_y}{(1-p_x)(1-p_y)} + B},\) with \(B = \sum_q 2 \log (1 - p_q).\) | |
'eskin' | 1 on a match, \(\frac{n^2}{n^2 + 2}\) on a mismatch. | |
'gambaryan' | \(-\left(p_x \log_2 p_x + (1-p_x) \log_2 (1-p_x)\right)\) on a match, which peaks where p is 0.5. | |
'goodall1' | On a match \(1 - \sum_{q:\, p_q \le p_x} \mathit{p2}_q.\) | |
'goodall2' | On a match \(1 - \sum_{q:\, p_q \ge p_x} \mathit{p2}_q.\) | |
'goodall3' | \(1 - \mathit{p2}_x\) on a match. | |
'goodall4' | \(\mathit{p2}_x\) on a match. | |
'iof' | 1 on a match, \(\frac{1}{1 + \log f_x \log f_y}\) on a mismatch. | |
'lin' | \(2 \log p_x\) on a match, \(2 \log (p_x + p_y)\) on a mismatch, over the sum of \(\log p_x + \log p_y.\) | |
'lin1' | On a match
\(\log \sum_{q:\, p_q = p_x} p_q;\)
on a mismatch
\(2 \log \sum_{q:\, \min (p_x, p_y) \le p_q \le \max (p_x, p_y)} p_q;\)
over the same sum as 'lin'. | |
'of' | 1 on a match, \(\frac{1}{1 + \log (N/f_x) \log (N/f_y)}\) on a mismatch. | |
'sm' | 1 on a match, the simple matching coefficient. | |
'smirnov' | On a match \(2 + \frac{N - f_x}{f_x} + \sum_{q \ne x} \frac{f_q}{N - f_q};\) on a mismatch \(\sum_{q \ne x, y} \frac{f_q}{N - f_q}.\) | |
've' | On a match the entropy of the variable, \(-\frac{1}{\log n} \sum_q p_q \log p_q.\) | |
'vm' | On a match the Gini index of the variable, \(\frac{n}{n-1} \left(1 - \sum_q p_q^2\right).\) |
The scores of a pair are summed over the variables, each times its weight,
and divided by the sum of the weights, except where the table names another
divisor, and 'gambaryan' and 'smirnov' divide by the number
of levels over every variable. With S that similarity,
'eskin', 'iof', 'lin', 'lin1' and
'of' return 1/S - 1, 'smirnov' returns
1/(1+S), and every other measure 1 - S. Several measures,
the four Goodall ones among them, place two identical rows at a nonzero
dissimilarity, a match on a common level saying less than one on a rare
level. 'lin' and 'lin1' return Inf for a pair of
similarity 0, such as two rows differing in every variable where each
variable holds two levels.
D = nomdist (…, Name, Value) takes the
following options.
| Name | Value | |
|---|---|---|
'Weights' | One weight per variable, each from 0 to
1 and at least one positive; a variable of weight 0 is left out. Not taken
by 'anderberg', 'gambaryan' and 'smirnov'. | |
'Reference' | A sample to count level frequencies over in place of X, of the same kind: a matrix with as many columns, the same type in each, or a table holding every variable of X, read by name. It must hold every level X holds. Given the whole data, it measures a subsample by how rare its levels are in the whole. |
Where nomclust answers otherwise, the answer here is the limit of
the measure’s formula. 'lin' and 'lin1' sum the counts
of levels, so a sum over every level is exactly 1 and a similarity of 0 is
exactly 0, where nomclust rounds it to about 1e16 or
replaces it with one more than the largest finite dissimilarity.
'gambaryan' scores a level every row holds as 0, where
nomclust returns NaN. Where every variable holds one level,
'lin' returns 0 and 'lin1' 1, where nomclust
returns NaN.
References:
See also: nomdist2, pdist, squareform, linkage
Source Code: nomdist