P-values for classification

dc.creatorDuembgen, Lutz
dc.creatorIgl, Bernd-Wolfgang
dc.creatorMunk, Axel
dc.date2008-01-18
dc.date2008-06-26
dc.date.accessioned2026-07-07T09:46:31Z
dc.date.available2026-07-07T09:46:31Z
dc.descriptionLet $(X,Y)$ be a random variable consisting of an observed feature vector $X\in \mathcal{X}$ and an unobserved class label $Y\in \{1,2,...,L\}$ with unknown joint distribution. In addition, let $\mathcal{D}$ be a training data set consisting of $n$ completely observed independent copies of $(X,Y)$. Usual classification procedures provide point predictors (classifiers) $\widehat{Y}(X,\mathcal{D})$ of $Y$ or estimate the conditional distribution of $Y$ given $X$. In order to quantify the certainty of classifying $X$ we propose to construct for each $θ=1,2,...,L$ a p-value $π_θ(X,\mathcal{D})$ for the null hypothesis that $Y=θ$, treating $Y$ temporarily as a fixed parameter. In other words, the point predictor $\widehat{Y}(X,\mathcal{D})$ is replaced with a prediction region for $Y$ with a certain confidence. We argue that (i) this approach is advantageous over traditional approaches and (ii) any reasonable classifier can be modified to yield nonparametric p-values. We discuss issues such as optimality, single use and multiple use validity, as well as computational and graphical aspects.
dc.descriptionPublished in at http://dx.doi.org/10.1214/08-EJS245 the Electronic Journal of Statistics (http://www.i-journals.org/ejs/) by the Institute of Mathematical Statistics (http://www.imstat.org)
dc.identifierhttps://arxiv.org/abs/0801.2934
dc.identifierhttp://arxiv.org/abs/0801.2934
dc.identifierElectronic Journal of Statistics 2008, Vol. 2, 468-493
dc.identifierdoi:10.1214/08-EJS245
dc.identifier.urihttp://salesiana.dossiersoluciones.com/handle/123456789/163557
dc.subjectStatistics Theory
dc.subjectMachine Learning
dc.subject62C05, 62F25, 62G09, 62G15, 62H30 (Primary)
dc.titleP-values for classification
dc.typetext

Files

Collections