Skip to main navigation Skip to search Skip to main content

Extended Methods to Handle Classification Biases

Research output: Chapter in Book / Report / Conference proceedingConference contributionAcademicpeer-review

Abstract

Classifiers can provide counts of items per class, but systematic classification errors yield biases (e.g., if a class is often misclassified as another, its size may be under-estimated). To handle classification biases, the statistics and epidemiology domains devised methods for estimating unbiased class sizes (or class probabilities) without identifying which individual items are misclassified. These bias correction methods are applicable to machine learning classifiers, but in some cases yield high result variance and increased biases. We present the applicability and drawbacks of existing methods and extend them with three novel methods. Our Sample-to-Sample method provides accurate confidence intervals for the bias correction results. Our Maximum Determinant method predicts which classifier yields the least result variance. Our Ratio-to-TP method details the error decomposition in classifier outputs (i.e., how many items classified as class C_y truly belong to C_x, for all possible classes) and has properties of interest for applying the Maximum Determinant method. Our methods are demonstrated empirically, and we discuss the need for establishing theory and guidelines for choosing the methods and classifier to apply.
Original languageEnglish
Title of host publication2017 IEEE International Conference on Data Science and Advanced Analytics (DSAA)
Subtitle of host publication[Proceedings]
PublisherIEEE
Pages765-774
Number of pages10
ISBN (Electronic)9781509050048
ISBN (Print)9781509050055
DOIs
Publication statusPublished - 2017

Funding

We are grateful to Arjen P. de Vries, Nishant Mehta, Erik Quaeghebeur and Rebecca Holman for their invaluable remarks and suggestions. Part of this research was funded by the Fish4Knowledge project (EU FP7 Grant 257024).

FundersFunder number
Seventh Framework Programme257024
Seventh Framework Programme

    Fingerprint

    Dive into the research topics of 'Extended Methods to Handle Classification Biases'. Together they form a unique fingerprint.

    Cite this