IDEAS home Printed from https://ideas.repec.org/a/oup/biomet/v96y2009i2p469-478.html
   My bibliography  Save this article

Scale adjustments for classifiers in high-dimensional, low sample size settings

Author

Listed:
  • Yao-Ban Chan
  • Peter Hall

Abstract

Distance-based classifiers are generally considered to be effective at discriminating between populations that differ in location. Indeed, nearest-neighbour methods and the support vector machine are frequently used in very high-dimensional problems involving gene expression data, where it is believed that elevated levels of expression convey much of the information for classification. However, one problem inherent to distance-based classifiers is that scale differences can mask location differences. In consequence, such classifiers can have poor performance if the information for classification accumulates through a large number of relatively small location differences in data components, rather than via large differences. In this paper, we show that a simple adjustment for scale, applicable to a variety of distance-based classifiers, can remedy the problem. For some classifiers, such as those based on the support vector machine or the centroid method, scale corrections are important primarily in the case of small training-sample sizes. However, for other classifiers, including those based on nearest-neighbour and average-distance methods, scale adjustments are helpful more generally. Copyright 2009, Oxford University Press.

Suggested Citation

  • Yao-Ban Chan & Peter Hall, 2009. "Scale adjustments for classifiers in high-dimensional, low sample size settings," Biometrika, Biometrika Trust, vol. 96(2), pages 469-478.
  • Handle: RePEc:oup:biomet:v:96:y:2009:i:2:p:469-478
    as

    Download full text from publisher

    File URL: http://hdl.handle.net/10.1093/biomet/asp007
    Download Restriction: Access to full text is restricted to subscribers.
    ---><---

    As the access to this document is restricted, you may want to search for a different version of it.

    Citations

    Citations are extracted by the CitEc Project, subscribe to its RSS feed for this item.
    as


    Cited by:

    1. Makoto Aoshima & Kazuyoshi Yata, 2014. "A distance-based, misclassification rate adjusted classifier for multiclass, high-dimensional data," Annals of the Institute of Statistical Mathematics, Springer;The Institute of Statistical Mathematics, vol. 66(5), pages 983-1010, October.
    2. Rauf Ahmad, M. & Pavlenko, Tatjana, 2018. "A U-classifier for high-dimensional data under non-normality," Journal of Multivariate Analysis, Elsevier, vol. 167(C), pages 269-283.
    3. Makoto Aoshima & Kazuyoshi Yata, 2019. "High-Dimensional Quadratic Classifiers in Non-sparse Settings," Methodology and Computing in Applied Probability, Springer, vol. 21(3), pages 663-682, September.
    4. Makoto Aoshima & Kazuyoshi Yata, 2019. "Distance-based classifier by data transformation for high-dimension, strongly spiked eigenvalue models," Annals of the Institute of Statistical Mathematics, Springer;The Institute of Statistical Mathematics, vol. 71(3), pages 473-503, June.
    5. Ishii, Aki & Yata, Kazuyoshi & Aoshima, Makoto, 2022. "Geometric classifiers for high-dimensional noisy data," Journal of Multivariate Analysis, Elsevier, vol. 188(C).
    6. Yugo Nakayama & Kazuyoshi Yata & Makoto Aoshima, 2020. "Bias-corrected support vector machine with Gaussian kernel in high-dimension, low-sample-size settings," Annals of the Institute of Statistical Mathematics, Springer;The Institute of Statistical Mathematics, vol. 72(5), pages 1257-1286, October.

    More about this item

    Statistics

    Access and download statistics

    Corrections

    All material on this site has been provided by the respective publishers and authors. You can help correct errors and omissions. When requesting a correction, please mention this item's handle: RePEc:oup:biomet:v:96:y:2009:i:2:p:469-478. See general information about how to correct material in RePEc.

    If you have authored this item and are not yet registered with RePEc, we encourage you to do it here. This allows to link your profile to this item. It also allows you to accept potential citations to this item that we are uncertain about.

    We have no bibliographic references for this item. You can help adding them by using this form .

    If you know of missing items citing this one, you can help us creating those links by adding the relevant references in the same way as above, for each refering item. If you are a registered author of this item, you may also want to check the "citations" tab in your RePEc Author Service profile, as there may be some citations waiting for confirmation.

    For technical questions regarding this item, or to correct its authors, title, abstract, bibliographic or download information, contact: Oxford University Press (email available below). General contact details of provider: https://academic.oup.com/biomet .

    Please note that corrections may take a couple of weeks to filter through the various RePEc services.

    IDEAS is a RePEc service. RePEc uses bibliographic data supplied by the respective publishers.