IDEAS home Printed from https://ideas.repec.org/a/eee/infome/v10y2016i2p454-470.html
   My bibliography  Save this article

Are the discretised lognormal and hooked power law distributions plausible for citation data?

Author

Listed:
  • Thelwall, Mike

Abstract

There is no agreement over which statistical distribution is most appropriate for modelling citation count data. This is important because if one distribution is accepted then the relative merits of different citation-based indicators, such as percentiles, arithmetic means and geometric means, can be more fully assessed. In response, this article investigates the plausibility of the discretised lognormal and hooked power law distributions for modelling the full range of citation counts, with an offset of 1. The citation counts from 23 Scopus subcategories were fitted to hooked power law and discretised lognormal distributions but both distributions failed a Kolmogorov–Smirnov goodness of fit test in over three quarters of cases. The discretised lognormal distribution also seems to have the wrong shape for citation distributions, with too few zeros and not enough medium values for all subjects. The cause of poor fits could be the impurity of the subject subcategories or the presence of interdisciplinary research. Although it is possible to test for subject subcategory purity indirectly through a goodness of fit test in theory with large enough sample sizes, it is probably not possible in practice. Hence it seems difficult to get conclusive evidence about the theoretically most appropriate statistical distribution.

Suggested Citation

  • Thelwall, Mike, 2016. "Are the discretised lognormal and hooked power law distributions plausible for citation data?," Journal of Informetrics, Elsevier, vol. 10(2), pages 454-470.
  • Handle: RePEc:eee:infome:v:10:y:2016:i:2:p:454-470
    DOI: 10.1016/j.joi.2016.03.001
    as

    Download full text from publisher

    File URL: http://www.sciencedirect.com/science/article/pii/S1751157716300074
    Download Restriction: Full text for ScienceDirect subscribers only

    File URL: https://libkey.io/10.1016/j.joi.2016.03.001?utm_source=ideas
    LibKey link: if access is restricted and if your library uses this service, LibKey will redirect you to where you can use your library subscription to access this item
    ---><---

    As the access to this document is restricted, you may want to search for a different version of it.

    References listed on IDEAS

    as
    1. T. S. Evans & N. Hopkins & B. S. Kaube, 2012. "Universality of performance indicators based on citation and reference counts," Scientometrics, Springer;Akadémiai Kiadó, vol. 93(2), pages 473-495, November.
    2. Thelwall, Mike, 2016. "The discretised lognormal and hooked power law distributions for complete citation data: Best options for modelling and regression," Journal of Informetrics, Elsevier, vol. 10(2), pages 336-346.
    3. Thelwall, Mike & Wilson, Paul, 2014. "Regression for citation data: An evaluation of different methods," Journal of Informetrics, Elsevier, vol. 8(4), pages 963-971.
    4. Vincent Larivière & Yves Gingras, 2010. "On the relationship between interdisciplinarity and scientific impact," Journal of the Association for Information Science & Technology, Association for Information Science & Technology, vol. 61(1), pages 126-131, January.
    5. Filippo Radicchi & Claudio Castellano, 2012. "A Reverse Engineering Approach to the Suppression of Citation Biases Reveals Universal Properties of Citation Distributions," PLOS ONE, Public Library of Science, vol. 7(3), pages 1-9, March.
    6. Michael D. Gordon, 1982. "Citation ranking versus subjective evaluation in the determination of journal hierachies in the social sciences," Journal of the American Society for Information Science, Association for Information Science & Technology, vol. 33(1), pages 55-57, January.
    7. Waltman, Ludo & van Eck, Nees Jan & van Leeuwen, Thed N. & Visser, Martijn S. & van Raan, Anthony F.J., 2011. "Towards a new crown indicator: Some theoretical considerations," Journal of Informetrics, Elsevier, vol. 5(1), pages 37-47.
    8. Michel Zitt, 2012. "The journal impact factor: angel, devil, or scapegoat? A comment on J.K. Vanclay’s article 2011," Scientometrics, Springer;Akadémiai Kiadó, vol. 92(2), pages 485-503, August.
    9. Éric Archambault & David Campbell & Yves Gingras & Vincent Larivière, 2009. "Comparing bibliometric statistics obtained from the Web of Science and Scopus," Journal of the American Society for Information Science and Technology, Association for Information Science & Technology, vol. 60(7), pages 1320-1326, July.
    10. Vuong, Quang H, 1989. "Likelihood Ratio Tests for Model Selection and Non-nested Hypotheses," Econometrica, Econometric Society, vol. 57(2), pages 307-333, March.
    11. Michael J. Stringer & Marta Sales-Pardo & Luís A. Nunes Amaral, 2010. "Statistical validation of a global model for the distribution of the ultimate number of citations accrued by papers published in a scientific journal," Journal of the Association for Information Science & Technology, Association for Information Science & Technology, vol. 61(7), pages 1377-1385, July.
    12. Vieira, E.S. & Gomes, J.A.N.F., 2010. "Citations to scientific articles: Its distribution and dependence on the article features," Journal of Informetrics, Elsevier, vol. 4(1), pages 1-13.
    13. Wilson, Paul, 2015. "The misuse of the Vuong test for non-nested models to test for zero-inflation," Economics Letters, Elsevier, vol. 127(C), pages 51-53.
    14. Vincent Larivière & Yves Gingras, 2010. "On the relationship between interdisciplinarity and scientific impact," Journal of the American Society for Information Science and Technology, Association for Information Science & Technology, vol. 61(1), pages 126-131, January.
    15. Derek De Solla Price, 1976. "A general theory of bibliometric and other cumulative advantage processes," Journal of the American Society for Information Science, Association for Information Science & Technology, vol. 27(5), pages 292-306, September.
    16. Ludo Waltman & Nees Jan van Eck & Anthony F. J. van Raan, 2012. "Universality of citation distributions revisited," Journal of the Association for Information Science & Technology, Association for Information Science & Technology, vol. 63(1), pages 72-77, January.
    17. Ludo Waltman & Nees Jan Eck & Thed N. Leeuwen & Martijn S. Visser & Anthony F. J. Raan, 2011. "Towards a new crown indicator: an empirical analysis," Scientometrics, Springer;Akadémiai Kiadó, vol. 87(3), pages 467-481, June.
    18. Young-Ho Eom & Santo Fortunato, 2011. "Characterizing and Modeling Citation Dynamics," PLOS ONE, Public Library of Science, vol. 6(9), pages 1-7, September.
    19. Pedro Albarrán & Antonio Perianes-Rodríguez & Javier Ruiz-Castillo, 2015. "Differences in citation impact across countries," Journal of the Association for Information Science & Technology, Association for Information Science & Technology, vol. 66(3), pages 512-525, March.
    20. Thelwall, Mike, 2016. "The precision of the arithmetic mean, geometric mean and percentiles for citation data: An experimental simulation modelling approach," Journal of Informetrics, Elsevier, vol. 10(1), pages 110-123.
    21. Gillespie, Colin S., 2015. "Fitting Heavy Tailed Distributions: The poweRlaw Package," Journal of Statistical Software, Foundation for Open Access Statistics, vol. 64(i02).
    22. López-Illescas, Carmen & de Moya-Anegón, Félix & Moed, Henk F., 2008. "Coverage and citation impact of oncological journals in the Web of Science and Scopus," Journal of Informetrics, Elsevier, vol. 2(4), pages 304-316.
    23. Wallace, Matthew L. & Larivière, Vincent & Gingras, Yves, 2009. "Modeling a century of citation distributions," Journal of Informetrics, Elsevier, vol. 3(4), pages 296-303.
    24. Ludo Waltman & Nees Jan van Eck & Anthony F. J. van Raan, 2012. "Universality of citation distributions revisited," Journal of the American Society for Information Science and Technology, Association for Information Science & Technology, vol. 63(1), pages 72-77, January.
    25. Thelwall, Mike & Wilson, Paul, 2014. "Distributions for cited articles from individual subjects and years," Journal of Informetrics, Elsevier, vol. 8(4), pages 824-839.
    Full references (including those not matched with items on IDEAS)

    Citations

    Citations are extracted by the CitEc Project, subscribe to its RSS feed for this item.
    as


    Cited by:

    1. Guillermo Armando Ronda-Pupo, 2017. "The citation-based impact of complex innovation systems scales with the size of the system," Scientometrics, Springer;Akadémiai Kiadó, vol. 112(1), pages 141-151, July.
    2. Banshal, Sumit Kumar & Gupta, Solanki & Lathabai, Hiran H & Singh, Vivek Kumar, 2022. "Power Laws in altmetrics: An empirical analysis," Journal of Informetrics, Elsevier, vol. 16(3).
    3. Vîiu, Gabriel-Alexandru, 2018. "The lognormal distribution explains the remarkable pattern documented by characteristic scores and scales in scientometrics," Journal of Informetrics, Elsevier, vol. 12(2), pages 401-415.
    4. Guillermo Armando Ronda-Pupo & J. Sylvan Katz, 2018. "The power law relationship between citation impact and multi-authorship patterns in articles in Information Science & Library Science journals," Scientometrics, Springer;Akadémiai Kiadó, vol. 114(3), pages 919-932, March.
    5. Gabriel-Alexandru Vȋiu & Mihai Păunescu, 2021. "The lack of meaningful boundary differences between journal impact factor quartiles undermines their independent use in research evaluation," Scientometrics, Springer;Akadémiai Kiadó, vol. 126(2), pages 1495-1525, February.
    6. Kaile Gong & Juan Xie & Ying Cheng & Vincent Larivière & Cassidy R. Sugimoto, 2019. "The citation advantage of foreign language references for Chinese social science papers," Scientometrics, Springer;Akadémiai Kiadó, vol. 120(3), pages 1439-1460, September.
    7. Gómez-Déniz, Emilio & Dorta-González, Pablo, 2024. "Modeling citation concentration through a mixture of Leimkuhler curves," Journal of Informetrics, Elsevier, vol. 18(2).
    8. Thelwall, Mike, 2016. "Citation count distributions for large monodisciplinary journals," Journal of Informetrics, Elsevier, vol. 10(3), pages 863-874.
    9. Thelwall, Mike, 2017. "Three practical field normalised alternative indicator formulae for research evaluation," Journal of Informetrics, Elsevier, vol. 11(1), pages 128-151.
    10. Mike Thelwall, 2019. "The influence of highly cited papers on field normalised indicators," Scientometrics, Springer;Akadémiai Kiadó, vol. 118(2), pages 519-537, February.
    11. Cena, Anna & Gagolewski, Marek & Siudem, Grzegorz & Żogała-Siudem, Barbara, 2022. "Validating citation models by proxy indices," Journal of Informetrics, Elsevier, vol. 16(2).
    12. Thelwall, Mike & Fairclough, Ruth, 2017. "The accuracy of confidence intervals for field normalised indicators," Journal of Informetrics, Elsevier, vol. 11(2), pages 530-540.
    13. Confraria, Hugo & Mira Godinho, Manuel & Wang, Lili, 2017. "Determinants of citation impact: A comparative analysis of the Global South versus the Global North," Research Policy, Elsevier, vol. 46(1), pages 265-279.
    14. Mike Thelwall & Kayvan Kousha, 2017. "ResearchGate versus Google Scholar: Which finds more early citations?," Scientometrics, Springer;Akadémiai Kiadó, vol. 112(2), pages 1125-1131, August.
    15. CholMyong Pak & Guang Yu & Weibin Wang, 2018. "A study on the citation situation within the citing paper: citation distribution of references according to mention frequency," Scientometrics, Springer;Akadémiai Kiadó, vol. 114(3), pages 905-918, March.
    16. Mrowinski, Maciej J. & Gagolewski, Marek & Siudem, Grzegorz, 2022. "Accidentality in journal citation patterns," Journal of Informetrics, Elsevier, vol. 16(4).
    17. Guillermo Armando Ronda-Pupo & J. Sylvan Katz, 2017. "The scaling relationship between degree centrality of countries and their citation-based performance on Management Information Systems," Scientometrics, Springer;Akadémiai Kiadó, vol. 112(3), pages 1285-1299, September.
    18. Thelwall, Mike, 2016. "Are there too many uncited articles? Zero inflated variants of the discretised lognormal and hooked power law distributions," Journal of Informetrics, Elsevier, vol. 10(2), pages 622-633.

    Most related items

    These are the items that most often cite the same works as this one and are cited by the same works as this one.
    1. Thelwall, Mike, 2016. "Are there too many uncited articles? Zero inflated variants of the discretised lognormal and hooked power law distributions," Journal of Informetrics, Elsevier, vol. 10(2), pages 622-633.
    2. Thelwall, Mike, 2016. "The precision of the arithmetic mean, geometric mean and percentiles for citation data: An experimental simulation modelling approach," Journal of Informetrics, Elsevier, vol. 10(1), pages 110-123.
    3. Thelwall, Mike, 2016. "Citation count distributions for large monodisciplinary journals," Journal of Informetrics, Elsevier, vol. 10(3), pages 863-874.
    4. Vîiu, Gabriel-Alexandru, 2018. "The lognormal distribution explains the remarkable pattern documented by characteristic scores and scales in scientometrics," Journal of Informetrics, Elsevier, vol. 12(2), pages 401-415.
    5. Thelwall, Mike & Sud, Pardeep, 2016. "National, disciplinary and temporal variations in the extent to which articles with more authors have more impact: Evidence from a geometric field normalised citation indicator," Journal of Informetrics, Elsevier, vol. 10(1), pages 48-61.
    6. Waltman, Ludo, 2016. "A review of the literature on citation impact indicators," Journal of Informetrics, Elsevier, vol. 10(2), pages 365-391.
    7. Fairclough, Ruth & Thelwall, Mike, 2015. "More precise methods for national research citation impact comparisons," Journal of Informetrics, Elsevier, vol. 9(4), pages 895-906.
    8. Thelwall, Mike, 2016. "The discretised lognormal and hooked power law distributions for complete citation data: Best options for modelling and regression," Journal of Informetrics, Elsevier, vol. 10(2), pages 336-346.
    9. Thelwall, Mike, 2017. "Three practical field normalised alternative indicator formulae for research evaluation," Journal of Informetrics, Elsevier, vol. 11(1), pages 128-151.
    10. Mike Thelwall, 2016. "Interpreting correlations between citation counts and other indicators," Scientometrics, Springer;Akadémiai Kiadó, vol. 108(1), pages 337-347, July.
    11. S. R. Goldberg & H. Anthony & T. S. Evans, 2015. "Modelling citation networks," Scientometrics, Springer;Akadémiai Kiadó, vol. 105(3), pages 1577-1604, December.
    12. Thelwall, Mike & Fairclough, Ruth, 2017. "The accuracy of confidence intervals for field normalised indicators," Journal of Informetrics, Elsevier, vol. 11(2), pages 530-540.
    13. Rodríguez-Navarro, Alonso & Brito, Ricardo, 2018. "Double rank analysis for research assessment," Journal of Informetrics, Elsevier, vol. 12(1), pages 31-41.
    14. Abramo, Giovanni & Cicero, Tindaro & D’Angelo, Ciriaco Andrea, 2012. "How important is choice of the scaling factor in standardizing citations?," Journal of Informetrics, Elsevier, vol. 6(4), pages 645-654.
    15. Thelwall, Mike, 2018. "Do females create higher impact research? Scopus citations and Mendeley readers for articles from five countries," Journal of Informetrics, Elsevier, vol. 12(4), pages 1031-1041.
    16. Brito, Ricardo & Navarro, Alonso Rodríguez, 2021. "The inconsistency of h-index: A mathematical analysis," Journal of Informetrics, Elsevier, vol. 15(1).
    17. Copiello, Sergio, 2019. "Peer and neighborhood effects: Citation analysis using a spatial autoregressive model and pseudo-spatial data," Journal of Informetrics, Elsevier, vol. 13(1), pages 238-254.
    18. Brito, Ricardo & Rodríguez-Navarro, Alonso, 2018. "Research assessment by percentile-based double rank analysis," Journal of Informetrics, Elsevier, vol. 12(1), pages 315-329.
    19. Alonso Rodríguez-Navarro & Ricardo Brito, 2019. "Probability and expected frequency of breakthroughs: basis and use of a robust method of research assessment," Scientometrics, Springer;Akadémiai Kiadó, vol. 119(1), pages 213-235, April.
    20. Bouyssou, Denis & Marchant, Thierry, 2016. "Ranking authors using fractional counting of citations: An axiomatic approach," Journal of Informetrics, Elsevier, vol. 10(1), pages 183-199.

    Corrections

    All material on this site has been provided by the respective publishers and authors. You can help correct errors and omissions. When requesting a correction, please mention this item's handle: RePEc:eee:infome:v:10:y:2016:i:2:p:454-470. See general information about how to correct material in RePEc.

    If you have authored this item and are not yet registered with RePEc, we encourage you to do it here. This allows to link your profile to this item. It also allows you to accept potential citations to this item that we are uncertain about.

    If CitEc recognized a bibliographic reference but did not link an item in RePEc to it, you can help with this form .

    If you know of missing items citing this one, you can help us creating those links by adding the relevant references in the same way as above, for each refering item. If you are a registered author of this item, you may also want to check the "citations" tab in your RePEc Author Service profile, as there may be some citations waiting for confirmation.

    For technical questions regarding this item, or to correct its authors, title, abstract, bibliographic or download information, contact: Catherine Liu (email available below). General contact details of provider: http://www.elsevier.com/locate/joi .

    Please note that corrections may take a couple of weeks to filter through the various RePEc services.

    IDEAS is a RePEc service. RePEc uses bibliographic data supplied by the respective publishers.