Author
Listed:
- Eric Marcon
(UMR AMAP - Botanique et Modélisation de l'Architecture des Plantes et des Végétations - Cirad - Centre de Coopération Internationale en Recherche Agronomique pour le Développement - CNRS - Centre National de la Recherche Scientifique - IRD [Occitanie] - Institut de Recherche pour le Développement - délégation Occitanie - IRD - Institut de Recherche pour le Développement - INRAE - Institut National de Recherche pour l’Agriculture, l’Alimentation et l’Environnement - UM - Université de Montpellier, AgroParisTech)
- Florence Puech
(UMR PSAE - Paris-Saclay Applied Economics - AgroParisTech - Université Paris-Saclay - INRAE - Institut National de Recherche pour l’Agriculture, l’Alimentation et l’Environnement)
Abstract
Since agglomeration is a core question in regional science, spatial concentration measures are widely employed in that field to evaluate the spatial distribution of activities. Increasing access to large geo-referenced datasets, coupled with the development of computing power, has encouraged the search for suitable spatial statistical tools. Distance-based methods have been extensively developed to detect spatial concentration, dispersion or independence of entities at any distance and without any bias. Recently, Tidu et al. (2024) highlighted the qualities of Marcon and Puech's M function, a relative distance-based measure, and also expressed reservations about the computation time required. Herein, we explore two possible ways to reduce the computation burden of large geo-located datasets: approximating the position of points and thinning the point pattern. In both cases, the deterioration extent of the M results is estimated and discussed as the gains it provides in computation time, using the R software. We also discuss implications of these findings in the field of regional science. We notably provide evidence that the individual location approximation generates information loss at substantially small distances, implying a trade-off between the smallest distance at which spatial interactions could be detected and computing performance. We also give support that thinning is an efficient method to analyze large datasets with very good accuracy. The R code used in the article is given for the reproducibility of our results.
Suggested Citation
Eric Marcon & Florence Puech, 2026.
"Computation of Large Spatial Datasets with the M function,"
Working Papers
hal-05512154, HAL.
Handle:
RePEc:hal:wpaper:hal-05512154
Note: View the original document on HAL open archive server: https://hal.science/hal-05512154v2
Download full text from publisher
Corrections
All material on this site has been provided by the respective publishers and authors. You can help correct errors and omissions. When requesting a correction, please mention this item's handle: RePEc:hal:wpaper:hal-05512154. See general information about how to correct material in RePEc.
If you have authored this item and are not yet registered with RePEc, we encourage you to do it here. This allows to link your profile to this item. It also allows you to accept potential citations to this item that we are uncertain about.
We have no bibliographic references for this item. You can help adding them by using this form .
If you know of missing items citing this one, you can help us creating those links by adding the relevant references in the same way as above, for each refering item. If you are a registered author of this item, you may also want to check the "citations" tab in your RePEc Author Service profile, as there may be some citations waiting for confirmation.
For technical questions regarding this item, or to correct its authors, title, abstract, bibliographic or download information, contact: CCSD (email available below). General contact details of provider: https://hal.archives-ouvertes.fr/ .
Please note that corrections may take a couple of weeks to filter through
the various RePEc services.