Standardization of interval symbolic data based on the empirical descriptive statistics

Standardization of interval symbolic data based on the empirical descriptive statistics

Author

Listed:

Guo, Junpeng
Li, Wenhua
Li, Chenhua
Gao, Sa

Abstract

In many statistical analysis methods, standardization of the sample data is usually recommended to prevent the results from being strongly affected by the scale of measurement of the variables. This paper focuses on the standardization of interval data obtained by symbolic data analysis (SDA). SDA is a new data analysis technique which captures the value of a variable with a symbolic representation. The empirical descriptive statistics of the interval symbolic variable are studied first. We then proposed the standardization method of interval symbolic data and conducted a simulation study to evaluate our standardization method by using clustering analysis. An application example on e-shops of several major cities in China is given at the end of the paper. Differing from previous research, we do not require the assumption of uniformly distributed data in the interval. Our method makes the best use of the original sample information.

Suggested Citation

Guo, Junpeng & Li, Wenhua & Li, Chenhua & Gao, Sa, 2012. "Standardization of interval symbolic data based on the empirical descriptive statistics," Computational Statistics & Data Analysis, Elsevier, vol. 56(3), pages 602-610.

Handle: RePEc:eee:csdana:v:56:y:2012:i:3:p:602-610
DOI: 10.1016/j.csda.2011.09.006

Download full text from publisher

As the access to this document is restricted, you may want to

for a different version of it.

References listed on IDEAS

Lawrence Hubert & Phipps Arabie, 1985. "Comparing partitions," Journal of Classification, Springer;The Classification Society, vol. 2(1), pages 193-218, December.
Billard L. & Diday E., 2003. "From the Statistics of Data to the Statistics of Knowledge: Symbolic Data Analysis," Journal of the American Statistical Association, American Statistical Association, vol. 98, pages 470-487, January.
Francisco Carvalho & Paula Brito & Hans-Hermann Bock, 2006. "Dynamic clustering for interval data based on L 2 distance," Computational Statistics, Springer, vol. 21(2), pages 231-250, June.

Full references (including those not matched with items on IDEAS)

Citations

Citations are extracted by the CitEc Project, subscribe to its RSS feed for this item.

Cited by:

Wenyang Huang & Huiwen Wang & Shanshan Wang, 2024. "A structural VAR and VECM modeling method for open-high-low-close data contained in candlestick chart," Financial Innovation, Springer;Southwestern University of Finance and Economics, vol. 10(1), pages 1-29, December.
Wenhua Li & Junpeng Guo & Ying Chen & Minglu Wang, 2016. "A New Representation of Interval Symbolic Data and Its Application in Dynamic Clustering," Journal of Classification, Springer;The Classification Society, vol. 33(1), pages 149-165, April.

Most related items

These are the items that most often cite the same works as this one and are cited by the same works as this one.

A. Pedro Duarte Silva & Peter Filzmoser & Paula Brito, 2018. "Outlier detection in interval data," Advances in Data Analysis and Classification, Springer;German Classification Society - Gesellschaft für Klassifikation (GfKl);Japanese Classification Society (JCS);Classification and Data Analysis Group of the Italian Statistical Society (CLADAG);International Federation of Classification Societies (IFCS), vol. 12(3), pages 785-822, September.
Fei Liu & L. Billard, 2022. "Partition of Interval-Valued Observations Using Regression," Journal of Classification, Springer;The Classification Society, vol. 39(1), pages 55-77, March.
Maia, André Luis Santiago & de Carvalho, Francisco de A.T., 2011. "Holt's exponential smoothing and neural network models for forecasting interval-valued time series," International Journal of Forecasting, Elsevier, vol. 27(3), pages 740-759, July.
Nataša Kejžar & Simona Korenjak-Černe & Vladimir Batagelj, 2021. "Clustering of modal-valued symbolic data," Advances in Data Analysis and Classification, Springer;German Classification Society - Gesellschaft für Klassifikation (GfKl);Japanese Classification Society (JCS);Classification and Data Analysis Group of the Italian Statistical Society (CLADAG);International Federation of Classification Societies (IFCS), vol. 15(2), pages 513-541, June.
Maia, André Luis Santiago & de Carvalho, Francisco de A.T., 2011. "Holt’s exponential smoothing and neural network models for forecasting interval-valued time series," International Journal of Forecasting, Elsevier, vol. 27(3), pages 740-759.
M. Rosário Oliveira & Margarida Azeitona & António Pacheco & Rui Valadas, 2022. "Association measures for interval variables," Advances in Data Analysis and Classification, Springer;German Classification Society - Gesellschaft für Klassifikation (GfKl);Japanese Classification Society (JCS);Classification and Data Analysis Group of the Italian Statistical Society (CLADAG);International Federation of Classification Societies (IFCS), vol. 16(3), pages 491-520, September.
Ana Belén Ramos-Guajardo, 2022. "A hierarchical clustering method for random intervals based on a similarity measure," Computational Statistics, Springer, vol. 37(1), pages 229-261, March.
Mihaela Voinea, 2021. "Support Teacher as Key Factor of Integration Children with Special Education Needs in Mainstream School," European Journal of Social Sciences Education and Research Articles, Revistia Research and Publishing, vol. 8, September.
Salvatore D. Tomarchio & Michael P. B. Gallaugher, 2026. "Mixtures of regressions using matrix-variate heavy-tailed distributions," Advances in Data Analysis and Classification, Springer;German Classification Society - Gesellschaft für Klassifikation (GfKl);Japanese Classification Society (JCS);Classification and Data Analysis Group of the Italian Statistical Society (CLADAG);International Federation of Classification Societies (IFCS), vol. 20(1), pages 11-38, March.
Yunpeng Zhao & Qing Pan & Chengan Du, 2019. "Logistic regression augmented community detection for network data with application in identifying autism‐related gene pathways," Biometrics, The International Biometric Society, vol. 75(1), pages 222-234, March.
Miguel de Carvalho & Gabriel Martos, 2022. "Modeling interval trendlines: Symbolic singular spectrum analysis for interval time series," Journal of Forecasting, John Wiley & Sons, Ltd., vol. 41(1), pages 167-180, January.
Wu, Han-Ming & Tien, Yin-Jing & Chen, Chun-houh, 2010. "GAP: A graphical environment for matrix visualization and cluster analysis," Computational Statistics & Data Analysis, Elsevier, vol. 54(3), pages 767-778, March.
José E. Chacón, 2021. "Explicit Agreement Extremes for a 2 × 2 Table with Given Marginals," Journal of Classification, Springer;The Classification Society, vol. 38(2), pages 257-263, July.
F. Marta L. Di Lascio & Andrea Menapace & Roberta Pappadà, 2024. "A spatially‐weighted AMH copula‐based dissimilarity measure for clustering variables: An application to urban thermal efficiency," Environmetrics, John Wiley & Sons, Ltd., vol. 35(1), February.
- F. Marta L. Di Lascio & Andrea Menapace & Roberta Pappadà, 2021. "A spatially-weighted AMH copula-based dissimilarity measure for clustering variables: An application to urban thermal efficiency," BEMPS - Bozen Economics & Management Paper Series BEMPS89, Faculty of Economics and Management at the Free University of Bozen.
Yifan Zhu & Chongzhi Di & Ying Qing Chen, 2019. "Clustering Functional Data with Application to Electronic Medication Adherence Monitoring in HIV Prevention Trials," Statistics in Biosciences, Springer;International Chinese Statistical Association, vol. 11(2), pages 238-261, July.
Zhen-Hao Guo & De-Shuang Huang & Shihua Zhang, 2026. "Multi-species integration, alignment and annotation of single-cell RNA-seq data with CAMEX," Nature Communications, Nature, vol. 17(1), pages 1-17, December.
Irene Vrbik & Paul McNicholas, 2015. "Fractionally-Supervised Classification," Journal of Classification, Springer;The Classification Society, vol. 32(3), pages 359-381, October.
Maurizio Vichi & Carlo Cavicchia & Patrick J. F. Groenen, 2022. "Hierarchical Means Clustering," Journal of Classification, Springer;The Classification Society, vol. 39(3), pages 553-577, November.
Batool, Fatima & Hennig, Christian, 2021. "Clustering with the Average Silhouette Width," Computational Statistics & Data Analysis, Elsevier, vol. 158(C).
Patrick D. Shay & Stephen S. Farnsworth Mick, 2017. "Clustered and distinct: a taxonomy of local multihospital systems," Health Care Management Science, Springer, vol. 20(3), pages 303-315, September.

More about this item

Keywords

; ; ; ;

Statistics

Access and download statistics

Corrections

All material on this site has been provided by the respective publishers and authors. You can help correct errors and omissions. When requesting a correction, please mention this item's handle: RePEc:eee:csdana:v:56:y:2012:i:3:p:602-610. See general information about how to correct material in RePEc.

If you have authored this item and are not yet registered with RePEc, we encourage you to do it here. This allows to link your profile to this item. It also allows you to accept potential citations to this item that we are uncertain about.

If CitEc recognized a bibliographic reference but did not link an item in RePEc to it, you can help with this form .

If you know of missing items citing this one, you can help us creating those links by adding the relevant references in the same way as above, for each refering item. If you are a registered author of this item, you may also want to check the "citations" tab in your RePEc Author Service profile, as there may be some citations waiting for confirmation.

For technical questions regarding this item, or to correct its authors, title, abstract, bibliographic or download information, contact: Catherine Liu (email available below). General contact details of provider: http://www.elsevier.com/locate/csda .

Please note that corrections may take a couple of weeks to filter through the various RePEc services.

IDEAS is a RePEc service. RePEc uses bibliographic data supplied by the respective publishers.

Browse Econ Literature

More features

Standardization of interval symbolic data based on the empirical descriptive statistics

Author

Abstract

Suggested Citation

Download full text from publisher

References listed on IDEAS

Citations

Most related items

More about this item

Keywords

Statistics

Corrections

More services and features

MyIDEAS

Author registration

Rankings

RePEc Genealogy

RePEc Biblio

MPRA

New papers by email

EconAcademics

Plagiarism

About RePEc

RePEc home

Blog

Help/FAQ

RePEc team

Participating archives

Privacy statement

Help us

Corrections

Volunteers

Get papers listed

Open a RePEc archive

Get RePEc data