Abstract: Background: Cluster analysis of high dimensional data poses various challenges including problems associated with the curse of dimensionality, presence of irrelevant attributes, and high computational costs. Conventional GMM based techniques do not perform well with data having thousands of attributes and small number of observations, e.g., Coil20, GLIOMA, and Lung datasets.
Materials and.....
Key Word: Data clustering, Ensemble clustering, GMM, Adaptive subspace generation, Reliability-aware learning, Cross-partition consensus, Random sampling
[1].
A. Saxena, M. Prasad, A. Gupta, et al., “A review of clustering techniques and developments,” Neurocomputing, vol. 267, pp. 664–681, 2017.
[2].
M. Mittal, L. M. Goyal, D. J. Hemanth, et al., “Clustering approaches for high-dimensional databases: A review,” WIREs Data Mining and Knowledge Discovery, vol. 9, no. 3, article no. e1300, 2019.
[3].
G. T. Reddy, M. P. K. Reddy, K. Lakshmanna, et al., “Analysis of dimensionality reduction techniques on big data,” IEEE Access, vol. 8, pp. 54776–54788, 2020.[4].
D. Wang, X. Y. Guo, S. Li, et al., “Robust high dimensional expectation maximization algorithm via trimmed hard thresholding,” Machine Learning, vol. 109, no. 12, pp. 2283–2311, 2020.
[5].
J. L. Liu, D. Cai, and X. F. He, “Gaussian mixture model with local consistency,” in Proceedings of the 24th AAAI Conference on Artificial Intelligence, Atlanta, GA, USA, pp. 512–517, 2010.