A New Similarity Function for Spectral Clustering with Application to Plant Phenotypic Data (2312.14920v3)
Abstract: Clustering species of the same plant into different groups is an important step in developing new species of the concerned plant. Phenotypic (or physical) characteristics of plant species are commonly used to perform clustering. Hierarchical Clustering (HC) is popularly used for this task, and this algorithm suffers from low accuracy. In one of the recent works (Shastri et al., 2021), the authors have used the standard Spectral Clustering (SC) algorithm to improve the clustering accuracy. They have demonstrated the efficacy of their algorithm on soybean species. In the SC algorithm, one of the crucial steps is building the similarity matrix. A Gaussian similarity function is the standard choice to build this matrix. In the past, many works have proposed variants of the Gaussian similarity function to improve the performance of the SC algorithm, however, all have focused on the variance or scaling of the Gaussian. None of the past works have investigated upon the choice of base "e" (Euler's number) of the Gaussian similarity function (natural exponential function). Based upon spectral graph theory, specifically the Cheeger's inequality, in this work we propose use of a base "a" exponential function as the similarity function. We also integrate this new approach with the notion of "local scaling" from one of the first works that experimented with the scaling of the Gaussian similarity function (Zelnik-Manor et al., 2004). Using an eigenvalue analysis, we theoretically justify that our proposed algorithm should work better than the existing one. With evaluation on 2376 soybean species and 1865 rice species, we experimentally demonstrate that our new SC is 35% and 11% better than the standard SC, respectively.
- Unequal probability sampling without replacement through a splitting method. Biometrika 85, 89–101.
- Cheeger’s inequality and the sparse cut problem. Lecture notes on recent advances in approximation algorithms (University of Washington).
- Cheeger’s inequality continued, spectral clustering. Lecture notes on design and analysis of algorithms I (University of Washington).
- Diversity and population structure of red rice germplasm in Bangladesh. PLoS One 13, e0196096.
- Genetic variability and cluster analysis for phenological traits of thai indigenous upland rice (oryza sativa l.). Indian Journal of Agricultural Research 54.
- Cube sampled K-prototype clustering for featured data, in: 2021 IEEE 18th India Council International Conference (INDICON), pp. 1–6.
- Cluster analysis in common bean genotypes (Phaseolus Vulgaris L.). Turkish Journal of Agricultural and Natural Sciences 1, 1030–1035.
- Multiway spectral partitioning and higher-order cheeger inequalities. Journal of the ACM (JACM) 61, 1–30.
- The international rice information system. A platform for meta-analysis of rice crop data. Plant Physiology 139, 637–642.
- Determination of the optimal number of clusters using a spectral clustering optimization. Expert systems with applications 65, 304–314.
- On spectral clustering: Analysis and an algorithm, in: Advances in neural information processing systems, MIT Press. pp. 849–856.
- Molecular and morphological characterization of indian farmers rice varieties (oryza sativa l.). Australian Journal of Crop Science 7, 923.
- Clustering analysis of soybean germplasm (glycine max l. merrill). The Pharma Innovation Journal 7, 781–786.
- Assessing genetic variation for heat tolerance in synthetic wheat lines using phenotypic data and molecular markers. Australian Journal of Crop Science 8, 515–522.
- Probabilistically sampled and spectrally clustered plant species using phenotypic characteristics. PeerJ 9, e11927.
- Vector quantized spectral clustering applied to whole genome sequences of plants. Evolutionary Bioinformatics 15, 1–7.
- Genetic diversity is indispensable for plant breeding to improve crops. Crop Science 61, 839–852.
- Variability assessment for root and drought tolerance traits and genetic diversity analysis of rice germplasm using SSR markers. Scientific reports 9, 16513.
- A tutorial on spectral clustering. Statistics and computing 17, 395–416.
- Self-tuning spectral clustering., in: Advances in neural information processing systems, pp. 1601–1608.