| Citation: | LIN Xiaochun, ZHOU Limin, ZHANG Baoke, YUAN Xin, ZHENG Xu, CHEN Xiaoran. Discrimination of Rare Earth Element Enrichment in Carbonatites Using Chained Preprocessing and a Self-Training Support Vector MachineJ. Rock and Mineral Analysis. DOI: 10.15898/j.ykcs.202605300149 |
In the global evaluation of carbonatite-related rare earth element (REE) resources, a large number of geochemical samples lack mineralization labels, while the data exhibit heterogeneous distributions and substantial dimensional differences between major and trace elements. Existing studies mainly rely on supervised learning with limited labeled samples, making it difficult to fully exploit massive unlabeled data and thereby constraining the identification of REE enrichment patterns and the assessment of mineralization potential. Based on a global whole-rock geochemical dataset comprising 1581 carbonatite samples (136 labeled samples and 1,445 unlabeled samples), this study developed a chained preprocessing workflow consisting of log1p transformation, K-nearest neighbor imputation, and RobustScaler normalization, and established a semi-supervised REE enrichment discrimination model using a self-training support vector machine (Self-training SVM). This approach effectively utilizes information from unlabeled samples under conditions of limited labeled data. The results show that the chained preprocessing effectively improves the skewed distribution and scale differences of geochemical data, thereby enhancing the comparability of multi-element features. Compared with the unprocessed data, the proposed model required only seven iterations instead of nine to obtain 1404 high-confidence pseudo-labels, significantly improving pseudo-label utilization. For the labeled samples, both five-fold stratified cross-validation and leave-one-out cross-validation achieved an accuracy of 91.2%, with areas under the receiver operating characteristic curve of 0.967 and 0.982, respectively, demonstrating good model stability and generalization ability. Principal component analysis indicates that REE-fertile and barren carbonatites are well separated in the high-dimensional geochemical feature space, revealing distinct multi-element geochemical assemblage characteristics between the two groups. The proposed method, integrating chained preprocessing with a self-training support vector machine, fully exploits the information contained in massive unlabeled geochemical data to achieve intelligent discrimination of REE enrichment status. It provides a new technical approach for evaluating the mineralization potential of carbonatite-related REE deposits and for the intelligent analysis of geological and mineral testing data.