基于自训练支持向量机的碳酸岩稀土富集判别方法研究

Discrimination of Rare Earth Element Enrichment in Carbonatites Using Chained Preprocessing and a Self-Training Support Vector Machine

  • 摘要: 针对全球碳酸岩型稀土资源评价中大量地球化学样品缺乏成矿标签、数据分布异质且主微量元素量纲差异显著等问题,现有研究多依赖小样本监督学习,难以充分利用海量无标签数据,制约了稀土富集规律的揭示与成矿潜力评价。本文基于全球1581件碳酸岩全岩地球化学数据(136件有标签样品,1445件无标签样品),构建“log1p变换-K近邻插值-RobustScaler标准化”链式预处理流程,并结合自训练支持向量机建立了稀土富集状态半监督判别模型,在标签稀缺条件下实现了无标签样品信息的有效利用。结果表明:链式预处理有效改善了地球化学数据的偏态分布与尺度差异,增强了多元素特征的可比性;与未作预处理数据相比,模型由9轮迭代缩短至7轮迭代即可获得1404个高置信度伪标签,显著提高了伪标签利用率;标注样品经5折分层交叉验证和留一法验证的准确率均为91.2%,ROC曲线下面积分别达到0.967和0.982,表明模型具有良好的稳定性与泛化能力。主成分分析表明稀土富集与贫瘠碳酸岩在高维地球化学特征空间中具有良好的可分性,揭示了二者存在显著差异的多元素地球化学组合特征。本文提出的链式预处理与自训练支持向量机方法,能够充分挖掘大量无标签地球化学数据的信息,实现稀土富集状态的智能判别,为碳酸岩型稀土矿化潜力评价和岩矿测试数据智能分析提供了新的技术方法。

     

    Abstract: In the global evaluation of carbonatite-related rare earth element (REE) resources, a large number of geochemical samples lack mineralization labels, while the data exhibit heterogeneous distributions and substantial dimensional differences between major and trace elements. Existing studies mainly rely on supervised learning with limited labeled samples, making it difficult to fully exploit massive unlabeled data and thereby constraining the identification of REE enrichment patterns and the assessment of mineralization potential. Based on a global whole-rock geochemical dataset comprising 1581 carbonatite samples (136 labeled samples and 1,445 unlabeled samples), this study developed a chained preprocessing workflow consisting of log1p transformation, K-nearest neighbor imputation, and RobustScaler normalization, and established a semi-supervised REE enrichment discrimination model using a self-training support vector machine (Self-training SVM). This approach effectively utilizes information from unlabeled samples under conditions of limited labeled data. The results show that the chained preprocessing effectively improves the skewed distribution and scale differences of geochemical data, thereby enhancing the comparability of multi-element features. Compared with the unprocessed data, the proposed model required only seven iterations instead of nine to obtain 1404 high-confidence pseudo-labels, significantly improving pseudo-label utilization. For the labeled samples, both five-fold stratified cross-validation and leave-one-out cross-validation achieved an accuracy of 91.2%, with areas under the receiver operating characteristic curve of 0.967 and 0.982, respectively, demonstrating good model stability and generalization ability. Principal component analysis indicates that REE-fertile and barren carbonatites are well separated in the high-dimensional geochemical feature space, revealing distinct multi-element geochemical assemblage characteristics between the two groups. The proposed method, integrating chained preprocessing with a self-training support vector machine, fully exploits the information contained in massive unlabeled geochemical data to achieve intelligent discrimination of REE enrichment status. It provides a new technical approach for evaluating the mineralization potential of carbonatite-related REE deposits and for the intelligent analysis of geological and mineral testing data.

     

/

返回文章
返回