Inferring gene regulatory networks (GRNs) from single-cell RNA sequencing (scRNA-seq) data is fundamentally challenged by severe data sparsity, where pervasive dropout events obscure true regulatory signals and compromise the reliability of downstream inference. Existing supervised methods, while leveraging prior network structures, remain highly susceptible to this noise due to their end-to-end learning paradigm. To address this bottleneck, we propose SGMHA, a novel two-stage framework that decouples representation learning from link prediction. Specifically, SGMHA first employs a self-supervised graph masked autoencoder (GraphMAE) to learn robust gene representations by reconstructing randomly masked expression values, thereby mitigating sparsity-induced distortions. Subsequently, an MHA (multi-head attention)-based fine-tuning module integrates these pre-trained representations with raw expression data to accurately infer directed regulatory links. Extensive benchmarking across seven scRNA-seq datasets demonstrates that SGMHA consistently outperforms eight state-of-the-art methods in both area under the receiver operating characteristic curve (AUROC) and area under the precision-recall curve (AUPRC). Applying SGMHA to breast cancer metastasis revealed context-specific GRNs and identified 26 high-confidence candidate drivers. Among these, six (NDUFAF4, ENY2, CCT5, PGK1, DCTPP1, and H2AFZ) were validated as prognostic biomarkers, with their mechanistic roles in metastatic adaptation detailed through multi-omics integration. Collectively, SGMHA provides an accurate, scalable, and biologically interpretable tool for GRN inference, holding strong promise for biomarker discovery in complex diseases.
山东省济南市章丘区文博路2号
齐鲁师范学院 genelibs生信实验室
山东省济南市高新区舜华路750号
大学科技园北区F座4单元2楼
电话: 0531-88819269