RNA sequencing datasets in the gene expression omnibus (GEO) increasingly include NCBI-generated count matrices, enabling streamlined signature gene discovery. We present ERAPID, a computational framework that automatically processes this public data for robust differential expression analysis. The pipeline integrates metadata harmonization, surrogate variable analysis (SVA) to capture latent technical variation for batch correction, dual differential expression methods (DESeq2 and dream), gene set enrichment analysis (GSEA), and evidence-based gene prioritization via automated literature mining. ERAPID delivers interactive dashboards (volcano/MA plots, heatmaps, enrichment reports, searchable DEG tables) and supports an optional meta‑analysis step. Applied to a neuropsychiatric cohort (GSE80655), ERAPID completed analysis on a standard laptop in under an hour, recapitulating the reported association of EGR1 with schizophrenia. In an Alzheimer's disease (AD) case study integrating four GEO datasets, ERAPID identified 17 DEGs consistently altered across all AD‑versus‑control comparisons (e.g., ADCYAP1, PPEF1, VGF, and CRH), with KEGG Alzheimer's disease and oxidative phosphorylation pathways showing negative enrichment. Thus, ERAPID lowers the barrier to reusing public transcriptomes for signature gene discovery and biological interpretation.
山东省济南市章丘区文博路2号
齐鲁师范学院 genelibs生信实验室
山东省济南市高新区舜华路750号
大学科技园北区F座4单元2楼
电话: 0531-88819269