Home LiteratureArticle Details
PMID: 31221194 Published · epublish English Journal Article Research Support, Non-U.S. Gov't Review

Essential guidelines for computational method benchmarking.

Genome biology ·Vol. 20 ·No. 1 ·2019-00-20 ·Pages 125

Weber LM, Saelens W, Cannoodt R, Soneson C, Hapfelmeier A, Gardner PP, Boulesteix AL, Saeys Y, Robinson MD

Abstract

In computational biology and other sciences, researchers are frequently faced with a choice between several computational methods for performing data analyses. Benchmarking studies aim to rigorously compare the performance of different methods using well-characterized benchmark datasets, to determine the strengths of each method or to provide recommendations regarding suitable choices of methods for an analysis. However, benchmarking studies must be carefully designed and implemented to provide accurate, unbiased, and informative results. Here, we summarize key practical guidelines and recommendations for performing high-quality benchmarking analyses, based on our experiences in computational biology.

MeSH Terms
Benchmarking Computational Biology/standards Datasets as Topic Guidelines as Topic Publishing Research Design Software
Authors & Affiliations
9 authors, click to expand affiliations / ORCID
Weber Lukas M
Institute of Molecular Life Sciences, University of Zurich, 8057, Zurich, Switzerland. | SIB Swiss Institute of Bioinformatics, University of Zurich, 8057, Zurich, Switzerland.
Saelens Wouter
Data Mining and Modelling for Biomedicine, VIB Center for Inflammation Research, 9052, Ghent, Belgium. | Department of Applied Mathematics, Computer Science and Statistics, Ghent University, 9000, Ghent, Belgium.
Cannoodt Robrecht
Data Mining and Modelling for Biomedicine, VIB Center for Inflammation Research, 9052, Ghent, Belgium. | Department of Applied Mathematics, Computer Science and Statistics, Ghent University, 9000, Ghent, Belgium.
Soneson Charlotte
Institute of Molecular Life Sciences, University of Zurich, 8057, Zurich, Switzerland. | SIB Swiss Institute of Bioinformatics, University of Zurich, 8057, Zurich, Switzerland. | Present address: Friedrich Miescher Institute for Biomedical Research and SIB Swiss Institute of Bioinformatics, 4058, Basel, Switzerland.
Hapfelmeier Alexander
Institute of Medical Informatics, Statistics and Epidemiology, Technical University of Munich, 81675, Munich, Germany.
Gardner Paul P
Department of Biochemistry, University of Otago, Dunedin, 9016, New Zealand.
Boulesteix Anne-Laure
Institute for Medical Information Processing, Biometry and Epidemiology, Ludwig-Maximilians-University, 81377, Munich, Germany.
Saeys Yvan
Data Mining and Modelling for Biomedicine, VIB Center for Inflammation Research, 9052, Ghent, Belgium. [email protected]. | Department of Applied Mathematics, Computer Science and Statistics, Ghent University, 9000, Ghent, Belgium. [email protected].
Robinson Mark D ORCID
Institute of Molecular Life Sciences, University of Zurich, 8057, Zurich, Switzerland. [email protected]. | SIB Swiss Institute of Bioinformatics, University of Zurich, 8057, Zurich, Switzerland. [email protected].
References (97)
97 references, click to expand
  1. Issues in bioinformatics benchmarking: the case study of multiple sequence alignment.
    Nucleic Acids Res. 2010 Nov;38(21):7353-63 PMID: 20639539
  2. Highly parallel direct RNA sequencing on an array of nanopores.
    Nat Methods. 2018 Mar;15(3):201-206 PMID: 29334379
  3. Crowdsourced analysis of clinical trial data to predict amyotrophic lateral sclerosis progression.
    Nat Biotechnol. 2015 Jan;33(1):51-7 PMID: 25362243
  4. A comparison of single-cell trajectory inference methods.
    Nat Biotechnol. 2019 May;37(5):547-554 PMID: 30936559
  5. Ten simple rules for a community computational challenge.
    PLoS Comput Biol. 2015 Apr 23;11(4):e1004150 PMID: 25906249
  6. Systematic benchmarking of omics computational tools.
    Nat Commun. 2019 Mar 27;10(1):1393 PMID: 30918265
  7. Parameter tuning is a key part of dimensionality reduction via deep variational autoencoders for single cell RNA transcriptomics.
    Pac Symp Biocomput. 2019;24:362-373 PMID: 30963075
  8. Detection and accurate false discovery rate control of differentially methylated regions from whole genome bisulfite sequencing.
    Biostatistics. 2019 Jul 1;20(3):367-383 PMID: 29481604
  9. The evaluation of tools used to predict the impact of missense variants is hindered by two types of circularity.
    Hum Mutat. 2015 May;36(5):513-23 PMID: 25684150
  10. Meta-analysis methods for genome-wide association studies and beyond.
    Nat Rev Genet. 2013 Jun;14(6):379-89 PMID: 23657481
  11. Comparison of Affymetrix GeneChip expression measures.
    Bioinformatics. 2006 Apr 1;22(7):789-94 PMID: 16410320
  12. Towards evidence-based computational statistics: lessons from clinical research on the role and design of real-data benchmark studies.
    BMC Med Res Methodol. 2017 Sep 9;17(1):138 PMID: 28888225
  13. Do count-based differential expression methods perform poorly when genes are expressed in only one condition?
    Genome Biol. 2015 Oct 08;16:222 PMID: 26450178
  14. The challenges of designing a benchmark strategy for bioinformatics pipelines in the identification of antimicrobial resistance determinants using next generation sequencing technologies.
    F1000Res. 2018 Apr 13;7:459 PMID: 30026930
  15. Sensitive detection of rare disease-associated cell subsets via representation learning.
    Nat Commun. 2017 Apr 06;8:14825 PMID: 28382969
  16. DataPackageR: Reproducible data preprocessing, standardization and sharing using R/Bioconductor for collaborative data analysis.
    Gates Open Res. 2018 Jul 10;2:31 PMID: 30234197
  17. Critical Assessment of Metagenome Interpretation-a benchmark of metagenomics software.
    Nat Methods. 2017 Nov;14(11):1063-1071 PMID: 28967888
  18. Why most published research findings are false.
    PLoS Med. 2005 Aug;2(8):e124 PMID: 16060722
  19. Comparison of mapping algorithms used in high-throughput sequencing: application to Ion Torrent data.
    BMC Genomics. 2014 Apr 05;15:264 PMID: 24708189
  20. An evaluation of the accuracy and speed of metagenome analysis tools.
    Sci Rep. 2016 Jan 18;6:19233 PMID: 26778510
  21. A Web Resource for Standardized Benchmark Datasets, Metrics, and Rosetta Protocols for Macromolecular Modeling and Design.
    PLoS One. 2015 Sep 03;10(9):e0130433 PMID: 26335248
  22. The Galaxy platform for accessible, reproducible and collaborative biomedical analyses: 2018 update.
    Nucleic Acids Res. 2018 Jul 2;46(W1):W537-W544 PMID: 29790989
  23. A plea for neutral comparison studies in computational sciences.
    PLoS One. 2013 Apr 24;8(4):e61562 PMID: 23637855
  24. iCOBRA: open, reproducible, standardized and live method benchmarking.
    Nat Methods. 2016 Apr;13(4):283 PMID: 27027585
  25. A comprehensive evaluation of module detection methods for gene expression data.
    Nat Commun. 2018 Mar 15;9(1):1090 PMID: 29545622
  26. Towards unified quality verification of synthetic count data with countsimQC.
    Bioinformatics. 2018 Feb 15;34(4):691-692 PMID: 29028961
  27. Exploring genomic dark matter: a critical assessment of the performance of homology search methods on noncoding RNA.
    Genome Res. 2007 Jan;17(1):117-25 PMID: 17151342
  28. Quantitative evaluation of software packages for single-molecule localization microscopy.
    Nat Methods. 2015 Aug;12(8):717-24 PMID: 26076424
  29. Exploring the single-cell RNA-seq analysis landscape with the scRNA-tools database.
    PLoS Comput Biol. 2018 Jun 25;14(6):e1006245 PMID: 29939984
  30. Snakemake--a scalable bioinformatics workflow engine.
    Bioinformatics. 2012 Oct 1;28(19):2520-2 PMID: 22908215
  31. Identifying accurate metagenome and amplicon software via a meta-analysis of sequence to taxonomy benchmarking studies.
    PeerJ. 2019 Jan 4;7:e6160 PMID: 30631651
  32. Robustly detecting differential expression in RNA sequencing data using observation weights.
    Nucleic Acids Res. 2014 Jun;42(11):e91 PMID: 24753412
  33. Comparing de novo genome assembly: the long and short of it.
    PLoS One. 2011 Apr 29;6(4):e19175 PMID: 21559467
  34. A comprehensive assessment of RNA-seq accuracy, reproducibility and information content by the Sequencing Quality Control Consortium.
    Nat Biotechnol. 2014 Sep;32(9):903-14 PMID: 25150838
  35. Comparison of clustering methods for high-dimensional single-cell flow and mass cytometry data.
    Cytometry A. 2016 Dec;89(12):1084-1096 PMID: 27992111
  36. miRNA-Seq normalization comparisons need improvement.
    RNA. 2013 Jun;19(6):733-4 PMID: 23616640
  37. ArrayExpress update--simplifying data submissions.
    Nucleic Acids Res. 2015 Jan;43(Database issue):D1113-6 PMID: 25361974
  38. Putting benchmarks in their rightful place: The heart of computational biology.
    PLoS Comput Biol. 2018 Nov 8;14(11):e1006494 PMID: 30408027
  39. Critical assessment of automated flow cytometry data analysis techniques.
    Nat Methods. 2013 Mar;10(3):228-38 PMID: 23396282
  40. A practical guide to methods controlling false discoveries in computational biology.
    Genome Biol. 2019 Jun 4;20(1):118 PMID: 31164141
  41. A comparison of meta-analysis methods for detecting differentially expressed genes in microarray experiments.
    Bioinformatics. 2008 Feb 1;24(3):374-82 PMID: 18204063
  42. Reproducible research in statistics: A review and guidelines for the Biometrical Journal.
    Biom J. 2016 Mar;58(2):416-27 PMID: 26711717
  43. Reproducible and replicable comparisons using SummarizedBenchmark.
    Bioinformatics. 2019 Jan 1;35(1):137-139 PMID: 30016409
  44. Reproducible research in computational science.
    Science. 2011 Dec 2;334(6060):1226-7 PMID: 22144613
  45. diffcyt: Differential discovery in high-dimensional cytometry via high-resolution clustering.
    Commun Biol. 2019 May 14;2:183 PMID: 31098416
  46. Inferring causal molecular networks: empirical assessment through a community-based effort.
    Nat Methods. 2016 Apr;13(4):310-8 PMID: 26901648
  47. NCBI GEO: archive for functional genomics data sets--update.
    Nucleic Acids Res. 2013 Jan;41(Database issue):D991-5 PMID: 23193258
  48. Bias, robustness and scalability in single-cell differential expression analysis.
    Nat Methods. 2018 Apr;15(4):255-261 PMID: 29481549
  49. A benchmark for evaluation of algorithms for identification of cellular correlates of clinical outcomes.
    Cytometry A. 2016 Jan;89(1):16-21 PMID: 26447924
  50. Differential analyses for RNA-seq: transcript-level estimates improve gene-level inferences.
    F1000Res. 2015 Dec 30;4:1521 PMID: 26925227
  51. Data-Driven Phenotypic Dissection of AML Reveals Progenitor-like Cells that Correlate with Prognosis.
    Cell. 2015 Jul 2;162(1):184-97 PMID: 26095251
  52. Making complex prediction rules applicable for readers: Current practice in random forest literature and recommendations.
    Biom J. 2019 Sep;61(5):1314-1328 PMID: 30069934
  53. Assemblathon 2: evaluating de novo methods of genome assembly in three vertebrate species.
    Gigascience. 2013 Jul 22;2(1):10 PMID: 23870653
  54. Assemblathon 1: a competitive assessment of de novo short read assembly methods.
    Genome Res. 2011 Dec;21(12):2224-41 PMID: 21926179
  55. A benchmark for Affymetrix GeneChip expression measures.
    Bioinformatics. 2004 Feb 12;20(3):323-31 PMID: 14960458
  56. The MicroArray Quality Control (MAQC)-II study of common practices for the development and validation of microarray-based predictive models.
    Nat Biotechnol. 2010 Aug;28(8):827-38 PMID: 20676074
  57. Best practices for benchmarking germline small-variant calls in human genomes.
    Nat Biotechnol. 2019 May;37(5):555-560 PMID: 30858580
  58. Benchmarking: contexts and details matter.
    Genome Biol. 2017 Jul 5;18(1):129 PMID: 28679434
  59. On the necessity and design of studies comparing statistical methods.
    Biom J. 2018 Jan;60(1):216-218 PMID: 29193206
  60. voom: Precision weights unlock linear model analysis tools for RNA-seq read counts.
    Genome Biol. 2014 Feb 03;15(2):R29 PMID: 24485249
  61. The BRaliBase dent-a tale of benchmark design and interpretation.
    Brief Bioinform. 2017 Mar 1;18(2):306-311 PMID: 26984616
  62. A comparison of methods for differential expression analysis of RNA-seq data.
    BMC Bioinformatics. 2013 Mar 09;14:91 PMID: 23497356
  63. Evaluation of methods for modeling transcription factor sequence specificity.
    Nat Biotechnol. 2013 Feb;31(2):126-34 PMID: 23354101
  64. Toward better benchmarking: challenge-based methods assessment in cancer genomics.
    Genome Biol. 2014 Sep 17;15(9):462 PMID: 25314947
  65. The self-assessment trap: can we all be better than average?
    Mol Syst Biol. 2011 Oct 11;7:537 PMID: 21988833
  66. Bioconda: sustainable and comprehensive software distribution for the life sciences.
    Nat Methods. 2018 Jul;15(7):475-476 PMID: 29967506
  67. QUAST: quality assessment tool for genome assemblies.
    Bioinformatics. 2013 Apr 15;29(8):1072-5 PMID: 23422339
  68. Comparison of clustering tools in R for medium-sized 10x Genomics single-cell RNA-sequencing data.
    F1000Res. 2018 Aug 15;7:1297 PMID: 30228881
  69. DRIMSeq: a Dirichlet-multinomial framework for multivariate count outcomes in genomics.
    F1000Res. 2016 Jun 13;5:1356 PMID: 28105305
  70. Synthetic spike-in standards for RNA-seq experiments.
    Genome Res. 2011 Sep;21(9):1543-51 PMID: 21816910
  71. Meta-research: Why research on research matters.
    PLoS Biol. 2018 Mar 13;16(3):e2005468 PMID: 29534060
  72. Benchmarking single cell RNA-sequencing analysis pipelines using mixture control experiments.
    Nat Methods. 2019 Jun;16(6):479-487 PMID: 31133762
  73. Combining tumor genome simulation with crowdsourcing to benchmark somatic single-nucleotide-variant detection.
    Nat Methods. 2015 Jul;12(7):623-30 PMID: 25984700
  74. Critical assessment of methods of protein structure prediction (CASP)-Round XII.
    Proteins. 2018 Mar;86 Suppl 1:7-15 PMID: 29082672
  75. A systematic performance evaluation of clustering methods for single-cell RNA-seq data.
    F1000Res. 2018 Jul 26;7:1141 PMID: 30271584
  76. Ten simple rules for reproducible computational research.
    PLoS Comput Biol. 2013 Oct;9(10):e1003285 PMID: 24204232
  77. Mortality Risk for Acute Cholangitis (MAC): a risk prediction model for in-hospital mortality in patients with acute cholangitis.
    BMC Gastroenterol. 2016 Feb 09;16:15 PMID: 26860903
  78. Comprehensive evaluation of differential gene expression analysis methods for RNA-seq data.
    Genome Biol. 2013;14(9):R95 PMID: 24020486
  79. mockrobiota: a Public Resource for Microbiome Bioinformatics Benchmarking.
    mSystems. 2016 Oct 18;1(5): PMID: 27822553
  80. Over-optimism in bioinformatics: an illustration.
    Bioinformatics. 2010 Aug 15;26(16):1990-8 PMID: 20581402
  81. Random forest versus logistic regression: a large-scale benchmark experiment.
    BMC Bioinformatics. 2018 Jul 17;19(1):270 PMID: 30016950
  82. Critical assessment of methods of protein structure prediction: Progress and new directions in round XI.
    Proteins. 2016 Sep;84 Suppl 1:4-14 PMID: 27171127
  83. Synthetic data sets for the identification of key ingredients for RNA-seq differential analysis.
    Brief Bioinform. 2018 Jan 1;19(1):65-76 PMID: 27742662
  84. Single-cell transcriptomics of 20 mouse organs creates a Tabula Muris.
    Nature. 2018 Oct;562(7727):367-372 PMID: 30283141
  85. The precision-recall plot is more informative than the ROC plot when evaluating binary classifiers on imbalanced datasets.
    PLoS One. 2015 Mar 04;10(3):e0118432 PMID: 25738806
  86. Comparing the performance of biomedical clustering methods.
    Nat Methods. 2015 Nov;12(11):1033-8 PMID: 26389570
  87. FlowRepository: a resource of annotated flow cytometry datasets associated with peer-reviewed publications.
    Cytometry A. 2012 Sep;81(9):727-31 PMID: 22887982
  88. Using simulation studies to evaluate statistical methods.
    Stat Med. 2019 May 20;38(11):2074-2102 PMID: 30652356
  89. Massively parallel digital transcriptional profiling of single cells.
    Nat Commun. 2017 Jan 16;8:14049 PMID: 28091601
  90. Simulation-based comprehensive benchmarking of RNA-seq aligners.
    Nat Methods. 2017 Feb;14(2):135-139 PMID: 27941783
  91. Ten simple rules for reducing overoptimistic reporting in methodological computational research.
    PLoS Comput Biol. 2015 Apr 23;11(4):e1004191 PMID: 25905639
  92. Improving the usability and archival stability of bioinformatics software.
    Genome Biol. 2019 Feb 27;20(1):47 PMID: 30813962
  93. Genomic landscape of human allele-specific DNA methylation.
    Proc Natl Acad Sci U S A. 2012 May 8;109(19):7332-7 PMID: 22523239
  94. A community effort to assess and improve drug sensitivity prediction algorithms.
    Nat Biotechnol. 2014 Dec;32(12):1202-12 PMID: 24880487
  95. A comprehensive evaluation of normalization methods for Illumina high-throughput RNA sequencing data analysis.
    Brief Bioinform. 2013 Nov;14(6):671-83 PMID: 22988256
  96. The MicroArray Quality Control (MAQC) project shows inter- and intraplatform reproducibility of gene expression measurements.
    Nat Biotechnol. 2006 Sep;24(9):1151-61 PMID: 16964229
  97. Comparative assessment of methods for the computational inference of transcript isoform abundance from RNA-seq data.
    Genome Biol. 2015 Jul 23;16:150 PMID: 26201343
Article Info
Journal
Genome biology
Abbr.
Genome Biol
ISSN
1474-760X
Published
2019-00-20
Epub
2019-00-20
Pages
125
Language
English
Region
England
NLM ID
100960660
PMCID
PMC6584985
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: [email protected]