17  References

Bankevich, Anton, Sergey Nurk, Dmitry Antipov, et al. 2012. SPAdes: A New Genome Assembly Algorithm and Its Applications to Single-Cell Sequencing.” Journal of Computational Biology 19 (5): 455–77. https://doi.org/10.1089/cmb.2012.0021.
Bolger, Anthony M., Marc Lohse, and Bjoern Usadel. 2014. “Trimmomatic: A Flexible Trimmer for Illumina Sequence Data.” Bioinformatics 30 (15): 2114–20. https://doi.org/10.1093/bioinformatics/btu170.
Bonfield, James K., John Marshall, Petr Danecek, et al. 2021. HTSlib: C Library for Reading/Writing High-Throughput Sequencing Data.” GigaScience 10 (2): giab007. https://doi.org/10.1093/gigascience/giab007.
Bray, Nicolas L., Harold Pimentel, Páll Melsted, and Lior Pachter. 2016. “Near-Optimal Probabilistic RNA-seq Quantification.” Nature Biotechnology 34 (5): 525–27. https://doi.org/10.1038/nbt.3519.
Chen, Shifu, Yanqing Zhou, Yaru Chen, and Jia Gu. 2018. “Fastp: An Ultra-Fast All-in-One FASTQ Preprocessor.” Bioinformatics 34 (17): i884–90. https://doi.org/10.1093/bioinformatics/bty560.
Cheng, Haoyu, Gregory T. Concepcion, Xiaowen Feng, Haowen Zhang, and Heng Li. 2021. “Haplotype-Resolved de Novo Assembly Using Phased Assembly Graphs with Hifiasm.” Nature Methods 18 (2): 170–75. https://doi.org/10.1038/s41592-020-01056-5.
Cleary, John G., Ross Braithwaite, Kurt Gaastra, et al. 2015. “Comparing Variant Call Files for Performance Benchmarking of Next-Generation Sequencing Variant Calling Pipelines.” bioRxiv, ahead of print. https://doi.org/10.1101/023754.
Cock, Peter J. A., Tiago Antao, Jeffrey T. Chang, et al. 2009. “Biopython: Freely Available Python Tools for Computational Molecular Biology and Bioinformatics.” Bioinformatics 25 (11): 1422–23. https://doi.org/10.1093/bioinformatics/btp163.
Cooke, Daniel P., David C. Wedge, and Gerton Lunter. 2021. “A Unified Haplotype-Based Method for Accurate and Comprehensive Variant Calling.” Nature Biotechnology 39 (7): 885–92. https://doi.org/10.1038/s41587-021-00861-3.
Crusoe, Michael R., Sanne Abeln, Alexandru Iosup, et al. 2022. “Methods Included: Standardizing Computational Reuse and Portability with the Common Workflow Language.” Communications of the ACM 65 (6): 54–63. https://doi.org/10.1145/3486897.
Danecek, Petr, Adam Auton, Goncalo Abecasis, et al. 2011. “The Variant Call Format and VCFtools.” Bioinformatics 27 (15): 2156–58. https://doi.org/10.1093/bioinformatics/btr330.
Danecek, Petr, James K. Bonfield, Jennifer Liddle, et al. 2021. “Twelve Years of SAMtools and BCFtools.” GigaScience 10 (2): giab008. https://doi.org/10.1093/gigascience/giab008.
De Coster, Wouter, and Rosa Rademakers. 2023. NanoPack2: Population-Scale Evaluation of Long-Read Sequencing Data.” Bioinformatics 39 (5): btad311. https://doi.org/10.1093/bioinformatics/btad311.
Di Tommaso, Paolo, Maria Chatzou, Evan W. Floden, Pablo Prieto Barja, Emilio Palumbo, and Cedric Notredame. 2017. “Nextflow Enables Reproducible Computational Workflows.” Nature Biotechnology 35 (4): 316–19. https://doi.org/10.1038/nbt.3820.
Dobin, Alexander, Carrie A. Davis, Felix Schlesinger, et al. 2013. STAR: Ultrafast Universal RNA-seq Aligner.” Bioinformatics 29 (1): 15–21. https://doi.org/10.1093/bioinformatics/bts635.
Ewels, Philip, Måns Magnusson, Sverker Lundin, and Max Käller. 2016. MultiQC: Summarize Analysis Results for Multiple Tools and Samples in a Single Report.” Bioinformatics 32 (19): 3047–48. https://doi.org/10.1093/bioinformatics/btw354.
Ferragina, Paolo, and Giovanni Manzini. 2000. “Opportunistic Data Structures with Applications.” Proceedings 41st Annual Symposium on Foundations of Computer Science (FOCS), 390–98. https://doi.org/10.1109/SFCS.2000.892127.
Garrison, Erik, and Gabor Marth. 2012. Haplotype-Based Variant Detection from Short-Read Sequencing. https://arxiv.org/abs/1207.3907.
Gurevich, Alexey, Vladislav Saveliev, Nikolay Vyahhi, and Glenn Tesler. 2013. QUAST: Quality Assessment Tool for Genome Assemblies.” Bioinformatics 29 (8): 1072–75. https://doi.org/10.1093/bioinformatics/btt086.
Hsi-Yang Fritz, Markus, Rasko Leinonen, Guy Cochrane, and Ewan Birney. 2011. “Efficient Storage of High Throughput DNA Sequencing Data Using Reference-Based Compression.” Genome Research 21 (5): 734–40. https://doi.org/10.1101/gr.114819.110.
Huang, Neng, and Heng Li. 2023. “Compleasm: A Faster and More Accurate Reimplementation of BUSCO.” Bioinformatics 39 (10): btad595. https://doi.org/10.1093/bioinformatics/btad595.
Kim, Daehwan, Joseph M. Paggi, Chanhee Park, Christopher Bennett, and Steven L. Salzberg. 2019. “Graph-Based Genome Alignment and Genotyping with HISAT2 and HISAT-genotype.” Nature Biotechnology 37 (8): 907–15. https://doi.org/10.1038/s41587-019-0201-4.
Kim, Sangtae, Konrad Scheffler, Aaron L. Halpern, et al. 2018. “Strelka2: Fast and Accurate Calling of Germline and Somatic Variants.” Nature Methods 15 (8): 591–94. https://doi.org/10.1038/s41592-018-0051-x.
Kolmogorov, Mikhail, Jeffrey Yuan, Yu Lin, and Pavel A. Pevzner. 2019. “Assembly of Long, Error-Prone Reads Using Repeat Graphs.” Nature Biotechnology 37 (5): 540–46. https://doi.org/10.1038/s41587-019-0072-8.
Koren, Sergey, Brian P. Walenz, Konstantin Berlin, Jason R. Miller, Nicholas H. Bergman, and Adam M. Phillippy. 2017. “Canu: Scalable and Accurate Long-Read Assembly via Adaptive k-Mer Weighting and Repeat Separation.” Genome Research 27 (5): 722–36. https://doi.org/10.1101/gr.215087.116.
Köster, Johannes, and Sven Rahmann. 2012. “Snakemake — a Scalable Bioinformatics Workflow Engine.” Bioinformatics 28 (19): 2520–22. https://doi.org/10.1093/bioinformatics/bts480.
Langmead, Ben, and Steven L. Salzberg. 2012. “Fast Gapped-Read Alignment with Bowtie 2.” Nature Methods 9 (4): 357–59. https://doi.org/10.1038/nmeth.1923.
Law, Charity W., Yunshun Chen, Wei Shi, and Gordon K. Smyth. 2014. “Voom: Precision Weights Unlock Linear Model Analysis Tools for RNA-seq Read Counts.” Genome Biology 15 (2): R29. https://doi.org/10.1186/gb-2014-15-2-r29.
Lawrence, Michael, Robert Gentleman, and Vincent Carey. 2009. “Rtracklayer: An R Package for Interfacing with Genome Browsers.” Bioinformatics 25 (14): 1841–42. https://doi.org/10.1093/bioinformatics/btp328.
Li, Bo, and Colin N. Dewey. 2011. RSEM: Accurate Transcript Quantification from RNA-Seq Data with or Without a Reference Genome.” BMC Bioinformatics 12 (1): 323. https://doi.org/10.1186/1471-2105-12-323.
Li, Heng. 2011. “Tabix: Fast Retrieval of Sequence Features from Generic TAB-Delimited Files.” Bioinformatics 27 (5): 718–19. https://doi.org/10.1093/bioinformatics/btq671.
Li, Heng. 2013. Aligning Sequence Reads, Clone Sequences and Assembly Contigs with BWA-MEM. https://arxiv.org/abs/1303.3997.
Li, Heng. 2018. “Minimap2: Pairwise Alignment for Nucleotide Sequences.” Bioinformatics 34 (18): 3094–100. https://doi.org/10.1093/bioinformatics/bty191.
Li, Heng, and Richard Durbin. 2009. “Fast and Accurate Short Read Alignment with Burrows–Wheeler Transform.” Bioinformatics 25 (14): 1754–60. https://doi.org/10.1093/bioinformatics/btp324.
Li, Heng, Bob Handsaker, Alec Wysoker, et al. 2009. “The Sequence Alignment/Map Format and SAMtools.” Bioinformatics 25 (16): 2078–79. https://doi.org/10.1093/bioinformatics/btp352.
Liao, Yang, Gordon K. Smyth, and Wei Shi. 2014. featureCounts: An Efficient General Purpose Program for Assigning Sequence Reads to Genomic Features.” Bioinformatics 30 (7): 923–30. https://doi.org/10.1093/bioinformatics/btt656.
Love, Michael I., Wolfgang Huber, and Simon Anders. 2014. “Moderated Estimation of Fold Change and Dispersion for RNA-seq Data with DESeq2.” Genome Biology 15 (12): 550. https://doi.org/10.1186/s13059-014-0550-8.
Manni, Mosè, Matthew R. Berkeley, Mathieu Seppey, Felipe A. Simão, and Evgeny M. Zdobnov. 2021. BUSCO Update: Novel and Streamlined Workflows Along with Broader and Deeper Phylogenetic Coverage for Scoring of Eukaryotic, Prokaryotic, and Viral Genomes.” Molecular Biology and Evolution 38 (10): 4647–54. https://doi.org/10.1093/molbev/msab199.
Martin, Marcel. 2011. “Cutadapt Removes Adapter Sequences from High-Throughput Sequencing Reads.” EMBnet.journal 17 (1): 10. https://doi.org/10.14806/ej.17.1.200.
McKenna, Aaron, Matthew Hanna, Eric Banks, et al. 2010. “The Genome Analysis Toolkit: A MapReduce Framework for Analyzing Next-Generation DNA Sequencing Data.” Genome Research 20 (9): 1297–303. https://doi.org/10.1101/gr.107524.110.
Mölder, Felix, Kim Philipp Jablonski, Brice Letcher, et al. 2025. “Sustainable Data Analysis with Snakemake.” F1000Research 10: 33. https://doi.org/10.12688/f1000research.29032.3.
Needleman, Saul B., and Christian D. Wunsch. 1970. “A General Method Applicable to the Search for Similarities in the Amino Acid Sequence of Two Proteins.” Journal of Molecular Biology 48 (3): 443–53. https://doi.org/10.1016/0022-2836(70)90057-4.
Neph, Shane, M. Scott Kuehn, Alex P. Reynolds, et al. 2012. BEDOPS: High-Performance Genomic Feature Operations.” Bioinformatics 28 (14): 1919–20. https://doi.org/10.1093/bioinformatics/bts277.
Patro, Rob, Geet Duggal, Michael I. Love, Rafael A. Irizarry, and Carl Kingsford. 2017. “Salmon Provides Fast and Bias-Aware Quantification of Transcript Expression.” Nature Methods 14 (4): 417–19. https://doi.org/10.1038/nmeth.4197.
Poplin, Ryan, Pi-Chuan Chang, David Alexander, et al. 2018. “A Universal SNP and Small-Indel Variant Caller Using Deep Neural Networks.” Nature Biotechnology 36 (10): 983–87. https://doi.org/10.1038/nbt.4235.
Quinlan, Aaron R., and Ira M. Hall. 2010. BEDTools: A Flexible Suite of Utilities for Comparing Genomic Features.” Bioinformatics 26 (6): 841–42. https://doi.org/10.1093/bioinformatics/btq033.
Rautiainen, Mikko, Sergey Nurk, Brian P. Walenz, et al. 2023. “Telomere-to-Telomere Assembly of Diploid Chromosomes with Verkko.” Nature Biotechnology 41 (10): 1474–82. https://doi.org/10.1038/s41587-023-01662-6.
Rhie, Arang, Brian P. Walenz, Sergey Koren, and Adam M. Phillippy. 2020. “Merqury: Reference-Free Quality, Completeness, and Phasing Assessment for Genome Assemblies.” Genome Biology 21 (1): 245. https://doi.org/10.1186/s13059-020-02134-9.
Ritchie, Matthew E., Belinda Phipson, Di Wu, et al. 2015. “Limma Powers Differential Expression Analyses for RNA-Sequencing and Microarray Studies.” Nucleic Acids Research 43 (7): e47. https://doi.org/10.1093/nar/gkv007.
Roberts, Michael, Wayne Hayes, Brian R. Hunt, Stephen M. Mount, and James A. Yorke. 2004. “Reducing Storage Requirements for Biological Sequence Comparison.” Bioinformatics 20 (18): 3363–69. https://doi.org/10.1093/bioinformatics/bth408.
Robinson, Mark D., Davis J. McCarthy, and Gordon K. Smyth. 2010. edgeR: A Bioconductor Package for Differential Expression Analysis of Digital Gene Expression Data.” Bioinformatics 26 (1): 139–40. https://doi.org/10.1093/bioinformatics/btp616.
Sahlin, Kristoffer. 2022. “Strobealign: Flexible Seed Size Enables Ultra-Fast and Accurate Read Alignment.” Genome Biology 23 (1): 260. https://doi.org/10.1186/s13059-022-02831-7.
Sena Brandine, Guilherme de, and Andrew D. Smith. 2019. “Falco: High-Speed FastQC Emulation for Quality Control of Sequencing Data.” F1000Research 8: 1874. https://doi.org/10.12688/f1000research.21142.1.
Shen, Wei, Shuai Le, Yan Li, and Fuquan Hu. 2016. SeqKit: A Cross-Platform and Ultrafast Toolkit for FASTA/Q File Manipulation.” PLOS ONE 11 (10): e0163962. https://doi.org/10.1371/journal.pone.0163962.
Smith, Temple F., and Michael S. Waterman. 1981. “Identification of Common Molecular Subsequences.” Journal of Molecular Biology 147 (1): 195–97. https://doi.org/10.1016/0022-2836(81)90087-5.
Smith, Tom, Andreas Heger, and Ian Sudbery. 2017. UMI-tools: Modeling Sequencing Errors in Unique Molecular Identifiers to Improve Quantification Accuracy.” Genome Research 27 (3): 491–99. https://doi.org/10.1101/gr.209601.116.
Soneson, Charlotte, Michael I. Love, and Mark D. Robinson. 2016. “Differential Analyses for RNA-seq: Transcript-Level Estimates Improve Gene-Level Inferences.” F1000Research 4: 1521. https://doi.org/10.12688/f1000research.7563.2.
Tan, Adrian, Gonçalo R. Abecasis, and Hyun Min Kang. 2015. “Unified Representation of Genetic Variants.” Bioinformatics 31 (13): 2202–4. https://doi.org/10.1093/bioinformatics/btv112.
Vasimuddin, Md., Sanchit Misra, Heng Li, and Srinivas Aluru. 2019. “Efficient Architecture-Aware Acceleration of BWA-MEM for Multicore Systems.” 2019 IEEE International Parallel and Distributed Processing Symposium (IPDPS), 314–24. https://doi.org/10.1109/IPDPS.2019.00041.
Wingett, Steven W., and Simon Andrews. 2018. FastQ Screen: A Tool for Multi-Genome Mapping and Quality Control.” F1000Research 7: 1338. https://doi.org/10.12688/f1000research.15931.2.
Zheng, Zhenxian, Shumin Li, Junhao Su, Amy Wing-Sze Leung, Tak-Wah Lam, and Ruibang Luo. 2022. “Symphonizing Pileup and Full-Alignment for Deep Learning-Based Long-Read Variant Calling.” Nature Computational Science 2 (12): 797–803. https://doi.org/10.1038/s43588-022-00387-x.
Zook, Justin M., Jennifer McDaniel, Nathan D. Olson, et al. 2019. “An Open Resource for Accurately Benchmarking Small Variant and Reference Calls.” Nature Biotechnology 37 (5): 561–66. https://doi.org/10.1038/s41587-019-0074-6.