CS Seminar: “Sequence analysis at scale: Accurate algorithms for indexing, comparing, and decontaminating the world’s biological sequence data”, Harun Mustafa (Johns Hopkins University), 1:30PM October 9, 2026 (EN)

Sequence analysis at scale: Accurate algorithms for indexing, comparing, and decontaminating the world’s biological sequence data

Dr. Harun Mustafa
Johns Hopkins University

Abstract: Public and controlled-access repositories currently hold petabytes of biological sequence data spanning every domain of life, from curated reference genomes to error-prone collections of unassembled sequencing reads, and continue to grow exponentially. Classical sequence comparison algorithms were each designed for a specific regime: aligning two short sequences, searching for a short query in a reference genome database, or finding locally similar regions between two genomes. With these tools, searching entire archives of unassembled reads and accurately comparing complete genomes end to end have until recently been out of reach for much of the genomics community. I will present nearly a decade of work on succinct data structures and alignment algorithms that make such analyses computationally feasible and accurate. MetaGraph indexes sequence collections as annotated de Bruijn graphs, whose lossless succinct representations enable full-text search across vast public archives using either fast exact-word matching or accurate sequence-to-graph alignment. Exploiting the annotation structure, the alignment model can be extended to accommodate sequence content imputation from related database entries during search. I will demonstrate these indexes on large-cohort analyses such as antimicrobial resistance tracing. Further extensions allow alignment models to better account for the complex rearrangements that accumulate over evolutionary time. These models can provide more comprehensive maps of the genetic differences between humans and, coupled with succinct indexes, uncover long regions of similarity between unrelated species that indicate reference genome contamination. As part of this effort, our rearrangement-aware alignment approach balances speed and accuracy, revealing greater divergence between closely related human haplotypes than widely cited statistics suggest. This work demonstrates that more accurate sequence comparison need not come at an unacceptable cost in running time.

Biography: Harun Mustafa is an assistant research scientist in the Salzberg Lab at Johns Hopkins University. Previously, he was a doctoral student, postdoctoral fellow, and established researcher in the Biomedical Informatics Group at ETH Zurich, advised by Prof. Gunnar Rätsch. His work is primarily in the development and application of algorithms for large-scale sequence comparison and analysis, and the development of novel alignment algorithms that better model complex genomic features. His research has been supported by fellowships awarded by the Swiss National Science Foundation and the ETH Board, and his work on scalable indexing was recognised as a Remarkable Output of 2025 by the Swiss Institute of Bioinformatics.

DATE: October 09, 2026 Friday @ 13:30

Place: EA 502