New platform enables ultrafast, reference-free sequence searches across millions of single cells
From Pepkio Team · 3 September 2026 · 3 min read
Scientists report today in Nature the launch of Malva, a computational platform that allows researchers to search for any RNA sequence – from viral genomes to cancer mutations – across millions of single cells in seconds, without needing a reference genome or downloading massive datasets. The work, led by senior author Nikolaus Rajewsky at the Max Delbrück Center (MDC) in Berlin, with first author Daniel León-Periñán, transforms single-cell atlases from static gene-count tables into dynamic, sequence-resolved resources.
Single-cell and spatial transcriptomics generate petabytes of raw sequencing data each year, but standard pipelines discard sequence information by mapping reads to a reference genome and retaining only gene counts. This makes it impossible to search for sequences that don’t match the reference – such as tumour-specific mutations, viral RNAs, or unannotated isoforms – without reprocessing the entire dataset.
Malva indexes raw sequencing reads by breaking them into short subsequences (k-mers) and storing them alongside the cell barcode of origin. The current Malva Index comprises around 74 million cells from thousands of experiments, covering healthy and diseased human tissues as well as mouse samples. Queries take milliseconds for short probes and seconds for full transcripts, using a fraction of the memory required by existing tools.
To demonstrate Malva’s capabilities, the team showed that it can:
- Detect viral sequences (e.g., HERV-K, lentiviral vectors) and common lab contaminants such as Mycoplasma.
- Recapitulate germline SNP frequencies from large populations.
- Identify cell-type-specific isoform usage (e.g., CD45 isoforms in immune cells).
- Find circular RNAs (e.g., CDR1as) in brain cells.
- Detect cancer somatic mutations across 16 tumour types with high sensitivity and specificity.
Beyond search, Malva can perform reference-free clustering of cells based purely on sequence composition, and even assemble marker sequences that reveal unannotated transcripts or microbial signals missed by gene-centric analyses.
“Malva provides a practical discovery layer for asking sequence-defined questions in single-cell biology,” the authors write. The platform is accessible via a public API, and the index is continuously expanded as new public data become available.
Limitations: The tool uses exact k-mer matching, so sensitivity drops with very high sequencing error rates (above 2%). It reports pseudocounts rather than absolute molecule counts, and results may need validation. The current index mostly covers short-read, 3′-biased data.
The authors envision Malva as a bridge between human and machine reasoning about biology, enabling AI models to retrieve real cellular evidence during analysis. As single-cell datasets grow, such searchable indexes could become a standard way to reuse public data – turning massive archives into live molecular resources.
Reference:
León-Periñán, D., Karaiskos, N. & Rajewsky, N. Ultrafast and reference-free sequence discovery in single-cell data. Nature (2026). DOI: 10.1038/s41586-026-10975-w
Explore Pepkio
- Bioinformatics CRO
Reproducible, publication-style analyses with full source code and methods for academic labs and biotech teams.
- The bioinformatics outsourcing playbook
Cost, timelines, vendor selection, and reproducibility for labs weighing whether to outsource bioinformatics.
- Free AI-assisted lab tools
Browser calculators for serial dilutions, molarity, PCR setup, plate readers, and more — no account required.