Publications

Peer-reviewed papers and preprints from Nygen Analytics and its research collaborators across single-cell RNA-seq, multi-omics, and computational biology. Each entry links to its DOI.

← Resources

Transcriptional Roadmap of the Human Airway Epithelium Identifying HLF as a Novel Regulator of Basal Stem Cell Function

Pavan Prabhala, Sofia Freiman, Nika Gvazava, Jiten Sharma, Sofia Wijk, Karina Kanzenbach, Stefan Lang, Rebecka Cattani, Jenny Wigén, David Bryder, Shamit Soneji, Johan Flygare, Ellen Tufvesson, Leif Bjermer, Darcy Wagner, Gunilla Westergren Thorsson, Mattias Magnusson

bioRxiv 2025-05-07 DOI: 10.1101/2025.05.06.650933

Abstract licence: bioRxiv preprint, no reuse licence set

Multi-agent AI enables evidence-based cell annotation in single-cell transcriptomics

Gautam Ahuja, Alex Antill, Yi Su, Giovanni Marco Dall’Olio, Sukhitha Basnayake, Göran Karlsson, Parashar Dhapola

bioRxiv 2025-11-07 DOI: 10.1101/2025.11.06.686964

Abstract Cell type annotation remains a critical bottleneck, with current methods often inaccurate and requiring extensive manual validation, particularly in disease contexts. While large language models (LLMs) show promise, they can be unreliable due to hallucinations. We developed CyteType, a multi-agent framework that generates competing hypotheses grounded in full expression data and study context, validates against external databases, and iteratively self-evaluates. Comprehensive benchmarking demonstrates that CyteType substantially outperforms reference-based and LLM-based methods, with self-generated confidence scores reliably identifying trustworthy annotations. CyteType transforms cell type annotation from label assignment into evidence-grounded biological discovery. Python (AnnData compatible): https://github.com/NygenAnalytics/CyteType R (Seurat compatible): https://github.com/NygenAnalytics/CyteTypeR

Abstract licence: CC BY-NC 4.0

Alphavirus replicon particle expressing IL-12 reprograms tumor-associated macrophages and neutrophils and induces anti-tumor immunity

Momoko Ishikawa, Dhanya K Nambiar, Jaya Sastri, Forrest Bowling, Keiko Ishimoto, Fumiya Tao, Koyo Takahashi, Ivan Stephanek, Hongbin Cao, Davis Leitner, Kenta Matsuda, Fred M Baik, John B. Sunwoo, Quynh-Thu Le, Jonathan F Smith, Wataru Akahata

bioRxiv 2025-08-11 DOI: 10.1101/2025.08.07.666585

Abstract licence: bioRxiv preprint, no reuse licence set

CTEC: a cross-tabulation ensemble clustering approach for single-cell RNA sequencing data analysis

Liang Wang, Chenyang Hong, Jiangning Song, Jianhua Yao

Bioinformatics 2024-03-29 Vol. 40, btae130 DOI: 10.1093/bioinformatics/btae130

Abstract Motivation. Cell-type clustering is a crucial first step for single-cell RNA-seq data analysis. However, existing clustering methods often provide different results on cluster assignments with respect to their own data pre-processing, choice of distance metrics, and strategies of feature extraction, thereby limiting their practical applications. Results. We propose Cross-Tabulation Ensemble Clustering (CTEC) method that formulates two re-clustering strategies (distribution- and outlier-based) via cross-tabulation. Benchmarking experiments on five scRNA-Seq datasets illustrate that the proposed CTEC method offers significant improvements over the individual clustering methods. Moreover, CTEC-DB outperforms the state-of-the-art ensemble methods for single-cell data clustering, with 45.4% and 17.1% improvement over the single-cell aggregated from ensemble clustering method (SAFE) and the single-cell aggregated clustering via Mixture model ensemble method (SAME), respectively, on the two-method ensemble test. Availability and implementation. The source code of the benchmark in this work is available at the GitHub repository https://github.com/LWCHN/CTEC.git.

Abstract licence: CC BY 4.0

EpiCarousel: memory- and time-efficient identification of metacells for atlas-level single-cell chromatin accessibility data

Sijie Li, Yuxi Li, Yu Sun, Yaru Li, Xiaoyang Chen, Songming Tang, Shengquan Chen

Bioinformatics 2024-03-29 Vol. 40, btae191 DOI: 10.1093/bioinformatics/btae191

Abstract Summary. Recent technical advancements in single-cell chromatin accessibility sequencing (scCAS) have brought new insights to the characterization of epigenetic heterogeneity. As single-cell genomics experiments scale up to hundreds of thousands of cells, the demand for computational resources for downstream analysis grows intractably large and exceeds the capabilities of most researchers. Here, we propose EpiCarousel, a tailored Python package based on lazy loading, parallel processing, and community detection for memory- and time-efficient identification of metacells, i.e. the emergence of homogenous cells, in large-scale scCAS data. Through comprehensive experiments on five datasets of various protocols, sample sizes, dimensions, number of cell types, and degrees of cell-type imbalance, EpiCarousel outperformed baseline methods in systematic evaluation of memory usage, computational time, and multiple downstream analyses including cell type identification. Moreover, EpiCarousel executes preprocessing and downstream cell clustering on the atlas-level dataset with 707 043 cells and 1 154 611 peaks within 2 h consuming <75 GB of RAM and provides superior performance for characterizing cell heterogeneity than state-of-the-art methods. Availability and implementation. The EpiCarousel software is well-documented and freely available at https://github.com/biox-nku/epicarousel. It can be seamlessly interoperated with extensive scCAS analysis toolkits.

Abstract licence: CC BY 4.0

Single cell multi-omics analysis of chronic myeloid leukemia links cellular heterogeneity to therapy response

Rebecca Warfvinge, Linda Geironson Ulfsson, Parashar Dhapola, Fatemeh Safi, Mikael Sommarin, Shamit Soneji, Henrik Hjorth-Hansen, Satu Mustjoki, Johan Richter, Ram Krishna Thakur, Göran Karlsson

eLife 2024-11-06 Vol. 12, RP92074 DOI: 10.7554/eLife.92074

Abstract The advent of tyrosine kinase inhibitors (TKIs) as treatment of chronic myeloid leukemia (CML) is a paradigm in molecularly targeted cancer therapy. Nonetheless, TKI-insensitive leukemia stem cells (LSCs) persist in most patients even after years of treatment and are imperative for disease progression as well as recurrence during treatment-free remission (TFR). Here, we have generated high-resolution single-cell multiomics maps from CML patients at diagnosis, retrospectively stratified by BCR::ABL1IS (%) following 12 months of TKI therapy. Simultaneous measurement of global gene expression profiles together with >40 surface markers from the same cells revealed that each patient harbored a unique composition of stem and progenitor cells at diagnosis. The patients with treatment failure after 12 months of therapy had a markedly higher abundance of molecularly defined primitive cells at diagnosis compared to the optimal responders. The multiomic feature landscape enabled visualization of the primitive fraction as a mixture of molecularly distinct BCR::ABL1+ LSCs and BCR::ABL1-hematopoietic stem cells (HSCs) in variable ratio across patients, and guided their prospective isolation by a combination of CD26 and CD35 cell surface markers. We for the first time show that BCR::ABL1+ LSCs and BCR::ABL1- HSCs can be distinctly separated as CD26+CD35- and CD26-CD35+, respectively. In addition, we found the ratio of LSC/HSC to be higher in patients with prospective treatment failure compared to optimal responders, at diagnosis as well as following 3 months of TKI therapy. Collectively, this data builds a framework for understanding therapy response and adapting treatment by devising strategies to extinguish or suppress TKI-insensitive LSCs.

Abstract licence: CC BY 4.0

Tracking early mammalian organogenesis – prediction and validation of differentiation trajectories at whole organism scale

Ivan Imaz-Rosshandler, Christina Rode, Carolina Guibentif, Luke T. G. Harland, Mai-Linh N. Ton, Parashar Dhapola, Daniel Keitley, Ricard Argelaguet, Fernando J. Calero-Nieto, Jennifer Nichols, John C. Marioni, Marella F. T. R. de Bruijn, Berthold Göttgens

Development 2024-01-31 Vol. 151, dev201867 DOI: 10.1242/dev.201867

Abstract Early organogenesis represents a key step in animal development, during which pluripotent cells diversify to initiate organ formation. Here, we sampled 300,000 single-cell transcriptomes from mouse embryos between E8.5 and E9.5 in 6-h intervals and combined this new dataset with our previous atlas (E6.5-E8.5) to produce a densely sampled timecourse of >400,000 cells from early gastrulation to organogenesis. Computational lineage reconstruction identified complex waves of blood and endothelial development, including a new programme for somite-derived endothelium. We also dissected the E7.5 primitive streak into four adjacent regions, performed scRNA-seq and predicted cell fates computationally. Finally, we defined developmental state/fate relationships by combining orthotopic grafting, microscopic analysis and scRNA-seq to transcriptionally determine cell fates of grafted primitive streak regions after 24 h of in vitro embryo culture. Experimentally determined fate outcomes were in good agreement with computationally predicted fates, demonstrating how classical grafting experiments can be revisited to establish high-resolution cell state/fate relationships. Such interdisciplinary approaches will benefit future studies in developmental biology and guide the in vitro production of cells for organ regeneration and repair.

Abstract licence: CC BY 4.0

Spatiotemporal transcriptomic map of glial cell response in a mouse model of acute brain ischemia

Daniel Zucha, Pavel Abaffy, Denisa Kirdajova, Daniel Jirak, Mikael Kubista, Miroslava Anderova, Lukas Valihrach

Proceedings of the National Academy of Sciences 2024-11-05 Vol. 121, e2404203121 DOI: 10.1073/pnas.2404203121

Abstract The role of nonneuronal cells in the resolution of cerebral ischemia remains to be fully understood. To decode key molecular and cellular processes that occur after ischemia, we performed spatial and single-cell transcriptomic profiling of the male mouse brain during the first week of injury. Cortical gene expression was severely disrupted, defined by inflammation and cell death in the lesion core, and glial scar formation orchestrated by multiple cell types on the periphery. The glial scar was identified as a zone with intense cell–cell communication, with prominent ApoE-Trem2 signaling pathway modulating microglial activation. For each of the three major glial populations, an inflammatory-responsive state, resembling the reactive states observed in neurodegenerative contexts, was observed. The recovered spectrum of ischemia-induced oligodendrocyte states supports the emerging hypothesis that oligodendrocytes actively respond to and modulate the neuroinflammatory stimulus. The findings are further supported by analysis of other spatial transcriptomic datasets from different mouse models of ischemic brain injury. Collectively, we present a landmark transcriptomic dataset accompanied by interactive visualization that provides a comprehensive view of spatiotemporal organization of processes in the postischemic mouse brain.

Abstract licence: CC BY 4.0

Scarf enables a highly memory-efficient analysis of large-scale single-cell genomics data

Parashar Dhapola, Johan Rodhe, Rasmus Olofzon, Thomas Bonald, Eva Erlandsson, Shamit Soneji, Göran Karlsson

Nature Communications (Springer Nature) 2022-08-08 Vol. 13, 4616 DOI: 10.1038/s41467-022-32097-3

In plain language Single-cell datasets grew past the point where most labs could process them, and the binding constraint was memory rather than compute time. Analysing a few million cells meant holding them in RAM, which put atlas-scale reanalysis behind a hardware budget most groups do not have. Scarf keeps data on disk and processes it in chunks instead of loading everything into memory, so a laptop can do work that previously needed a cluster. It handles scRNA-seq, scATAC-seq and CITE-seq, interoperates with the rest of the Python single-cell stack, and includes a subsampling method that preserves rare populations and differentiation trajectories rather than sampling them away. Scarf is the engine underneath ScarfWeb.
Abstract As the scale of single-cell genomics experiments grows into the millions, the computational requirements to process this data are beyond the reach of many. Herein we present Scarf, a modularly designed Python package that seamlessly interoperates with other single-cell toolkits and allows for memory-efficient single-cell analysis of millions of cells on a laptop or low-cost devices like single-board computers. We demonstrate Scarf’s memory and compute-time efficiency by applying it to the largest existing single-cell RNA-Seq and ATAC-Seq datasets. Scarf wraps memory-efficient implementations of a graph-based t-stochastic neighbour embedding and hierarchical clustering algorithm. Moreover, Scarf performs accurate reference-anchored mapping of datasets while maintaining memory efficiency. By implementing a subsampling algorithm, Scarf additionally has the capacity to generate representative sampling of cells from a given dataset wherein rare cell populations and lineage differentiation trajectories are conserved. Together, Scarf provides a framework wherein any researcher can perform advanced processing, subsampling, reanalysis, and integration of atlas-scale datasets on standard laptop computers. Scarf is available on Github: https://github.com/parashardhapola/scarf .

Abstract licence: CC BY 4.0