Ram Krishna Thakur, Göran Karlsson
Blood
2026-01-29
Vol. 147
DOI: 10.1182/blood.2025029011
Abstract licence: CC BY-NC-ND 4.0
Pavan Prabhala, Sofia Freiman, Nika Gvazava, Jiten Sharma, Sofia Wijk, Karina Kanzenbach, Stefan Lang, Rebecka Cattani, Jenny Wigén, David Bryder, Shamit Soneji, Johan Flygare, Ellen Tufvesson, Leif Bjermer, Darcy Wagner, Gunilla Westergren Thorsson, Mattias Magnusson
bioRxiv
2025-05-07
DOI: 10.1101/2025.05.06.650933
Abstract licence: bioRxiv preprint, no reuse licence set
Sara Palo, Keiki Nagaharu, Mikael Sommarin, Rasmus Olofzon, Virginia Turati, Shamit Soneji, Göran Karlsson, Charlotta Böiers
bioRxiv
2025-12-04
DOI: 10.64898/2025.12.02.691283
Abstract licence: CC BY-NC-ND 4.0
Gautam Ahuja, Alex Antill, Yi Su, Giovanni Marco Dall’Olio, Sukhitha Basnayake, Göran Karlsson, Parashar Dhapola
bioRxiv
2025-11-07
DOI: 10.1101/2025.11.06.686964
Abstract
Cell type annotation remains a critical bottleneck, with current methods often inaccurate and requiring extensive manual validation, particularly in disease contexts. While large language models (LLMs) show promise, they can be unreliable due to hallucinations. We developed CyteType, a multi-agent framework that generates competing hypotheses grounded in full expression data and study context, validates against external databases, and iteratively self-evaluates. Comprehensive benchmarking demonstrates that CyteType substantially outperforms reference-based and LLM-based methods, with self-generated confidence scores reliably identifying trustworthy annotations. CyteType transforms cell type annotation from label assignment into evidence-grounded biological discovery. Python (AnnData compatible): https://github.com/NygenAnalytics/CyteType R (Seurat compatible): https://github.com/NygenAnalytics/CyteTypeR
Abstract licence: CC BY-NC 4.0
Momoko Ishikawa, Dhanya K Nambiar, Jaya Sastri, Forrest Bowling, Keiko Ishimoto, Fumiya Tao, Koyo Takahashi, Ivan Stephanek, Hongbin Cao, Davis Leitner, Kenta Matsuda, Fred M Baik, John B. Sunwoo, Quynh-Thu Le, Jonathan F Smith, Wataru Akahata
bioRxiv
2025-08-11
DOI: 10.1101/2025.08.07.666585
Abstract licence: bioRxiv preprint, no reuse licence set
Liang Wang, Chenyang Hong, Jiangning Song, Jianhua Yao
Bioinformatics
2024-03-29
Vol. 40, btae130
DOI: 10.1093/bioinformatics/btae130
Abstract
Motivation. Cell-type clustering is a crucial first step for single-cell RNA-seq data analysis. However, existing clustering methods often provide different results on cluster assignments with respect to their own data pre-processing, choice of distance metrics, and strategies of feature extraction, thereby limiting their practical applications. Results. We propose Cross-Tabulation Ensemble Clustering (CTEC) method that formulates two re-clustering strategies (distribution- and outlier-based) via cross-tabulation. Benchmarking experiments on five scRNA-Seq datasets illustrate that the proposed CTEC method offers significant improvements over the individual clustering methods. Moreover, CTEC-DB outperforms the state-of-the-art ensemble methods for single-cell data clustering, with 45.4% and 17.1% improvement over the single-cell aggregated from ensemble clustering method (SAFE) and the single-cell aggregated clustering via Mixture model ensemble method (SAME), respectively, on the two-method ensemble test. Availability and implementation. The source code of the benchmark in this work is available at the GitHub repository https://github.com/LWCHN/CTEC.git.
Abstract licence: CC BY 4.0
Sijie Li, Yuxi Li, Yu Sun, Yaru Li, Xiaoyang Chen, Songming Tang, Shengquan Chen
Bioinformatics
2024-03-29
Vol. 40, btae191
DOI: 10.1093/bioinformatics/btae191
Abstract
Summary. Recent technical advancements in single-cell chromatin accessibility sequencing (scCAS) have brought new insights to the characterization of epigenetic heterogeneity. As single-cell genomics experiments scale up to hundreds of thousands of cells, the demand for computational resources for downstream analysis grows intractably large and exceeds the capabilities of most researchers. Here, we propose EpiCarousel, a tailored Python package based on lazy loading, parallel processing, and community detection for memory- and time-efficient identification of metacells, i.e. the emergence of homogenous cells, in large-scale scCAS data. Through comprehensive experiments on five datasets of various protocols, sample sizes, dimensions, number of cell types, and degrees of cell-type imbalance, EpiCarousel outperformed baseline methods in systematic evaluation of memory usage, computational time, and multiple downstream analyses including cell type identification. Moreover, EpiCarousel executes preprocessing and downstream cell clustering on the atlas-level dataset with 707 043 cells and 1 154 611 peaks within 2 h consuming <75 GB of RAM and provides superior performance for characterizing cell heterogeneity than state-of-the-art methods. Availability and implementation. The EpiCarousel software is well-documented and freely available at https://github.com/biox-nku/epicarousel. It can be seamlessly interoperated with extensive scCAS analysis toolkits.
Abstract licence: CC BY 4.0
Radek Sindelka, Ravindra Naraine, Pavel Abaffy, Daniel Zucha, Daniel Kraus, Jiri Netusil, Karel Smetana, Lukas Lacina, Berwini Beduya Endaya, Jiri Neuzil, Martin Psenicka, Mikael Kubista
Genome Biology
2024-10-01
Vol. 25, 251
DOI: 10.1186/s13059-024-03396-3
Abstract licence: CC BY-NC-ND 4.0
Rebecca Warfvinge, Linda Geironson Ulfsson, Parashar Dhapola, Fatemeh Safi, Mikael Sommarin, Shamit Soneji, Henrik Hjorth-Hansen, Satu Mustjoki, Johan Richter, Ram Krishna Thakur, Göran Karlsson
eLife
2024-11-06
Vol. 12, RP92074
DOI: 10.7554/eLife.92074
Abstract
The advent of tyrosine kinase inhibitors (TKIs) as treatment of chronic myeloid leukemia (CML) is a paradigm in molecularly targeted cancer therapy. Nonetheless, TKI-insensitive leukemia stem cells (LSCs) persist in most patients even after years of treatment and are imperative for disease progression as well as recurrence during treatment-free remission (TFR). Here, we have generated high-resolution single-cell multiomics maps from CML patients at diagnosis, retrospectively stratified by BCR::ABL1IS (%) following 12 months of TKI therapy. Simultaneous measurement of global gene expression profiles together with >40 surface markers from the same cells revealed that each patient harbored a unique composition of stem and progenitor cells at diagnosis. The patients with treatment failure after 12 months of therapy had a markedly higher abundance of molecularly defined primitive cells at diagnosis compared to the optimal responders. The multiomic feature landscape enabled visualization of the primitive fraction as a mixture of molecularly distinct BCR::ABL1+ LSCs and BCR::ABL1-hematopoietic stem cells (HSCs) in variable ratio across patients, and guided their prospective isolation by a combination of CD26 and CD35 cell surface markers. We for the first time show that BCR::ABL1+ LSCs and BCR::ABL1- HSCs can be distinctly separated as CD26+CD35- and CD26-CD35+, respectively. In addition, we found the ratio of LSC/HSC to be higher in patients with prospective treatment failure compared to optimal responders, at diagnosis as well as following 3 months of TKI therapy. Collectively, this data builds a framework for understanding therapy response and adapting treatment by devising strategies to extinguish or suppress TKI-insensitive LSCs.
Abstract licence: CC BY 4.0
Ivan Imaz-Rosshandler, Christina Rode, Carolina Guibentif, Luke T. G. Harland, Mai-Linh N. Ton, Parashar Dhapola, Daniel Keitley, Ricard Argelaguet, Fernando J. Calero-Nieto, Jennifer Nichols, John C. Marioni, Marella F. T. R. de Bruijn, Berthold Göttgens
Development
2024-01-31
Vol. 151, dev201867
DOI: 10.1242/dev.201867
Abstract
Early organogenesis represents a key step in animal development, during which pluripotent cells diversify to initiate organ formation. Here, we sampled 300,000 single-cell transcriptomes from mouse embryos between E8.5 and E9.5 in 6-h intervals and combined this new dataset with our previous atlas (E6.5-E8.5) to produce a densely sampled timecourse of >400,000 cells from early gastrulation to organogenesis. Computational lineage reconstruction identified complex waves of blood and endothelial development, including a new programme for somite-derived endothelium. We also dissected the E7.5 primitive streak into four adjacent regions, performed scRNA-seq and predicted cell fates computationally. Finally, we defined developmental state/fate relationships by combining orthotopic grafting, microscopic analysis and scRNA-seq to transcriptionally determine cell fates of grafted primitive streak regions after 24 h of in vitro embryo culture. Experimentally determined fate outcomes were in good agreement with computationally predicted fates, demonstrating how classical grafting experiments can be revisited to establish high-resolution cell state/fate relationships. Such interdisciplinary approaches will benefit future studies in developmental biology and guide the in vitro production of cells for organ regeneration and repair.
Abstract licence: CC BY 4.0
Daniel Zucha, Pavel Abaffy, Denisa Kirdajova, Daniel Jirak, Mikael Kubista, Miroslava Anderova, Lukas Valihrach
Proceedings of the National Academy of Sciences
2024-11-05
Vol. 121, e2404203121
DOI: 10.1073/pnas.2404203121
Abstract
The role of nonneuronal cells in the resolution of cerebral ischemia remains to be fully understood. To decode key molecular and cellular processes that occur after ischemia, we performed spatial and single-cell transcriptomic profiling of the male mouse brain during the first week of injury. Cortical gene expression was severely disrupted, defined by inflammation and cell death in the lesion core, and glial scar formation orchestrated by multiple cell types on the periphery. The glial scar was identified as a zone with intense cell–cell communication, with prominent ApoE-Trem2 signaling pathway modulating microglial activation. For each of the three major glial populations, an inflammatory-responsive state, resembling the reactive states observed in neurodegenerative contexts, was observed. The recovered spectrum of ischemia-induced oligodendrocyte states supports the emerging hypothesis that oligodendrocytes actively respond to and modulate the neuroinflammatory stimulus. The findings are further supported by analysis of other spatial transcriptomic datasets from different mouse models of ischemic brain injury. Collectively, we present a landmark transcriptomic dataset accompanied by interactive visualization that provides a comprehensive view of spatiotemporal organization of processes in the postischemic mouse brain.
Abstract licence: CC BY 4.0
Mikael N. E. Sommarin, Rasmus Olofzon, Sara Palo, Parashar Dhapola, Shamit Soneji, Göran Karlsson, Charlotta Böiers
Blood Advances
2023-09-14
Vol. 7
DOI: 10.1182/bloodadvances.2023009808
Sagnik Yarlagadda, Todd D. Giorgio
The FEBS Journal
2024-01-20
Vol. 291
DOI: 10.1111/febs.17036
Abstract licence: Publisher terms
Parashar Dhapola, Johan Rodhe, Rasmus Olofzon, Thomas Bonald, Eva Erlandsson, Shamit Soneji, Göran Karlsson
Nature Communications (Springer Nature)
2022-08-08
Vol. 13, 4616
DOI: 10.1038/s41467-022-32097-3
In plain language
Single-cell datasets grew past the point where most labs could process them, and the binding constraint was memory rather than compute time. Analysing a few million cells meant holding them in RAM, which put atlas-scale reanalysis behind a hardware budget most groups do not have. Scarf keeps data on disk and processes it in chunks instead of loading everything into memory, so a laptop can do work that previously needed a cluster. It handles scRNA-seq, scATAC-seq and CITE-seq, interoperates with the rest of the Python single-cell stack, and includes a subsampling method that preserves rare populations and differentiation trajectories rather than sampling them away. Scarf is the engine underneath ScarfWeb.
Abstract
As the scale of single-cell genomics experiments grows into the millions, the computational requirements to process this data are beyond the reach of many. Herein we present Scarf, a modularly designed Python package that seamlessly interoperates with other single-cell toolkits and allows for memory-efficient single-cell analysis of millions of cells on a laptop or low-cost devices like single-board computers. We demonstrate Scarf’s memory and compute-time efficiency by applying it to the largest existing single-cell RNA-Seq and ATAC-Seq datasets. Scarf wraps memory-efficient implementations of a graph-based t-stochastic neighbour embedding and hierarchical clustering algorithm. Moreover, Scarf performs accurate reference-anchored mapping of datasets while maintaining memory efficiency. By implementing a subsampling algorithm, Scarf additionally has the capacity to generate representative sampling of cells from a given dataset wherein rare cell populations and lineage differentiation trajectories are conserved. Together, Scarf provides a framework wherein any researcher can perform advanced processing, subsampling, reanalysis, and integration of atlas-scale datasets on standard laptop computers. Scarf is available on Github: https://github.com/parashardhapola/scarf .
Abstract licence: CC BY 4.0