Tuesday, September 22, 2026
Science
No Result
View All Result
  • Login
  • HOME
  • SCIENCE NEWS
  • CONTACT US
  • HOME
  • SCIENCE NEWS
  • CONTACT US
No Result
View All Result
Scienmag
No Result
View All Result
Home Science News Biology

ParTIpy Brings Pareto Task Inference to Million-Cell Single-Cell Datasets

September 22, 2026
in Biology
Drew Townsend
By Drew Townsend Scienmag Editorial Profile - Cell Biology
Reading Time: 5 mins read
0
ParTIpy Brings Pareto Task Inference to Million-Cell Single-Cell Datasets

ParTIpy Brings Pareto Task Inference to Million-Cell Single-Cell Datasets

ParTIpy Brings Pareto Task Inference to Million-Cell Single-Cell Datasets

65
SHARES
587
VIEWS
Share on FacebookShare on Twitter
ADVERTISEMENT

Biology is a discipline of compromise. A bird’s beak cannot be simultaneously perfect for cracking seeds and snatching insects, and a single cell cannot execute every function written into its genome at once. Instead, living systems allocate finite resources across competing tasks, producing the specialization and division of labor that scientists observe at every scale, from entire organisms down to individual molecules. A research team led by Philipp Sven Lars Schäfer, Ricardo O Ramirez Flores, and Julio Saez-Rodriguez at Heidelberg University, together with collaborators at the Hebrew University of Jerusalem and EMBL-EBI, has now unveiled an open-source software package called ParTIpy that brings a powerful mathematical framework for studying these trade-offs into the era of million-cell single-cell datasets. Published in Molecular Systems Biology, the work promises to reshape how biologists interpret the continuous variability hidden within seemingly uniform cell populations.

The intellectual foundation of ParTIpy is Pareto Task Inference, or ParTI, a framework grounded in the theory of multi-objective optimality. The core idea is elegant: when evolution drives phenotypes toward Pareto optimality, any improvement in performance on one task necessarily degrades performance on another. Under this assumption, phenotypes such as organisms, proteins, or cells, described by traits that embody trade-offs, are predicted to occupy a simple geometric structure called a polytope, a generalized shape with corners, straight edges, and flat faces, such as a triangle or tetrahedron extended to any number of dimensions. Crucially, the number of vertices of this polytope equals the number of tasks the system is optimizing. Each vertex, known as an archetype, represents a phenotype perfectly specialized for one particular task.

To recover these archetypes from data, the framework relies on archetypal analysis, a dimensionality reduction technique first introduced by Cutler and Breiman in 1994. The method identifies extreme points within a dataset, the archetypes, and then models every observation as a nonnegative linear combination of them. Formally, the archetypes are defined as convex combinations of the data points, ensuring they lie within the convex hull of the observed distribution, and each data point is reconstructed as a convex combination of the archetypes. When applied to single-cell transcriptomics, this means fitting a polytope around cells in gene expression space, where each vertex corresponds to a transcriptional program optimized for a distinct biological task. Positioning a cell relative to the vertices then reveals its bias toward one or several functions.

Previous studies have demonstrated the value of this perspective, revealing trade-offs in hepatocytes, fibroblasts, and cancer cells, and even extending the approach to spatial transcriptomics and proteomics, where groups of co-localized cells cooperate on shared tasks. However, the original ParTI software, written in MATLAB, could not scale to the massive datasets generated by modern high-throughput single-cell technologies, and its dependence on a commercial license hindered broader adoption. ParTIpy addresses both limitations directly, implementing state-of-the-art initialization and optimization strategies in Python and integrating seamlessly with the scverse ecosystem of single-cell analysis tools, including scanpy and AnnData objects.

A central technical challenge in archetypal analysis is that the underlying optimization problem is non-convex, meaning solution quality is highly sensitive to how the algorithm is initialized. The team benchmarked three initialization algorithms and two optimization procedures across twenty-one single-cell transcriptome datasets, each representing a single cell type and drawn from three independent studies spanning multiple sclerosis lesions, spatial Xenium data, and lupus patient blood. They found that combining principal convex hull analysis, or PCHA, for optimization with the Archetypal++, or AA++, initialization scheme delivered the best trade-off between reconstruction accuracy and runtime. When speed is the priority, a more efficient Frank-Wolfe variant provides a practical alternative. The package also supports a convexity relaxation parameter that allows archetypes to extend beyond the observed data cloud, a crucial option when sampling near the boundaries of the true phenotypic space is sparse.

The most striking computational innovation, however, is ParTIpy’s use of coresets. A coreset is a small, cleverly weighted subset of the data such that optimizing on the subset yields results close to those obtained from the full dataset. The researchers incorporated coreset weights directly into the objective function and modified their optimization algorithms accordingly. Their benchmarks showed that the required coreset fraction shrinks as datasets grow: for datasets exceeding one hundred thousand cells, analyzing just one to ten percent of the data was sufficient on average, cutting runtime roughly fourfold. When fitting six archetypes on one million cells using a ten percent coreset, the optimization completed in forty-five to one hundred five seconds. Critically, archetype-specific gene expression profiles derived from coresets remained highly concordant with full-data results, with a median Pearson correlation above 0.9 across coreset fractions greater than one percent.

Because the correct number of archetypes is never known in advance, ParTIpy implements three complementary selection criteria. The first examines the fraction of variance explained as archetypes are added, looking for an elbow where returns diminish. The second employs an information-theoretic criterion that penalizes model complexity, favoring solutions that balance goodness of fit against the number of free parameters. The third assesses robustness through bootstrapping: archetypal analysis is repeated on multiple resampled datasets, archetype positions are aligned across replicates using Hungarian matching, and the positional variance of each archetype is computed. Low variance signals stable, well-defined archetypes, whereas high variance warns of over-parameterization. The team validated that ParTIpy’s results are highly concordant with the original MATLAB implementation across three distinct single-cell datasets.

To demonstrate the framework in action, the researchers applied ParTIpy to single-cell RNA sequencing data of hepatocytes sampled along the liver lobule’s porto-central axis. Three selection criteria converged on a four-archetype model, and pathway enrichment analysis using Reactome gene sets via the decoupler-py univariate linear model revealed distinct specializations: one archetype was associated with drug and xenobiotic metabolism, another with lipid metabolism. When the archetypal programs were projected onto spatial transcriptomics data, the metabolism archetype localized near the central vein while the lipid archetype enriched near the portal node, matching established liver zonation patterns and suggesting that spatial gradients of oxygen and metabolites drive this functional division of labor. The package further supports inference of archetype crosstalk networks by integrating archetype-specific expression profiles with curated ligand-receptor databases.

The second demonstration tackled a longstanding weakness of conventional cross-condition single-cell analysis. Standard approaches discretize cells into clusters and test for changes in cluster abundance, imposing sharp boundaries on what are often intrinsically continuous cell states. ParTIpy instead models cells as continuous mixtures of archetypal programs. Analyzing roughly 147,000 fibroblasts from sixteen non-failing and twenty-six cardiomyopathy donor hearts, the team fit a three-archetype model chosen for superior bootstrap stability. Donor-level pseudobulk profiles revealed significant shifts toward an activated fibrotic archetype in disease, marked by genes such as POSTN, FN1, and COL1A1 and by transcription factors including TWIST1 and SCX, while a quiescent, stress-responsive archetype was closer to healthy donors. Notably, even every non-failing heart contained a small fibroblast population near the fibrotic archetype, indicating that cardiomyopathy redistributes cells among existing programs rather than creating disease-specific states.

The researchers emphasize that ParTIpy’s scalability, downstream utilities, and interoperability make Pareto task inference accessible to the broader single-cell community and beyond, with tutorials covering spatial transcriptomics applications. Unlike deep archetypal analysis approaches that rely on encoder-decoder networks and amortized variational inference, which can introduce an amortization gap, ParTIpy achieves scale through coreset-based direct optimization while preserving archetypes as extremal points in the original trait space. Future integration with RNA velocity, pseudotime, and optimal transport methods could distinguish stable task specializations from transient states. The package, version 0.2.0, is freely available on PyPI and GitHub, with comprehensive documentation at partipy.readthedocs.io, positioning ParTIpy as a versatile new lens for exploring functional trade-offs in high-dimensional biological data.

Subject of Research: Scalable archetypal analysis and Pareto task inference for single-cell transcriptomic data

Article Title: ParTIpy: a scalable framework for archetypal analysis and Pareto task inference

Article References: ParTIpy: a scalable framework for archetypal analysis and Pareto task inference. (n.d.). https://doi.org/10.1038/s44320-026-00209-6

Image Credits: AI Generated

DOI: 10.1038/s44320-026-00209-6

Keywords: ParTIpy, Pareto task inference, archetypal analysis, single-cell transcriptomics, coresets, trade-offs, cell state variability, hepatocyte zonation, cardiac fibroblasts, scverse, polytope geometry, open-source software

Cite Scienmag News

Drew Townsend. (September 22, 2026). ParTIpy Brings Pareto Task Inference to Million-Cell Single-Cell Datasets. Scienmag. https://scienmag.com/partipy-brings-pareto-task-inference-to-million-cell-single-cell-datasets/

Drew Townsend. "ParTIpy Brings Pareto Task Inference to Million-Cell Single-Cell Datasets." Scienmag, 22 September 2026, https://scienmag.com/partipy-brings-pareto-task-inference-to-million-cell-single-cell-datasets/. Accessed 22 September 2026.

Drew Townsend. "ParTIpy Brings Pareto Task Inference to Million-Cell Single-Cell Datasets." Scienmag. September 22, 2026. https://scienmag.com/partipy-brings-pareto-task-inference-to-million-cell-single-cell-datasets/

Tags: applications of Pareto theory in biologyarchetypal analysisbiological trade-offs in resource allocationcardiac fibroblastscell specialization and division of laborcell state variabilitycoresetshepatocyte zonationhigh-dimensional single-cell data interpretationmathematical frameworks for cell function analysismulti-objective optimality in cellular traitsmulti-task inference in systems biologyopen-source softwareopen-source software for single-cell dataPareto task inferencePareto Task Inference in biologyParTIpypolytope geometryscversesingle-cell dataset analysissingle-cell heterogeneity and variabilitysingle-cell transcriptomicstrade-offs
Share26Tweet16
Previous Post

Lung Disease Strikes During Ulcerative Colitis Remission, Challenging Drug Explanation

Next Post

Natural Fibers Step Up as Plastic-Free Champions for Sustainable Food Packaging

Related Posts

Losing Genes in Sequence Propelled the Global Rise of Monophasic Salmonella ST34
Biology

Losing Genes in Sequence Propelled the Global Rise of Monophasic Salmonella ST34

September 22, 2026
Korean Red Ginseng Shields the Gut From Air Pollution Damage, Mouse Study Finds
Biology

Korean Red Ginseng Shields the Gut From Air Pollution Damage, Mouse Study Finds

September 22, 2026
Pangenome-Powered Tool SVPG Sharpens Structural Variant Detection and Speeds Graph Updates
Biology

Pangenome-Powered Tool SVPG Sharpens Structural Variant Detection and Speeds Graph Updates

September 22, 2026
Genetic Map of DMD Variants Emerges from Türkiye’s Black Sea Region
Biology

Genetic Map of DMD Variants Emerges from Türkiye’s Black Sea Region

September 22, 2026
Twisted Peptide Shape Reveals How Plants Reopen Stomata After Immune Alarm
Biology

Twisted Peptide Shape Reveals How Plants Reopen Stomata After Immune Alarm

September 22, 2026
Drug Combinations Show Broad Synergy Against Alpha, Beta, and Gamma Herpesviruses
Biology

Drug Combinations Show Broad Synergy Against Alpha, Beta, and Gamma Herpesviruses

September 22, 2026
Next Post
Natural Fibers Step Up as Plastic-Free Champions for Sustainable Food Packaging

Natural Fibers Step Up as Plastic-Free Champions for Sustainable Food Packaging

  • Mothers who receive childcare support from maternal grandparents show more optimized

    Mothers who receive childcare support from maternal grandparents show more parental warmth, finds NTU Singapore study

    27656 shares
    Share 11059 Tweet 6912
  • University of Seville Breaks 120-Year-Old Mystery, Revises a Key Einstein Concept

    1061 shares
    Share 424 Tweet 265
  • Bee body mass, pathogens and local climate influence heat tolerance

    682 shares
    Share 273 Tweet 171
  • Researchers record first-ever images and data of a shark experiencing a boat strike

    546 shares
    Share 218 Tweet 137
  • Groundbreaking Clinical Trial Reveals Lubiprostone Enhances Kidney Function

    531 shares
    Share 212 Tweet 133
Science

Embark on a thrilling journey of discovery with Scienmag.com—your ultimate source for cutting-edge breakthroughs. Immerse yourself in a world where curiosity knows no limits and tomorrow’s possibilities become today’s reality!

RECENT NEWS

  • Diffusion May Drive Earthquakes With Slip That Grows With Distance
  • Natural Fibers Step Up as Plastic-Free Champions for Sustainable Food Packaging
  • ParTIpy Brings Pareto Task Inference to Million-Cell Single-Cell Datasets
  • Lung Disease Strikes During Ulcerative Colitis Remission, Challenging Drug Explanation

Categories

  • Agriculture
  • Anthropology
  • Archaeology
  • Athmospheric
  • Biology
  • Biotechnology
  • Blog
  • Bussines
  • Cancer
  • Chemistry
  • Climate
  • Earth Science
  • Editorial Policy
  • Marine
  • Mathematics
  • Medicine
  • Pediatry
  • Policy
  • Psychology & Psychiatry
  • Science Education
  • Social Science
  • Space
  • Technology and Engineering

Subscribe to Blog via Email

Enter your email address to subscribe to this blog and receive notifications of new posts by email.

Join 5,151 other subscribers

© 2025 Scienmag - Science Magazine

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • HOME
  • SCIENCE NEWS
  • CONTACT US

© 2025 Scienmag - Science Magazine

Discover more from Science

Subscribe now to keep reading and get access to the full archive.

Continue reading