<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>cell state and identity &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/cell-state-and-identity/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Sun, 04 Oct 2026 06:03:06 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>cell state and identity &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>New Topic Model Traces Cell Identity Back to Regulatory Paths in Single-Cell Data</title>
		<link>https://scienmag.com/new-topic-model-traces-cell-identity-back-to-regulatory-paths-in-single-cell-data/</link>
		
		<dc:creator><![CDATA[Juliet Wilcox]]></dc:creator>
		<pubDate>Sun, 04 Oct 2026 06:03:06 +0000</pubDate>
				<category><![CDATA[Biology]]></category>
		<category><![CDATA[benchmarking]]></category>
		<category><![CDATA[bioinformatics]]></category>
		<category><![CDATA[biological structure integration]]></category>
		<category><![CDATA[cell state and identity]]></category>
		<category><![CDATA[cell type annotation]]></category>
		<category><![CDATA[computational framework for cell identity]]></category>
		<category><![CDATA[gene activity and cell classification]]></category>
		<category><![CDATA[gene regulatory mechanisms]]></category>
		<category><![CDATA[gene regulatory networks]]></category>
		<category><![CDATA[interpretability]]></category>
		<category><![CDATA[Latent Dirichlet Allocation]]></category>
		<category><![CDATA[non-negative matrix factorization]]></category>
		<category><![CDATA[RegPathTopic model]]></category>
		<category><![CDATA[regulatory gene networks]]></category>
		<category><![CDATA[reproducibility]]></category>
		<category><![CDATA[single-cell data analysis]]></category>
		<category><![CDATA[Single-Cell RNA Sequencing]]></category>
		<category><![CDATA[topic modeling]]></category>
		<category><![CDATA[topic modeling in genomics]]></category>
		<category><![CDATA[traceable regulatory vocabularies]]></category>
		<category><![CDATA[transcription factors]]></category>
		<category><![CDATA[Transcriptomics]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=233690</guid>

					<description><![CDATA[Researchers have developed RegPathTopic, a deterministic topic-modeling framework that builds traceable regulatory vocabularies from transcription factor-target graphs and benchmarks them honestly against matched controls for single-cell annotation.]]></description>
										<content:encoded><![CDATA[<p>Single-cell RNA sequencing has transformed biology by allowing researchers to measure gene activity in thousands of individual cells at once, but a persistent problem has shadowed the field: the labels that annotation tools assign to cells often say little about the regulatory mechanisms that actually define those cells. A newly published study in BMC Genomics introduces RegPathTopic, a computational framework designed to close this gap by building deterministic, traceable regulatory vocabularies from gene regulatory networks and using them in topic models for cell-type annotation. The work, led by Saidi Wang and colleagues at Henan University and the Henan Academy of Agricultural Sciences, was published on 7 September 2026 as an open-access research article.</p>
<p>The core insight behind RegPathTopic is that the vocabulary a topic model uses matters as much as the model itself. Topic models such as latent Dirichlet allocation and non-negative matrix factorization have long been applied to single-cell data, where they treat cell types as topics and genes as words. When suitable references are available, such approaches can reach high label accuracy, yet the assigned labels do not necessarily expose the regulatory features associated with cell identity or state. Graph-informed topic models have tried to inject biological structure by drawing on transcription factor-target networks, but they have faced a separate reproducibility problem: when regulatory vocabulary terms are generated stochastically, results can shift from run to run, undermining confidence in the outputs.</p>
<p>RegPathTopic tackles this reproducibility problem head-on by constructing vocabulary terms deterministically. The method converts transcription factor-target graphs into regulatory paths and transcription factor-centered modules, which serve as the vocabulary items that the topic model can draw upon. Rather than sampling these terms randomly, the framework scores and filters them according to explicit criteria, then calibrates the retained terms against degree-matched controls. This calibration step is critical: it ensures that any apparent advantage of the regulatory vocabulary is not simply an artifact of network degree, a common confounder in network-based analyses where highly connected genes can dominate results for reasons unrelated to biology.</p>
<p>Once the vocabulary is built and calibrated, RegPathTopic learns topic representations from three possible input matrices: a gene-only matrix, a regulatory-only matrix, or a combined matrix that merges both views. This flexibility allows researchers to compare how much signal the regulatory vocabulary contributes relative to conventional gene-level features. The topic representations can then be used for cell-type annotation in the standard way, but with a crucial difference: because the vocabulary terms are explicit paths and modules derived from known regulatory graphs, the resulting topics remain interpretable in terms of regulatory structure rather than being opaque lists of genes.</p>
<p>The authors evaluated their framework across seven benchmark settings, and the headline numbers are impressive on their face. In the filtered CellTypist internal benchmark, RegPathTopic achieved a macro-F1 score of 0.975, matching the best gene-only topic baseline. In CellTypist transfer and scIB transfer tasks, it reached macro-F1 values of 0.970 and 0.830 respectively, with positive margins over the displayed predecessor-style controls and over degree-matched graph-derived controls. These benchmarks, drawn from widely used resources including peripheral blood mononuclear cell datasets, represent some of the most demanding tests available for annotation methods, where models must generalize from reference data to new cells and new datasets.</p>
<p>What distinguishes this study from much of the machine-learning-for-biology literature, however, is its unusual candor about the limits of those results. The authors report that the positive margins did not extend to all comparator families. Gene-only transfer, several established annotation tools, and decoupler TRRUST activity baselines were stronger in relevant settings. In other words, RegPathTopic is not being presented as a superior annotation method or as a replacement for dedicated regulator-activity approaches such as AUCell-based scoring of regulator activity. The paper explicitly states that the evidence supports RegPathTopic as a reproducible framework for constructing and auditing graph-derived topic vocabularies and testing them against matched graph controls, and nothing more ambitious than that.</p>
<p>This careful framing extends to the biological interpretation of the vocabulary terms themselves. Vocabulary audits showed that retained paths and modules remained explicit and traceable through topic-level outputs, meaning that a researcher can follow a thread from a topic back to the specific regulatory paths that contributed to it. But the authors are emphatic that these analyses establish feature traceability rather than validated transcription factor activity or causal regulatory interpretation. A traced path in a topic is a hypothesis about regulation, not a demonstrated one. Perturbation-supported or orthogonal biological validation, such as experiments that directly perturb the regulators in question, would be required before the traced vocabulary terms could be interpreted as validated regulatory programs.</p>
<p>The methodological discipline on display here reflects a broader tension in computational biology. Interpretability is one of the most sought-after properties in machine learning applied to genomics, because biologists want models that not only classify cells correctly but also explain why. Yet interpretability claims are easy to overstate: a model that produces human-readable features may still be capturing statistical artifacts rather than biological mechanisms. By calibrating its regulatory vocabulary against degree-matched controls and reporting honestly where simpler or established methods outperform it, RegPathTopic offers a template for how graph-derived features should be evaluated before being celebrated as biologically meaningful.</p>
<p>The reproducibility angle is equally significant. Stochastic vocabulary generation in graph-informed topic models means that two researchers running the same pipeline on the same data could obtain different regulatory terms and therefore different topics, making results difficult to compare, audit, or build upon. Deterministic construction eliminates this source of variance. Combined with the explicit scoring, filtering, and calibration steps, it means that every vocabulary term in a RegPathTopic analysis has a documented provenance: which graph it came from, how it was scored, why it was retained, and how it compares to matched controls. For a field increasingly concerned with the reliability of computational pipelines, that kind of auditability is a substantive contribution in its own right.</p>
<p>Looking forward, the framework opens several avenues. Because the vocabulary is modular, future work could incorporate richer regulatory graphs as they become available, and the matched-control evaluation paradigm could be applied to other graph-derived feature constructions beyond topic modeling. The authors&#8217; funding acknowledgments, including support from the National Natural Science Foundation of China and several Henan provincial programs, indicate institutional investment in this line of research. For now, the study stands as a measured but meaningful advance: a demonstration that regulatory vocabularies can be made deterministic, traceable, and honestly benchmarked, even when the authors themselves insist that the hardest biological questions, the ones about causal regulation, remain open and await experimental validation.</p>
<p><strong>Subject of Research:</strong> Deterministic regulatory-path topic modeling for interpretable single-cell RNA-seq cell-type annotation</p>
<p><strong>Article Title:</strong> RegPathTopic: deterministic regulatory-path topic modeling with traceable regulatory vocabularies for single-cell annotation</p>
<p><strong>Article References:</strong> Wang, S., Sun, Z., Gao, J., &amp; Jiao, D. (2026). RegPathTopic: deterministic regulatory-path topic modeling with traceable regulatory vocabularies for single-cell annotation. <em>BMC Genomics</em>. <a href="https://doi.org/10.1186/s12864-026-13323-4" rel="noopener noreferrer">https://doi.org/10.1186/s12864-026-13323-4</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1186/s12864-026-13323-4" rel="noopener noreferrer">10.1186/s12864-026-13323-4</a></p>
<p><strong>Keywords:</strong> single-cell RNA sequencing, cell-type annotation, gene regulatory networks, topic modeling, transcription factors, latent Dirichlet allocation, non-negative matrix factorization, interpretability, reproducibility, benchmarking, transcriptomics, Bioinformatics</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">233690</post-id>	</item>
	</channel>
</rss>
