<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>perplexity &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/perplexity/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Fri, 02 Oct 2026 04:46:50 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>perplexity &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>Borrowing From Information Theory, Scientists Count the True Diversity of Human Isoforms</title>
		<link>https://scienmag.com/borrowing-from-information-theory-scientists-count-the-true-diversity-of-human-isoforms/</link>
		
		<dc:creator><![CDATA[Juliet Wilcox]]></dc:creator>
		<pubDate>Fri, 02 Oct 2026 04:46:50 +0000</pubDate>
				<category><![CDATA[Biology]]></category>
		<category><![CDATA[alternative splicing]]></category>
		<category><![CDATA[biological diversity metrics]]></category>
		<category><![CDATA[ENCODE4]]></category>
		<category><![CDATA[Genome Biology]]></category>
		<category><![CDATA[genome biology research]]></category>
		<category><![CDATA[genomics data analysis]]></category>
		<category><![CDATA[Hill numbers]]></category>
		<category><![CDATA[Human isoform diversity]]></category>
		<category><![CDATA[information theory in genomics]]></category>
		<category><![CDATA[isoform detection challenges]]></category>
		<category><![CDATA[isoform diversity]]></category>
		<category><![CDATA[long-read RNA sequencing]]></category>
		<category><![CDATA[measuring gene transcript diversity]]></category>
		<category><![CDATA[open reading frames]]></category>
		<category><![CDATA[PacBio]]></category>
		<category><![CDATA[PacBio sequencing technology]]></category>
		<category><![CDATA[perplexity]]></category>
		<category><![CDATA[RNA isoform cataloging]]></category>
		<category><![CDATA[Shannon entropy]]></category>
		<category><![CDATA[SQANTI3]]></category>
		<category><![CDATA[transcript structure analysis]]></category>
		<category><![CDATA[transcriptome]]></category>
		<category><![CDATA[transcriptome complexity]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=225758</guid>

					<description><![CDATA[Researchers have introduced perplexity, an information-theoretic measure of effective isoform number, as a threshold-free way to quantify transcript diversity across 55 human cell types.]]></description>
										<content:encoded><![CDATA[<p>Every biology student learns the central dogma: one gene, one protein. Modern genomics has long since demolished that tidy picture. A single human gene can be sliced and reassembled into dozens of distinct messenger RNA molecules, known as isoforms, each of which may encode a different protein or perform a different regulatory role. Long-read RNA sequencing technologies, such as those developed by PacBio, now make it possible to read these RNA molecules from end to end, revealing an astonishing menagerie of transcript structures that shorter-read methods simply cannot resolve. Yet a stubborn problem has persisted: once researchers have catalogued all these isoforms, how do they actually measure how diverse a gene&#8217;s transcript output really is?</p>
<p>A team led by Megan D. Schertzer and Stella H. Park of the New York Genome Center and Columbia University, together with Jiayu Su, Fairlie Reese, Gloria M. Sheynkman and senior author David A. Knowles, believes the field has been answering that question in the wrong way. In a study published in Genome Biology, the researchers propose replacing the crude filtering steps that dominate current pipelines with a single, principled number borrowed from information theory: perplexity. Their central claim is provocative. Instead of discarding rare isoforms as noise, every transcript a gene produces, no matter how scarce, should contribute proportionally to that gene&#8217;s measured diversity.</p>
<p>To understand why this matters, it helps to look at how isoform analysis currently works. Long-read sequencing experiments inevitably generate some technical artifacts, spurious transcript models that do not correspond to real RNA molecules. Pipelines such as SQANTI3, which the new study uses to classify transcript structures into quality categories, help researchers separate genuine isoforms from these errors. But after artifact removal, most workflows then apply an expression threshold, keeping only isoforms that exceed some minimum abundance, typically measured in transcripts per million. Those thresholds are, as the authors point out, arbitrary. An isoform sitting just below the cutoff is thrown away, even though it may be a bona fide transcript with a real biological function. The result is a systematic undercounting of diversity that varies from study to study depending on where each team happens to set its cutoff, undermining reproducibility across labs and datasets.</p>
<p>Perplexity offers a way out of this arbitrariness. The concept originates in Shannon entropy, the foundational measure of uncertainty in information theory. If you treat the isoforms of a gene as the possible outcomes of a random draw, weighted by their relative abundances, the Shannon entropy quantifies how uncertain you are about which isoform you will draw next. Exponentiating that entropy yields perplexity, which has an elegant and intuitive interpretation: it is the effective number of isoforms a gene produces. A gene that expresses one dominant transcript with only trace alternatives has a perplexity close to one. A gene that spreads its expression evenly across ten isoforms has a perplexity approaching ten. The metric is closely related to the Hill numbers that ecologists have long used to measure biodiversity, a theoretical connection the authors develop in their supplementary materials, drawing on the framework formalized by Tom Leinster.</p>
<p>The beauty of this formulation is that no isoform is ever discarded. A transcript representing one percent of a gene&#8217;s output contributes to the perplexity calculation in proportion to its abundance, exactly as its share of the expression distribution dictates. Rare isoforms nudge the effective count upward; abundant ones dominate it. The measurement is continuous, interpretable and independent of any researcher-chosen threshold. In effect, perplexity converts the messy, long-tailed distribution of isoform abundances that long-read sequencers reveal into a single number that means the same thing in every dataset, every tissue and every laboratory.</p>
<p>To demonstrate that the metric delivers on this promise, the team applied it to a substantial trove of public data: 124 PacBio long-read RNA sequencing datasets from the ENCODE4 project, spanning 55 human cell types. This resource, generated by the Encyclopedia of DNA Elements consortium, captures transcriptomes across a remarkable range of biological contexts, from immune cells to neuronal lineages. After processing the data through their custom pipeline, which includes artifact filtering with SQANTI3 and annotation of open reading frames, the authors computed perplexity for every gene in every sample and asked whether the resulting diversity measurements were interpretable and reproducible.</p>
<p>The answer, according to the study, is yes. Perplexity proved to be a stable and meaningful quantity across genes, regulatory levels and tissues. Genes whose isoform repertoires are governed by well-characterized regulatory mechanisms showed diversity profiles consistent with known biology, while the metric distinguished cell types in ways that reflect their underlying splicing and transcription programs. Because the calculation uses all detected isoforms rather than a thresholded subset, repeated measurements of the same biological condition converge far more reliably than threshold-based counts, addressing one of the most persistent complaints about long-read transcriptomics: that different pipelines, and even slightly different parameters within a single pipeline, can yield noticeably different catalogs of isoforms.</p>
<p>The work also extends beyond simply counting transcripts. The authors analyzed their detected isoforms at the level of open reading frames, the stretches of RNA that are actually translated into protein, and their supplementary analyses include quadrant classifications of open reading frames and tissue-specific assignments. This reflects a growing recognition in the field that isoform diversity matters not merely as a cataloging exercise but because different isoforms can carry different protein-coding potential, potentially producing distinct protein products from the same gene. A diversity metric that faithfully captures the full isoform landscape therefore has direct implications for proteomics and for interpreting the functional consequences of genetic variation.</p>
<p>The implications reach into medicine as well. Alternative splicing is increasingly implicated in cancer, neurological disease and genetic disorders, and dysregulated isoform production is a hallmark of many tumors. Gloria Sheynkman, whose laboratory at the University of Virginia focuses on connecting transcriptomics to proteomics, and David Knowles, whose group at Columbia and the New York Genome Center develops computational methods for genome interpretation, have both founded or advised companies in this space, underscoring the translational interest in reliable isoform quantification. A reproducible, threshold-free measure of isoform diversity could improve the statistical power of studies that compare diseased and healthy tissue, where subtle shifts in splicing regimes may be drowned out by the noise introduced by arbitrary filtering.</p>
<p>There is also a satisfying intellectual symmetry in the approach. Ecologists counting species in a rainforest, linguists measuring the richness of a vocabulary and now genomicists quantifying the diversity of a transcriptome are all grappling with the same mathematical structure: a distribution over many categories, where rare categories are real but hard to handle. Shannon entropy and its exponential, perplexity, provide a common language for all of these problems. By importing that language into transcriptomics, Schertzer, Park and their colleagues have turned a long-standing methodological weakness into an opportunity, giving the field a way to embrace the full, messy richness of the human transcriptome rather than trimming it to fit a threshold. As long-read sequencing continues its march toward routine clinical and research use, metrics like perplexity may well become as standard a vocabulary for RNA biology as expression levels are today.</p>
<p><strong>Subject of Research:</strong> An information-theoretic metric for quantifying isoform diversity in the human transcriptome</p>
<p><strong>Article Title:</strong> Perplexity as a metric for isoform diversity in the human transcriptome</p>
<p><strong>Article References:</strong> Schertzer, M. D., Park, S. H., Su, J., Reese, F., Sheynkman, G. M., &amp; Knowles, D. A. (2026). Perplexity as a metric for isoform diversity in the human transcriptome. <em>Genome Biology</em>. <a href="https://doi.org/10.1186/s13059-026-04260-2" rel="noopener noreferrer">https://doi.org/10.1186/s13059-026-04260-2</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1186/s13059-026-04260-2" rel="noopener noreferrer">10.1186/s13059-026-04260-2</a></p>
<p><strong>Keywords:</strong> perplexity, isoform diversity, long-read RNA sequencing, Shannon entropy, Hill numbers, transcriptome, ENCODE4, PacBio, alternative splicing, SQANTI3, open reading frames, Genome Biology</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">225758</post-id>	</item>
	</channel>
</rss>
