<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>seqlets &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/seqlets/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Thu, 08 Oct 2026 23:28:11 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.3</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>seqlets &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>Tangermeme toolkit turns black-box genomic AI into interpretable biology</title>
		<link>https://scienmag.com/tangermeme-toolkit-turns-black-box-genomic-ai-into-interpretable-biology/</link>
		
		<dc:creator><![CDATA[Juliet Wilcox]]></dc:creator>
		<pubDate>Thu, 08 Oct 2026 23:28:11 +0000</pubDate>
				<category><![CDATA[Biology]]></category>
		<category><![CDATA[AI model explanation in genomics]]></category>
		<category><![CDATA[bioinformatics Python packages]]></category>
		<category><![CDATA[cis-regulatory logic]]></category>
		<category><![CDATA[deep learning]]></category>
		<category><![CDATA[deep learning in chromatin accessibility]]></category>
		<category><![CDATA[DNA design]]></category>
		<category><![CDATA[feature attribution]]></category>
		<category><![CDATA[genomic deep learning interpretability]]></category>
		<category><![CDATA[genomics]]></category>
		<category><![CDATA[interpretability]]></category>
		<category><![CDATA[interpretable AI for genomic data]]></category>
		<category><![CDATA[Nature Methods]]></category>
		<category><![CDATA[neural network models for DNA sequence analysis]]></category>
		<category><![CDATA[open-source]]></category>
		<category><![CDATA[open-source bioinformatics tools]]></category>
		<category><![CDATA[regulatory grammar extraction in genomics]]></category>
		<category><![CDATA[seqlets]]></category>
		<category><![CDATA[single-cell and spatial genomics prediction]]></category>
		<category><![CDATA[software toolkit]]></category>
		<category><![CDATA[TANGERMEME toolkit for biological data]]></category>
		<category><![CDATA[transcription factor binding]]></category>
		<category><![CDATA[transparent neural networks in biology]]></category>
		<category><![CDATA[understanding transcription factor binding models]]></category>
		<category><![CDATA[variant effect prediction]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=250425</guid>

					<description><![CDATA[A new open-source software package provides fast, modular tools for extracting interpretable cis-regulatory patterns from deep learning models of the genome.]]></description>
										<content:encoded><![CDATA[<p>Deep learning has transformed genomics over the past decade. Neural networks trained directly on raw DNA sequence can now predict where transcription factors bind, where histones carry particular chemical marks, which stretches of chromatin lie open and accessible, how the genome folds in three dimensions, whether transcripts are spliced one way or another, and even how quickly messenger RNA degrades inside the cell. The most sophisticated of these models operate at single-base-pair resolution and, increasingly, at single-cell or spatial resolution. Yet for all this predictive firepower, a persistent frustration has shadowed the field: the models are spectacular at telling us what happens, and maddeningly opaque about why. A network that predicts MYC binding with exquisite accuracy does not automatically hand over the regulatory grammar it has memorized. Extracting that grammar has required bespoke, hand-rolled analysis code that every laboratory writes, debugs and maintains on its own.</p>
<p>A new study published in Nature Methods by Jacob Schreiber, of the Research Institute of Molecular Pathology at the Vienna BioCenter and UMass Chan Medical School, aims to close that gap. The paper introduces tangermeme, a free, open-source Python package described as an everything-but-the-model toolkit for genomic deep learning. The name of the design philosophy is deliberate. Rather than competing with the many frameworks that build, train and host neural networks, tangermeme concentrates exclusively on what happens after a model has been trained: making predictions efficiently, perturbing sequences, attributing importance to individual nucleotides, calling regulatory patterns and designing new DNA. The result, according to the paper, is a comprehensive, flexible and efficient Swiss Army knife for cis-regulatory pattern identification and analysis.</p>
<p>The rationale for this division of labor is grounded in how the field has actually evolved. Model architectures and training strategies churn constantly, with convolutions giving way to transformers and new optimizers arriving every year. Downstream analyses, by contrast, are remarkably stable and largely agnostic to architecture. Estimating the effect of a noncoding variant on a model&#8217;s prediction works essentially the same way whether the underlying network uses convolutional filters or self-attention, and whether it was trained with one optimizer or another. This modularity means that a well-engineered library of analysis operations can serve any model, old or new. Yet until now, no optimized repository of such methods existed that generalized across model types. Instead, researchers typically ship bespoke analysis code bundled with each released model, and existing packages either implement only a few related algorithms or focus on training and fine-tuning with limited post-training support.</p>
<p>Tangermeme&#8217;s central technical trick is the strict separation of sequence manipulations from model operations, which can then be freely stacked to create analyses. In silico marginalization compares predictions before and after substituting a short sequence into a template, quickly identifying which motifs from a database the model responds to. Ablation is the conceptual opposite, altering a region the model may be using and measuring the drop in prediction. Variant effect estimation compares predictions before and after one or a few noncontiguous substitutions, allowing fine-mapping of candidate regulatory variants. Crucially, the operation performed after a sequence manipulation need not be a simple prediction; it can be DeepLIFT/SHAP attribution, in silico saturation mutagenesis, or any custom operation a researcher writes. This composability lets scientists ask questions such as how surrounding sequence context influences transcription factor binding, whether a regulatory element is suitable for a synthetic construct, or how two variants at the same locus interact, all in a handful of commands.</p>
<p>The package also embraces DNA design, implementing several methods catalogued in a recent review of enhancer modeling. The simplest is screening: generate random sequences, score them against a user-defined objective, and keep the best. Greedy substitution, sometimes called directed evolution, iteratively improves a sequence by testing every possible single-nucleotide change and keeping the best one at each step. Motif implantation works similarly but swaps in entire motifs from a database rather than single bases. Most novel is a construct marginalization method that merges design with marginalization, creating short DNA inserts that reliably shift model predictions, such as patterns of predicted chromatin accessibility, averaged across many background sequences, without requiring prior knowledge of what the construct should contain. Such libraries of designed inserts could prove valuable for synthetic biology and therapeutic applications.</p>
<p>Engineering choices under the hood are where tangermeme distinguishes itself from prior efforts. Operations come with built-in batching, support every data type and device available in PyTorch, and work out of the box on models with multiple inputs and outputs, arbitrary internal units including convolutions, long short-term memory blocks and transformers, and any output format from a single binary value to a full base-pair-resolution profile. Speed benchmarks illustrate the payoff: tangermeme one-hot encodes the entirety of human chromosome 1 in under two seconds, roughly three times faster than the next fastest implementation tested. That may sound trivial, but fast encoding enables on-the-fly batch generation during analysis, and the same optimization philosophy runs throughout the codebase. Batching also solves a practical bottleneck in DeepLIFT/SHAP, which in many implementations requires all background sequences to fit in GPU memory simultaneously, a serious constraint for massive models.</p>
<p>Perhaps the most consequential finding to emerge from building the toolkit carefully is a subtle failure mode in widely used attribution code. DeepLIFT/SHAP works by overriding the backward pass of nonlinear operations to incorporate a reference sequence, and identifying which operations are nonlinear requires checking each computational layer against an internal dictionary. Existing implementations can silently fail when they encounter a nonlinear operation outside that dictionary or when an activation is reused, often without documentation or warnings. The telltale sign is convergence deltas, which should be zero in theory and near machine precision in practice, ballooning to indicate that the attributions are simply wrong. Tangermeme never fails silently: it raises warnings when convergence deltas are too high, can return those deltas for monitoring, and allows users to register custom nonlinear functions to avoid the problem entirely.</p>
<p>Beyond wrapping established algorithms, tangermeme introduces genuinely new methods for distilling learned cis-regulatory logic, centered on objects called seqlets, the contiguous spans of high-attribution nucleotides that attribution methods such as DeepLIFT/SHAP, in silico saturation mutagenesis, PISA or saliency produce. The package&#8217;s novel recursive seqlet caller is, according to the paper, the first to call variable-length seqlets directly, using a principled statistical definition: a seqlet is a span whose attribution sum is statistically significant against a null distribution, with the requirement that all subspans within it also pass the test. The algorithm bins attribution values, builds null distributions for each candidate length, converts them to P values, and decodes the longest non-overlapping significant spans. Compared against a reimplementation of TF-MoDISco&#8217;s caller on synthetic sequences with an inserted MYC motif, the recursive approach produced shorter, more precise calls and ran faster on modest-to-large numbers of sequences, while TF-MoDISco&#8217;s wider, more permissive calls reflect design choices made for its original role inside a larger pipeline.</p>
<p>The payoff of this machinery is demonstrated on real models. Applying the automatic seqlet-calling and annotation pipeline, which maps called seqlets to motifs from the JASPAR database and then counts motif occurrences, pairwise co-occurrences and spacing relationships, Schreiber compared two models ostensibly trained to predict the same thing: MYC binding. At the PLD6 promoter, attribution analysis revealed that BPNet&#8217;s predictions are driven almost exclusively by a MYC motif, whereas the much larger Beluga model relies on a broader lexicon of motifs also associated with chromatin accessibility and transcription initiation. Counting annotated seqlets across all MYC peaks and their pairwise occurrences made these differences in learned regulatory logic immediately visible, despite both models achieving similar predictive goals.</p>
<p>Tangermeme is released under the MIT license, installable with a single pip command and accompanied by extensive documentation and tutorials that reproduce the paper&#8217;s figures. Although the demonstrations focus on transcription factor binding and chromatin accessibility, the authors emphasize that the toolkit applies to models of any genomic modality, including alternative splicing, transcription, enhancer activity, RNA stability and RNA binding. Schreiber also points toward the future: as large language model-based coding agents become common in research, well-documented packages with tested, composable building blocks will let such systems reuse reliable code rather than reinventing analyses, reducing the burden of auditing generated code. Future development will focus on expanding functionality and deepening documentation for exactly that purpose, positioning tangermeme as infrastructure for an era in which interpreting the genome&#8217;s regulatory code may depend as much on trustworthy software as on the models themselves.</p>
<p><strong>Subject of Research:</strong> A deep learning interpretation toolkit for cis-regulatory genomics</p>
<p><strong>Article Title:</strong> Tangermeme: a toolkit for understanding cis-regulatory logic using deep learning models</p>
<p><strong>Article References:</strong> Schreiber, J. (2026). Tangermeme: a toolkit for understanding cis-regulatory logic using deep learning models. <em>Nature Methods</em>. <a href="https://doi.org/10.1038/s41592-026-03254-z" rel="noopener noreferrer">https://doi.org/10.1038/s41592-026-03254-z</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1038/s41592-026-03254-z" rel="noopener noreferrer">10.1038/s41592-026-03254-z</a></p>
<p><strong>Keywords:</strong> deep learning, genomics, cis-regulatory logic, software toolkit, feature attribution, seqlets, transcription factor binding, variant effect prediction, DNA design, interpretability, Nature Methods, open source</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">250425</post-id>	</item>
	</channel>
</rss>
