Thursday, October 8, 2026
Science
No Result
View All Result
  • Login
  • HOME
  • SCIENCE NEWS
  • CONTACT US
  • HOME
  • SCIENCE NEWS
  • CONTACT US
No Result
View All Result
Scienmag
No Result
View All Result
Home Science News Biology

Tangermeme toolkit turns black-box genomic AI into interpretable biology

October 8, 2026
in Biology
Juliet Wilcox
By Juliet Wilcox Scienmag Editorial Profile - Human Genetics
Reading Time: 6 mins read
0
Tangermeme toolkit turns black-box genomic AI into interpretable biology

Tangermeme toolkit turns black-box genomic AI into interpretable biology

65
SHARES
587
VIEWS
Share on FacebookShare on Twitter
ADVERTISEMENT

Deep learning has transformed genomics over the past decade. Neural networks trained directly on raw DNA sequence can now predict where transcription factors bind, where histones carry particular chemical marks, which stretches of chromatin lie open and accessible, how the genome folds in three dimensions, whether transcripts are spliced one way or another, and even how quickly messenger RNA degrades inside the cell. The most sophisticated of these models operate at single-base-pair resolution and, increasingly, at single-cell or spatial resolution. Yet for all this predictive firepower, a persistent frustration has shadowed the field: the models are spectacular at telling us what happens, and maddeningly opaque about why. A network that predicts MYC binding with exquisite accuracy does not automatically hand over the regulatory grammar it has memorized. Extracting that grammar has required bespoke, hand-rolled analysis code that every laboratory writes, debugs and maintains on its own.

A new study published in Nature Methods by Jacob Schreiber, of the Research Institute of Molecular Pathology at the Vienna BioCenter and UMass Chan Medical School, aims to close that gap. The paper introduces tangermeme, a free, open-source Python package described as an everything-but-the-model toolkit for genomic deep learning. The name of the design philosophy is deliberate. Rather than competing with the many frameworks that build, train and host neural networks, tangermeme concentrates exclusively on what happens after a model has been trained: making predictions efficiently, perturbing sequences, attributing importance to individual nucleotides, calling regulatory patterns and designing new DNA. The result, according to the paper, is a comprehensive, flexible and efficient Swiss Army knife for cis-regulatory pattern identification and analysis.

The rationale for this division of labor is grounded in how the field has actually evolved. Model architectures and training strategies churn constantly, with convolutions giving way to transformers and new optimizers arriving every year. Downstream analyses, by contrast, are remarkably stable and largely agnostic to architecture. Estimating the effect of a noncoding variant on a model’s prediction works essentially the same way whether the underlying network uses convolutional filters or self-attention, and whether it was trained with one optimizer or another. This modularity means that a well-engineered library of analysis operations can serve any model, old or new. Yet until now, no optimized repository of such methods existed that generalized across model types. Instead, researchers typically ship bespoke analysis code bundled with each released model, and existing packages either implement only a few related algorithms or focus on training and fine-tuning with limited post-training support.

Tangermeme’s central technical trick is the strict separation of sequence manipulations from model operations, which can then be freely stacked to create analyses. In silico marginalization compares predictions before and after substituting a short sequence into a template, quickly identifying which motifs from a database the model responds to. Ablation is the conceptual opposite, altering a region the model may be using and measuring the drop in prediction. Variant effect estimation compares predictions before and after one or a few noncontiguous substitutions, allowing fine-mapping of candidate regulatory variants. Crucially, the operation performed after a sequence manipulation need not be a simple prediction; it can be DeepLIFT/SHAP attribution, in silico saturation mutagenesis, or any custom operation a researcher writes. This composability lets scientists ask questions such as how surrounding sequence context influences transcription factor binding, whether a regulatory element is suitable for a synthetic construct, or how two variants at the same locus interact, all in a handful of commands.

The package also embraces DNA design, implementing several methods catalogued in a recent review of enhancer modeling. The simplest is screening: generate random sequences, score them against a user-defined objective, and keep the best. Greedy substitution, sometimes called directed evolution, iteratively improves a sequence by testing every possible single-nucleotide change and keeping the best one at each step. Motif implantation works similarly but swaps in entire motifs from a database rather than single bases. Most novel is a construct marginalization method that merges design with marginalization, creating short DNA inserts that reliably shift model predictions, such as patterns of predicted chromatin accessibility, averaged across many background sequences, without requiring prior knowledge of what the construct should contain. Such libraries of designed inserts could prove valuable for synthetic biology and therapeutic applications.

Engineering choices under the hood are where tangermeme distinguishes itself from prior efforts. Operations come with built-in batching, support every data type and device available in PyTorch, and work out of the box on models with multiple inputs and outputs, arbitrary internal units including convolutions, long short-term memory blocks and transformers, and any output format from a single binary value to a full base-pair-resolution profile. Speed benchmarks illustrate the payoff: tangermeme one-hot encodes the entirety of human chromosome 1 in under two seconds, roughly three times faster than the next fastest implementation tested. That may sound trivial, but fast encoding enables on-the-fly batch generation during analysis, and the same optimization philosophy runs throughout the codebase. Batching also solves a practical bottleneck in DeepLIFT/SHAP, which in many implementations requires all background sequences to fit in GPU memory simultaneously, a serious constraint for massive models.

Perhaps the most consequential finding to emerge from building the toolkit carefully is a subtle failure mode in widely used attribution code. DeepLIFT/SHAP works by overriding the backward pass of nonlinear operations to incorporate a reference sequence, and identifying which operations are nonlinear requires checking each computational layer against an internal dictionary. Existing implementations can silently fail when they encounter a nonlinear operation outside that dictionary or when an activation is reused, often without documentation or warnings. The telltale sign is convergence deltas, which should be zero in theory and near machine precision in practice, ballooning to indicate that the attributions are simply wrong. Tangermeme never fails silently: it raises warnings when convergence deltas are too high, can return those deltas for monitoring, and allows users to register custom nonlinear functions to avoid the problem entirely.

Beyond wrapping established algorithms, tangermeme introduces genuinely new methods for distilling learned cis-regulatory logic, centered on objects called seqlets, the contiguous spans of high-attribution nucleotides that attribution methods such as DeepLIFT/SHAP, in silico saturation mutagenesis, PISA or saliency produce. The package’s novel recursive seqlet caller is, according to the paper, the first to call variable-length seqlets directly, using a principled statistical definition: a seqlet is a span whose attribution sum is statistically significant against a null distribution, with the requirement that all subspans within it also pass the test. The algorithm bins attribution values, builds null distributions for each candidate length, converts them to P values, and decodes the longest non-overlapping significant spans. Compared against a reimplementation of TF-MoDISco’s caller on synthetic sequences with an inserted MYC motif, the recursive approach produced shorter, more precise calls and ran faster on modest-to-large numbers of sequences, while TF-MoDISco’s wider, more permissive calls reflect design choices made for its original role inside a larger pipeline.

The payoff of this machinery is demonstrated on real models. Applying the automatic seqlet-calling and annotation pipeline, which maps called seqlets to motifs from the JASPAR database and then counts motif occurrences, pairwise co-occurrences and spacing relationships, Schreiber compared two models ostensibly trained to predict the same thing: MYC binding. At the PLD6 promoter, attribution analysis revealed that BPNet’s predictions are driven almost exclusively by a MYC motif, whereas the much larger Beluga model relies on a broader lexicon of motifs also associated with chromatin accessibility and transcription initiation. Counting annotated seqlets across all MYC peaks and their pairwise occurrences made these differences in learned regulatory logic immediately visible, despite both models achieving similar predictive goals.

Tangermeme is released under the MIT license, installable with a single pip command and accompanied by extensive documentation and tutorials that reproduce the paper’s figures. Although the demonstrations focus on transcription factor binding and chromatin accessibility, the authors emphasize that the toolkit applies to models of any genomic modality, including alternative splicing, transcription, enhancer activity, RNA stability and RNA binding. Schreiber also points toward the future: as large language model-based coding agents become common in research, well-documented packages with tested, composable building blocks will let such systems reuse reliable code rather than reinventing analyses, reducing the burden of auditing generated code. Future development will focus on expanding functionality and deepening documentation for exactly that purpose, positioning tangermeme as infrastructure for an era in which interpreting the genome’s regulatory code may depend as much on trustworthy software as on the models themselves.

Subject of Research: A deep learning interpretation toolkit for cis-regulatory genomics

Article Title: Tangermeme: a toolkit for understanding cis-regulatory logic using deep learning models

Article References: Schreiber, J. (2026). Tangermeme: a toolkit for understanding cis-regulatory logic using deep learning models. Nature Methods. https://doi.org/10.1038/s41592-026-03254-z

Image Credits: AI Generated

DOI: 10.1038/s41592-026-03254-z

Keywords: deep learning, genomics, cis-regulatory logic, software toolkit, feature attribution, seqlets, transcription factor binding, variant effect prediction, DNA design, interpretability, Nature Methods, open source

Cite Scienmag News

Juliet Wilcox. (October 8, 2026). Tangermeme toolkit turns black-box genomic AI into interpretable biology. Scienmag. https://scienmag.com/tangermeme-toolkit-turns-black-box-genomic-ai-into-interpretable-biology/

Juliet Wilcox. "Tangermeme toolkit turns black-box genomic AI into interpretable biology." Scienmag, 8 October 2026, https://scienmag.com/tangermeme-toolkit-turns-black-box-genomic-ai-into-interpretable-biology/. Accessed 8 October 2026.

Juliet Wilcox. "Tangermeme toolkit turns black-box genomic AI into interpretable biology." Scienmag. October 8, 2026. https://scienmag.com/tangermeme-toolkit-turns-black-box-genomic-ai-into-interpretable-biology/

Tags: AI model explanation in genomicsbioinformatics Python packagescis-regulatory logicdeep learningdeep learning in chromatin accessibilityDNA designfeature attributiongenomic deep learning interpretabilitygenomicsinterpretabilityinterpretable AI for genomic dataNature Methodsneural network models for DNA sequence analysisopen-sourceopen-source bioinformatics toolsregulatory grammar extraction in genomicsseqletssingle-cell and spatial genomics predictionsoftware toolkitTANGERMEME toolkit for biological datatranscription factor bindingtransparent neural networks in biologyunderstanding transcription factor binding modelsvariant effect prediction
Share26Tweet16
Previous Post

Years on the Ward Shape Why Nurses Go Back to School, Study Finds

Next Post

How Much Is Your House Worth to a Flood? New Study Puts a Price on Every Building in Germany

Related Posts

Mossy Wetlands Defy Standard Methods for Splitting Water Losses
Biology

Mossy Wetlands Defy Standard Methods for Splitting Water Losses

October 8, 2026
Blood Tests Reveal Hidden Burden of a Neglected Parasite Across Ethiopian Schools
Biology

Blood Tests Reveal Hidden Burden of a Neglected Parasite Across Ethiopian Schools

October 8, 2026
Soybean Roots Recruit Phosphate-Solubilizing Bacteria Through a Starvation Signal
Biology

Soybean Roots Recruit Phosphate-Solubilizing Bacteria Through a Starvation Signal

October 8, 2026
Anesthesia and Surgery Rewire the Alzheimer’s Brain Differently, Mouse Study Suggests
Biology

Anesthesia and Surgery Rewire the Alzheimer’s Brain Differently, Mouse Study Suggests

October 8, 2026
Malaria Parasites Must Burrow Through Cells to Spark Vaccine Protection, Study Finds
Biology

Malaria Parasites Must Burrow Through Cells to Spark Vaccine Protection, Study Finds

October 8, 2026
Oyster Parasite Reveals Ancient Blueprint of Divergent Mitochondrial Machinery
Biology

Oyster Parasite Reveals Ancient Blueprint of Divergent Mitochondrial Machinery

October 8, 2026
Next Post
How Much Is Your House Worth to a Flood? New Study Puts a Price on Every Building in Germany

How Much Is Your House Worth to a Flood? New Study Puts a Price on Every Building in Germany

  • Mothers who receive childcare support from maternal grandparents show more optimized

    Mothers who receive childcare support from maternal grandparents show more parental warmth, finds NTU Singapore study

    27656 shares
    Share 11059 Tweet 6912
  • University of Seville Breaks 120-Year-Old Mystery, Revises a Key Einstein Concept

    1061 shares
    Share 424 Tweet 265
  • Bee body mass, pathogens and local climate influence heat tolerance

    682 shares
    Share 273 Tweet 171
  • Researchers record first-ever images and data of a shark experiencing a boat strike

    546 shares
    Share 218 Tweet 137
  • Groundbreaking Clinical Trial Reveals Lubiprostone Enhances Kidney Function

    531 shares
    Share 212 Tweet 133
Science

Embark on a thrilling journey of discovery with Scienmag.com—your ultimate source for cutting-edge breakthroughs. Immerse yourself in a world where curiosity knows no limits and tomorrow’s possibilities become today’s reality!

RECENT NEWS

  • Tiny Sun-Watching CubeSat Proves It Can Also Track Earth’s Climate Energy Balance
  • How Much Is Your House Worth to a Flood? New Study Puts a Price on Every Building in Germany
  • Tangermeme toolkit turns black-box genomic AI into interpretable biology
  • Years on the Ward Shape Why Nurses Go Back to School, Study Finds

Categories

  • Agriculture
  • Anthropology
  • Archaeology
  • Athmospheric
  • Biology
  • Biotechnology
  • Blog
  • Bussines
  • Cancer
  • Chemistry
  • Climate
  • Earth Science
  • Editorial Policy
  • Marine
  • Mathematics
  • Medicine
  • Pediatry
  • Policy
  • Psychology & Psychiatry
  • Science Education
  • Science News
  • Social Science
  • Space
  • Technology and Engineering

Subscribe to Blog via Email

Enter your email address to subscribe to this blog and receive notifications of new posts by email.

Join 5,150 other subscribers

© 2025 Scienmag - Science Magazine

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • HOME
  • SCIENCE NEWS
  • CONTACT US

© 2025 Scienmag - Science Magazine

Discover more from Science

Subscribe now to keep reading and get access to the full archive.

Continue reading