<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>large-scale experimental datasets in microbiology &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/large-scale-experimental-datasets-in-microbiology/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Thu, 08 Oct 2026 19:58:48 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.3</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>large-scale experimental datasets in microbiology &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>Machine learning predicts which phages kill which bacteria from genomes alone</title>
		<link>https://scienmag.com/machine-learning-predicts-which-phages-kill-which-bacteria-from-genomes-alone/</link>
		
		<dc:creator><![CDATA[Teresa Odom]]></dc:creator>
		<pubDate>Thu, 08 Oct 2026 19:58:48 +0000</pubDate>
				<category><![CDATA[Biology]]></category>
		<category><![CDATA[antibiotic resistance solutions]]></category>
		<category><![CDATA[Antimicrobial Resistance]]></category>
		<category><![CDATA[bacterial genome analysis]]></category>
		<category><![CDATA[bacteriophages]]></category>
		<category><![CDATA[CatBoost]]></category>
		<category><![CDATA[computational models for bacteriophage targeting]]></category>
		<category><![CDATA[genome-based infection prediction]]></category>
		<category><![CDATA[genomic data analysis in microbiology]]></category>
		<category><![CDATA[genomics]]></category>
		<category><![CDATA[large-scale experimental datasets in microbiology]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[Machine learning in phage therapy]]></category>
		<category><![CDATA[microbial genomics and machine learning]]></category>
		<category><![CDATA[microbiome engineering]]></category>
		<category><![CDATA[phage cocktails]]></category>
		<category><![CDATA[phage specificity and host range]]></category>
		<category><![CDATA[phage therapy]]></category>
		<category><![CDATA[phage therapy development]]></category>
		<category><![CDATA[phage-bacteria interaction datasets]]></category>
		<category><![CDATA[phage-host interaction prediction]]></category>
		<category><![CDATA[phage-host interactions]]></category>
		<category><![CDATA[protein families]]></category>
		<category><![CDATA[RB-TnSeq]]></category>
		<category><![CDATA[strain-level prediction]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=248969</guid>

					<description><![CDATA[Researchers have built a phylogeny-agnostic machine-learning framework that predicts strain-level phage–host interactions from genome sequences alone and validated it experimentally across millions of training runs.]]></description>
										<content:encoded><![CDATA[<p>Bacteriophages, the viruses that infect bacteria, are increasingly viewed as one of the most promising answers to the global crisis of antibiotic resistance. Yet their greatest strength—the exquisite specificity with which they attack individual bacterial strains—is also their greatest practical weakness. A phage that lyses one clinical isolate may be harmless to a genetically near-identical neighbor, and with phage banks now holding thousands of candidates, experimentally testing every phage against every target pathogen has become impractical. A new study published in Nature Microbiology presents a machine-learning framework that tackles this bottleneck head-on, predicting strain-level phage–host interactions directly from genome sequences without requiring any prior knowledge of bacterial phylogeny or the molecular mechanisms of infection.</p>
<p>The research, led by Avery J. C. Noonan and colleagues at Lawrence Berkeley National Laboratory and collaborating institutions, was built on an unusually rigorous foundation. The team systematically optimized their workflow over more than 13.2 million individual training runs, evaluating over 200 parameter settings across six published phage–host interaction datasets. Together these datasets encompass 115,037 experimentally determined interactions spanning 949 bacterial strains and 518 phages, drawn from Escherichia coli, Klebsiella, Pseudomonas and Vibrionaceae hosts. The proportion of infectious interactions within each dataset ranged from as little as 2.1 percent to 36.2 percent, a severe imbalance that mirrors the reality of phage biology, where most phages fail to infect most strains they encounter.</p>
<p>The core technical challenge the researchers faced is one familiar to anyone working at the intersection of genomics and machine learning: dimensionality. Each bacterial or phage genome can contribute thousands of numerical features—protein families, gene fragments or sequence k-mers—while the phenotypic data describing which phage infects which strain is sparse, binary and heavily skewed toward negative outcomes. This combination invites overfitting, in which a model memorizes quirks of its training data rather than learning generalizable rules. Many earlier tools sidestepped the problem by restricting features to genes already known to mediate infection, such as phage tail fibres or host capsule-synthesis loci. That strategy works within a single, well-studied genus but cannot transfer to emerging pathogens or poorly characterized phages, and it forecloses the discovery of novel infection determinants.</p>
<p>The new framework, by contrast, is deliberately phylogeny-agnostic. Genomes are represented as binary presence–absence profiles of protein families, clustered with MMSeqs2 at optimized similarity thresholds. Informative features are then selected through recursive feature elimination, and predictions are generated by an ensemble of gradient-boosted decision trees using the CatBoost algorithm, with class weights computed separately for each phage to compensate for imbalance. The team compared eight machine-learning algorithms, six feature-selection methods and multiple genome-representation schemes, finding that algorithmic choices mattered more than how genomes were encoded. Notably, the simplest representation—short 3-amino-acid k-mers—performed worst, while protein families, longer k-mers and hybrid schemes captured sufficient biological signal.</p>
<p>Performance was assessed through nested cross-validation under three increasingly demanding scenarios: predicting infections of unseen bacterial strains by known phages, infections of known strains by unseen phages, and interactions between entirely novel phage–strain pairs. Across datasets, the area under the receiver operating characteristic curve ranged from 0.67 to 0.94, with Matthews correlation coefficients between 0.13 and 0.54. For E. coli, the model achieved an AUROC of 0.87, essentially matching the 0.86 reported by a previous species-specific method that required phage-specific models and host features based on known E. coli infection mediators. The new approach reached comparable accuracy without any of those constraints. Performance correlated strongly with dataset size, and the weakest results came from the smallest, most imbalanced dataset—a Klebsiella matrix of only 3,658 interactions with 4.9 percent positives—suggesting the limits reflect data availability rather than flaws in the method.</p>
<p>The framework also translated its probability scores into practical phage-cocktail design. Phages were clustered into &#8216;activity groups&#8217; using HDBSCAN density-based clustering, and the top-ranked phage from each cluster was selected, balancing predicted efficacy against mechanistic diversity. In silico tests across twenty-fold cross-validation showed that model-guided selection outperformed the common heuristic of choosing the most promiscuous phages. For single-phage selection, the models identified active phages for 66.9 percent of E. coli strains, 66.7 percent of Klebsiella strains and 39.4 percent of Vibrionaceae strains—improvements of 1.1-fold, 3.1-fold and 2.7-fold over promiscuity-based choices. Five-phage cocktails achieved bacterial coverage of up to 97.5 percent for E. coli, 87.8 percent for Klebsiella and 57.5 percent for Vibrionaceae, consistently beating baseline strategies.</p>
<p>Crucially, the team did not stop at computational validation. Because the phages in their training data were not available for retesting, they experimentally challenged the model with 52 previously unseen phages from the BASEL collection, spotted against 25 E. coli strains from the ECOR reference collection. Of the 1,300 tested interactions, 1,240 yielded clear phenotypes, and the model&#8217;s predictions against this genuinely novel matrix achieved an AUROC of 0.84—only marginally below its cross-validation performance of 0.87. This demonstrated that the framework generalizes to phages it had never encountered and to interaction data generated in a different laboratory with different protocols, a stringent test that many published prediction tools have never undergone.</p>
<p>To verify that the model&#8217;s predictions rest on real biology rather than statistical artifacts, the researchers turned to genome-wide random barcode transposon-site sequencing, or RB-TnSeq. They screened a pooled library of 3,804 gene knockouts in E. coli ECOR27 against 19 phages, identifying 51 genes whose disruption altered susceptibility. When these genetically validated infection mediators were compared with the model&#8217;s top predictive features, ranked by SHAP values, 68.6 percent showed a connection: 12 mapped directly to predictive features, 10 lay within a three-gene window of one, and 13 were functionally linked through the STRING protein-association database. Direct matches were enriched 5.0-fold above random expectation. Interestingly, the model often picked out a single regulatory gene within a larger biosynthetic pathway—porin regulators such as tsx, fadL and nfrA, or individual genes in lipopolysaccharide and capsule clusters—suggesting it captures pathway-level regulation rather than merely listing every component.</p>
<p>Feature interpretation also surfaced both familiar and unexpected biology. Among the top 25 predictive features in the E. coli dataset were elements tied to cell wall biosynthesis, mobile genetic elements, restriction–modification systems and genes of viral origin, alongside eight features with no clear mechanistic link that may represent uncharacterized infection mediators. Network and synteny analyses pointed to candidate roles for putrescine catabolism and regulation of the known phage receptor fepA, hypotheses the authors flag for future experimental work. Defence systems generally carried negative impacts on infection probability, as expected, while anti-defence features trended positive.</p>
<p>The study is candid about its limitations. Cross-genus prediction remains poor, with leave-one-genus-out AUROC values hovering barely above random, indicating that interaction determinants are largely genus-specific or that protein-family features cannot capture conserved mechanisms across deep evolutionary distances. RB-TnSeq validation is restricted to non-essential genes with strong fitness effects, and the authors suggest CRISPR interference screens as a complement for essential genes. Performance also depends on large datasets of confirmed interactions, which are slow and expensive to produce. Nevertheless, the successful merging of Klebsiella datasets from different research groups points toward community-driven, standardized data generation as a viable path forward. As antimicrobial resistance escalates, the authors argue, this framework—released as the open-source GenoPHI package—offers a scalable computational platform for rational phage-therapy design and precision microbiome engineering across clinical, agricultural and industrial settings.</p>
<p><strong>Subject of Research:</strong> Machine-learning prediction of strain-level bacteriophage–host interactions from bacterial and phage genome sequences</p>
<p><strong>Article Title:</strong> Phylogeny-agnostic strain-level prediction of phage–host interactions from genomes using machine learning</p>
<p><strong>Article References:</strong> Noonan, A. J. C., Moriniere, L., Rivera-López, E. O., Patel, K., Pena, M., Svab, M., Kazakov, A., Deutschbauer, A., Dudley, E. G., Mutalik, V. K., &amp; Arkin, A. P. (2026). Phylogeny-agnostic strain-level prediction of phage–host interactions from genomes using machine learning. <em>Nature Microbiology</em>. <a href="https://doi.org/10.1038/s41564-026-02482-5" rel="noopener noreferrer">https://doi.org/10.1038/s41564-026-02482-5</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1038/s41564-026-02482-5" rel="noopener noreferrer">10.1038/s41564-026-02482-5</a></p>
<p><strong>Keywords:</strong> bacteriophages, phage therapy, machine learning, phage–host interactions, genomics, antimicrobial resistance, strain-level prediction, RB-TnSeq, phage cocktails, microbiome engineering, CatBoost, protein families</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">248969</post-id>	</item>
	</channel>
</rss>
