<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>Prokka &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/prokka/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Thu, 01 Oct 2026 14:51:16 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>Prokka &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>Massive benchmark of 156,000 genomes reveals which annotation tool to trust</title>
		<link>https://scienmag.com/massive-benchmark-of-156000-genomes-reveals-which-annotation-tool-to-trust/</link>
		
		<dc:creator><![CDATA[Juliet Wilcox]]></dc:creator>
		<pubDate>Thu, 01 Oct 2026 14:51:16 +0000</pubDate>
				<category><![CDATA[Biology]]></category>
		<category><![CDATA[archaea]]></category>
		<category><![CDATA[bacteria]]></category>
		<category><![CDATA[Bakta]]></category>
		<category><![CDATA[benchmarking]]></category>
		<category><![CDATA[bioinformatics]]></category>
		<category><![CDATA[computational biology]]></category>
		<category><![CDATA[EggNOG-mapper]]></category>
		<category><![CDATA[functional gene annotation in bacteria and archaea]]></category>
		<category><![CDATA[Gene Ontology]]></category>
		<category><![CDATA[genome annotation]]></category>
		<category><![CDATA[genome annotation tool performance]]></category>
		<category><![CDATA[large-scale genomic analysis]]></category>
		<category><![CDATA[metagenome-assembled genomes]]></category>
		<category><![CDATA[microbial gene annotation tools comparison]]></category>
		<category><![CDATA[microbial genomics research]]></category>
		<category><![CDATA[open-source genome annotation pipelines]]></category>
		<category><![CDATA[PGAP]]></category>
		<category><![CDATA[prokaryotic genome analysis for downstream applications]]></category>
		<category><![CDATA[prokaryotic genome annotation accuracy]]></category>
		<category><![CDATA[prokaryotic genome annotation benchmarking]]></category>
		<category><![CDATA[Prokka]]></category>
		<category><![CDATA[systematic benchmarking of genome annotation tools]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=223326</guid>

					<description><![CDATA[A landmark benchmark of four open-source genome annotation tools across 156,033 prokaryotic genomes shows that Bakta excels on high-quality bacterial genomes while PGAP performs best on archaea and messy assemblies.]]></description>
										<content:encoded><![CDATA[<p>Every time a microbiologist sequences a new bacterium or archaeon, the raw string of A, C, G and T letters must be translated into something biologically meaningful: where the genes start and stop, what proteins they encode, and what those proteins actually do. This process, called genome annotation, is the foundation on which nearly every downstream analysis in prokaryotic genomics rests. Yet despite its central importance, researchers have long lacked a rigorous, large-scale answer to a deceptively simple question: which annotation tool should they actually use? A new study published in Genome Biology by Mateusz Jundzill of Jena University Hospital and colleagues, including Martin Hölzer of the Robert Koch Institute and Serghei Mangul, now provides that answer with unprecedented scope, systematically benchmarking four of the most widely used open-source annotation pipelines across more than 156,000 prokaryotic genomes.</p>
<p>The four tools under scrutiny were Prokka, Bakta, EggNOG-mapper and PGAP, each representing a distinct philosophy of how to assign function to genes. Prokka, a long-standing workhorse of bacterial genomics, is prized for its speed and ease of use. Bakta, a more recent successor, extends the Prokka concept with expanded databases and richer functional annotation. EggNOG-mapper takes a different route, transferring functional information from precomputed orthology groups spanning thousands of organisms. PGAP, the NCBI&#8217;s Prokaryotic Genome Annotation Pipeline, is the pipeline behind the reference annotations deposited in public databases, giving it a special role in standardizing how prokaryotic genomes are described worldwide. Until now, no study had compared these tools head-to-head at anything approaching this scale.</p>
<p>The scale of the benchmark is what sets it apart. The team assembled a dataset of 156,033 genomes deliberately chosen to span the diversity and the messiness of real-world sequencing. At its core were Escherichia coli strains, which served as a baseline for performance because they are among the best-characterized organisms in biology, with densely curated gene catalogs against which predictions can be checked. Around that core, the researchers layered thousands of bacterial genomes from across the taxonomic spectrum, thousands of archaeal genomes representing a domain of life that is often underrepresented in benchmarking efforts, and, critically, categories of genomes that reflect the challenges of modern practice: artificially frameshifted genomes and metagenome-assembled genomes, known as MAGs, reconstructed from environmental sequencing data.</p>
<p>The inclusion of imperfect genomes is arguably the study&#8217;s most consequential design decision. In principle, annotation tools are developed and tuned on clean, complete, high-quality reference genomes. In practice, the genomes flowing through laboratories today are frequently fragmented assemblies, contaminated with foreign DNA, or stitched together computationally from short metagenomic reads, none of which ever saw a perfect reference. Frameshifts, in which small insertions or deletions disrupt the reading frame of a coding sequence, are a particular hazard, because a tool that fails to recognize a frameshifted gene may either miss the gene entirely or mislabel its function. By deliberately stress-testing the pipelines against these degraded inputs, the benchmark captures how the tools behave where it matters most: on the messy data that dominate real surveillance, clinical and environmental projects.</p>
<p>The headline result is that there is no single winner, and that the best choice depends on what kind of genome is on the bench. For high-quality bacterial genomes, Bakta emerged as the strongest performer, delivering the most reliable annotation when the input sequence was complete and well assembled. For archaeal genomes, however, the picture flipped: PGAP outperformed the alternatives, a finding with practical weight given that archaeal genomics has historically been underserved by tools optimized on bacteria. PGAP also proved the more robust option for challenging bacterial assemblies, including metagenome-assembled, fragmented and contaminated samples, suggesting that the NCBI pipeline&#8217;s underlying machinery handles imperfect input more gracefully than its competitors.</p>
<p>When the analysis turned to functional annotation, the assignment of Gene Ontology terms that describe what each gene product does, a further layer of nuance appeared. PGAP consistently provided broader term coverage, attaching Gene Ontology annotations to a wider share of features across the genome. EggNOG-mapper, by contrast, tended to assign more terms per annotated feature, painting a denser functional portrait of the genes it did annotate. These are not contradictory results but complementary ones, and they matter for anyone conducting enrichment analyses or building functional profiles from annotation output. A researcher whose analysis depends on maximizing the number of annotated genes may prefer broader coverage, while one seeking richer descriptions of individual proteins may favor the denser per-feature annotation that EggNOG-mapper provides.</p>
<p>The implications ripple outward through microbiology. Genome annotation is not a cosmetic step; it determines which genes are counted in resistance screens, which virulence factors are flagged in pathogen surveillance, which metabolic pathways are reconstructed from environmental MAGs, and which drug targets are proposed from newly sequenced isolates. If two laboratories annotate the same genome with different tools and reach different conclusions about its gene content, the reproducibility of the entire field is at stake. By quantifying exactly where the tools agree and diverge, and by tying those differences to genome quality, taxonomy and origin, the study gives researchers an evidence-based decision framework rather than a matter of habit or folklore. The authors also provide detailed supplementary data, including per-sample summaries for E. coli, archaeal, non-MAG bacterial and MAG bacterial groups, along with accession lists, so that other groups can apply the benchmark&#8217;s findings directly to their own data.</p>
<p>The study also carries a message for the developers of annotation software. Tools that excel on pristine reference genomes may quietly degrade on the fragmented and contaminated assemblies that increasingly dominate public archives, and a benchmark of this scale makes such weaknesses visible in a way that small-scale comparisons never could. The authors frame their work as informing future tool development, and the specific failure modes documented across frameshifted and metagenome-assembled genomes offer a concrete to-do list: better handling of disrupted reading frames, greater robustness to contamination, and improved performance on archaeal diversity are all areas where the benchmark identifies room for improvement. As sequencing costs continue to fall and MAG reconstruction becomes routine, the gap between the genomes tools were designed for and the genomes they actually receive will only widen unless development is guided by evidence of this kind.</p>
<p>Published open access in Genome Biology on 4 September 2026, the study arrives at a moment when the volume of prokaryotic sequencing data is growing explosively, driven by clinical surveillance, outbreak investigation, and large-scale environmental and host-associated microbiome projects. Every one of those projects begins with the same annotation step that this benchmark has now illuminated. For the working microbiologist, the practical guidance is straightforward: reach for Bakta when annotating high-quality bacterial genomes, trust PGAP for archaea and for assemblies that are fragmented, contaminated or metagenomic in origin, and weigh the trade-off between PGAP&#8217;s broader Gene Ontology coverage and EggNOG-mapper&#8217;s richer per-feature annotation when functional profiling is the goal. For the field as a whole, the study demonstrates that even the most established steps of the bioinformatics pipeline benefit from systematic, large-scale scrutiny, and it sets a template for how future annotation tools should be evaluated: not on a handful of model organisms, but across the full taxonomic and quality spectrum of the genomes scientists actually sequence.</p>
<p><strong>Subject of Research:</strong> Large-scale benchmarking of prokaryotic genome annotation tools across diverse bacterial and archaeal genomes</p>
<p><strong>Article Title:</strong> Large-scale benchmarking of prokaryotic annotation tools across thousands of species</p>
<p><strong>Article References:</strong> Jundzill, M., Hölzer, M., Mangul, S., Marquet, M., Ehricht, R., Lohde, M., Spott, R., Makarewicz, O., Pletz, M. W., &amp; Brandt, C. (2026). Large-scale benchmarking of prokaryotic annotation tools across thousands of species. <em>Genome Biology, 27</em>(1), Article 284. <a href="https://doi.org/10.1186/s13059-026-04262-0" rel="noopener noreferrer">https://doi.org/10.1186/s13059-026-04262-0</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1186/s13059-026-04262-0" rel="noopener noreferrer">10.1186/s13059-026-04262-0</a></p>
<p><strong>Keywords:</strong> genome annotation, bioinformatics, bacteria, archaea, benchmarking, Prokka, Bakta, PGAP, EggNOG-mapper, metagenome-assembled genomes, Gene Ontology, computational biology</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">223326</post-id>	</item>
	</channel>
</rss>
