<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>protein databases &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/protein-databases/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Fri, 09 Oct 2026 17:39:58 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.3</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>protein databases &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>Wild-AC: A Faster Way to Match Peptides to Proteins, Wildcards Included</title>
		<link>https://scienmag.com/wild-ac-a-faster-way-to-match-peptides-to-proteins-wildcards-included/</link>
		
		<dc:creator><![CDATA[Drew Townsend]]></dc:creator>
		<pubDate>Fri, 09 Oct 2026 17:39:58 +0000</pubDate>
				<category><![CDATA[Biology]]></category>
		<category><![CDATA[Aho-Corasick]]></category>
		<category><![CDATA[Aho-Corasick algorithm in bioinformatics]]></category>
		<category><![CDATA[algorithms]]></category>
		<category><![CDATA[bioinformatics]]></category>
		<category><![CDATA[bioinformatics algorithm development]]></category>
		<category><![CDATA[computational challenges in proteomics]]></category>
		<category><![CDATA[efficient pattern matching algorithms]]></category>
		<category><![CDATA[FM-index]]></category>
		<category><![CDATA[high-throughput proteomics data processing]]></category>
		<category><![CDATA[mass spectrometry data analysis]]></category>
		<category><![CDATA[open-source]]></category>
		<category><![CDATA[peptide fingerprinting techniques]]></category>
		<category><![CDATA[peptide fragment identification]]></category>
		<category><![CDATA[peptide mapping]]></category>
		<category><![CDATA[protein database search optimization]]></category>
		<category><![CDATA[protein databases]]></category>
		<category><![CDATA[protein sequence ambiguity resolution]]></category>
		<category><![CDATA[Proteomics]]></category>
		<category><![CDATA[proteomics peptide-to-protein matching]]></category>
		<category><![CDATA[string matching]]></category>
		<category><![CDATA[Wild-AC]]></category>
		<category><![CDATA[wildcard amino acid sequence search]]></category>
		<category><![CDATA[wildcards]]></category>
		<category><![CDATA[Wu-Manber]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=255153</guid>

					<description><![CDATA[A new open-source algorithm called Wild-AC extends the classic Aho-Corasick string matcher to handle ambiguous amino acids in protein databases, outperforming FM-index and Wu-Manber on realistic proteomics search workloads.]]></description>
										<content:encoded><![CDATA[<p>Every mass spectrometry experiment in modern proteomics ends with the same deceptively simple question: which proteins do these measured peptide fragments come from? Answering it means searching thousands, sometimes hundreds of thousands, of short amino acid sequences against enormous protein databases that are riddled with ambiguous characters. Those ambiguity codes, the wildcards of the protein world, have long been a computational headache, forcing researchers to choose between slow exhaustive searches and index structures that must be painstakingly precomputed. A new algorithm called Wild-AC, described in BMC Bioinformatics by Chris Bielow of Freie Universität Berlin, now promises to dissolve that dilemma with a clever twist on a fifty-year-old classic.</p>
<p>The classic in question is the Aho-Corasick algorithm, a cornerstone of computer science first published in 1975. Its genius lies in building a finite-state automaton from an entire set of search patterns at once, so that a single left-to-right pass over the text can locate every occurrence of every pattern simultaneously. For exact matching, it remains remarkably hard to beat: no preprocessing of the text is required, the automaton is built once from the patterns, and the scan proceeds at a pace independent of how many patterns are in play. That property, known as index-free operation, is precisely what makes it attractive for proteomics workflows, where the pattern set, the peptides, changes with every experiment while the text, the protein database, may be reused.</p>
<p>The trouble begins with wildcards. Protein databases routinely contain the ambiguity codes B, Z and X, which stand for &#8216;asparagine or aspartic acid&#8217;, &#8216;glutamine or glutamic acid&#8217; and &#8216;any amino acid&#8217;, respectively. When a peptide pattern collides with one of these characters in the text, a standard Aho-Corasick automaton simply has no transition to follow, and the match is lost. Existing solutions have been unsatisfying in different ways. FM-indexes, the compressed full-text indexes popularized by genomics, can handle wildcards but demand substantial preprocessing of the text and become awkward when the wildcard sits in the searched text rather than in the pattern. Other multi-pattern methods, such as the widely used Wu-Manber algorithm, handle exact matching quickly but offer no native wildcard semantics at all.</p>
<p>Wild-AC&#8217;s central idea is disarmingly elegant: when the scan encounters a wildcard character in the text, it does not halt or fall back to a slow secondary routine. Instead, it branches the search into parallel &#8216;scout&#8217; paths, one for each possible interpretation of the ambiguous residue, while the unmodified primary search continues unaffected through the common, wildcard-free case. Each scout carries its own state within the automaton and explores the subtree of possibilities that the wildcard opens up. Because wildcards are relatively rare, typically a few percent of database positions after standard masking, the scouts remain short-lived excursions rather than an exponential explosion. The primary automaton, meanwhile, behaves exactly like a textbook Aho-Corasick machine, so the overwhelming majority of matching work pays no wildcard tax whatsoever.</p>
<p>The engineering details matter as much as the concept. Wild-AC is implemented in C++ and released as open source, with support for multi-threading so that large peptide sets can be partitioned across cores. Because no index of the text is built, the algorithm starts working immediately on any database in plain sequence format, a meaningful advantage in laboratory pipelines where databases are swapped frequently, for example when searching against species-specific proteomes or contamination databases. The memory footprint stays modest, and the automaton construction scales gracefully with the number of patterns, which in practice ranges from a thousand peptides in a targeted experiment to half a million in a deep discovery run.</p>
<p>The benchmark results are where the algorithm earns its headline. Bielow tested Wild-AC across realistic proteomics workloads: pattern sets of 1,000 to 500,000 peptides with an average length of roughly 18 amino acids, searched against protein databases ranging from 3 million to 210 million characters. In the wildcard case, Wild-AC outperformed the FM-index outright. In pure exact matching, it matched or exceeded Wu-Manber, which is notable because Wu-Manber was designed specifically for fast multi-pattern exact search and has few peers. The advantage held across the entire realistic range of pattern counts, suggesting the algorithm is not a niche performer tuned to one benchmark configuration but a genuine workhorse.</p>
<p>One nuance deserves attention. When the databases were artificially masked to a wildcard rate of 5 percent, a scenario representing heavily annotated or deliberately ambiguous reference proteomes, the picture became more conditional. Wild-AC retained its advantage for large peptide sets, but the FM-index became the preferable choice for small ones. This makes intuitive sense: an index amortizes its construction cost over many queries, so if you plan to run only a handful of small searches against the same text, paying once for a prebuilt index can win. For the typical proteomics pattern of use, many peptides per search and databases that change often, Wild-AC&#8217;s index-free design tips the balance decisively the other way.</p>
<p>Why does this matter beyond the benchmark tables? Peptide-to-protein mapping sits at the foundation of nearly every downstream analysis in shotgun proteomics: protein identification, quantification, quality control and the detection of sequence variants all depend on it being fast and correct. Ambiguous amino acids are not an exotic edge case. They appear wherever sequences are incompletely characterized, in genomes assembled from noisy data, in databases that merge paralogous proteins, and in deliberately degenerate searches for modified or variant peptides. An algorithm that treats wildcards as a first-class citizen, without imposing index construction or sacrificing exact-search speed, removes a persistent friction from those pipelines. The author credits discussions with Sandro Andreotti on the algorithmic design and implementation input from Simon Gene Gottlieb, whose FM-index codebase and fuzzy amino acid matching implementation served as the benchmark comparison, with reviewer feedback from Ragnar Groot Koerkamp among others helping sharpen the final manuscript.</p>
<p>There is also a broader lesson for computational biology in how Wild-AC was built. Rather than inventing an entirely new data structure, the work takes a mature, well-understood algorithm and extends it minimally to cover a missing capability. The scout-path mechanism preserves the automaton&#8217;s behavior in the common case and pays for wildcard handling only where it is actually needed. This kind of surgical extension, validating performance against the strongest existing tools rather than straw men, is exactly the pattern that tends to produce software that gets adopted. The open-source C++ implementation, hosted publicly on GitHub, lowers the barrier for integration into existing search engines and quality-control tools, where Aho-Corasick machinery is often already present in some form.</p>
<p>For the proteomics community, the immediate takeaway is practical: if your pipeline maps large peptide sets against protein databases containing ambiguity codes, Wild-AC now offers the fastest known route, with no index to maintain and threads to spare. For small searches against a fixed, heavily masked database, the FM-index remains a sensible choice. For everyone else, the paper is a reminder that some of the most impactful bioinformatics advances still come from revisiting the classics of stringology and asking what happens when the text, not just the pattern, refuses to be unambiguous. As protein databases continue to swell with sequences of uneven certainty, algorithms that embrace that uncertainty at full speed will only grow in importance.</p>
<p><strong>Subject of Research:</strong> A wildcard-enabled multi-pattern string matching algorithm for peptide-to-protein mapping in proteomics</p>
<p><strong>Article Title:</strong> Wild-AC: A fast, index-free multi-pattern string matching algorithm with wildcard support for proteomics</p>
<p><strong>Article References:</strong> Bielow, C. (2026). Wild-AC: A fast, index-free multi-pattern string matching algorithm with wildcard support for proteomics. <em>BMC Bioinformatics</em>. <a href="https://doi.org/10.1186/s12859-026-06686-8" rel="noopener noreferrer">https://doi.org/10.1186/s12859-026-06686-8</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1186/s12859-026-06686-8" rel="noopener noreferrer">10.1186/s12859-026-06686-8</a></p>
<p><strong>Keywords:</strong> Wild-AC, Aho-Corasick, string matching, wildcards, proteomics, peptide mapping, protein databases, FM-index, Wu-Manber, algorithms, bioinformatics, open source</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">255153</post-id>	</item>
		<item>
		<title>New SPAID Database Maps Hidden Autoantigens Behind Autoimmune Diseases</title>
		<link>https://scienmag.com/new-spaid-database-maps-hidden-autoantigens-behind-autoimmune-diseases/</link>
		
		<dc:creator><![CDATA[Drew Townsend]]></dc:creator>
		<pubDate>Mon, 21 Sep 2026 00:03:51 +0000</pubDate>
				<category><![CDATA[Biology]]></category>
		<category><![CDATA[autoantigens]]></category>
		<category><![CDATA[autoantigens in autoimmune diseases]]></category>
		<category><![CDATA[autoimmune disease biomarkers]]></category>
		<category><![CDATA[autoimmune disease diagnostics]]></category>
		<category><![CDATA[autoimmune diseases]]></category>
		<category><![CDATA[bioinformatics in immunology]]></category>
		<category><![CDATA[biomarker discovery]]></category>
		<category><![CDATA[comprehensive autoantigen mapping]]></category>
		<category><![CDATA[epitopes]]></category>
		<category><![CDATA[HLA class I]]></category>
		<category><![CDATA[immune epitope validation]]></category>
		<category><![CDATA[Immunogenicity prediction]]></category>
		<category><![CDATA[mass spectrometry]]></category>
		<category><![CDATA[mass spectrometry in autoantigen discovery]]></category>
		<category><![CDATA[non-canonical proteins]]></category>
		<category><![CDATA[non-canonical proteins in autoimmunity]]></category>
		<category><![CDATA[non-coding genome translation]]></category>
		<category><![CDATA[novel autoantigen identification]]></category>
		<category><![CDATA[protein databases]]></category>
		<category><![CDATA[Proteomics]]></category>
		<category><![CDATA[rheumatoid arthritis]]></category>
		<category><![CDATA[SPAID database]]></category>
		<category><![CDATA[T-cell and MHC ligand assays]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=204364</guid>

					<description><![CDATA[Researchers have launched SPAID, a comprehensive database that maps both canonical and non-canonical candidate autoantigens across 14 autoimmune diseases by integrating validated epitope evidence with large-scale proteomic analysis.]]></description>
										<content:encoded><![CDATA[<p>Autoimmune diseases, in which the immune system turns against the body&#8217;s own tissues, affect hundreds of millions of people worldwide and remain notoriously difficult to diagnose early and precisely. At the heart of every autoimmune response lies a molecular trigger: an autoantigen, a self-protein or peptide that the immune system mistakenly recognizes as foreign. Yet despite decades of research, the full landscape of these triggers remains incomplete, in part because scientists have traditionally focused only on canonical, well-annotated protein-coding genes. Now, a research team led by scientists at Sun Yat-sen University and collaborating institutions in China has unveiled SPAID, a comprehensive database designed to systematically catalog candidate autoantigens across 14 autoimmune disorders, including both canonical proteins and a vast, largely unexplored universe of non-canonical proteins translated from non-coding regions of the genome.</p>
<p>SPAID, which is freely accessible online at spaid.renlab.cn, organizes its evidence into two distinct levels. The first, a validated level, contains proteins carrying experimentally confirmed epitopes drawn from the Immune Epitope Database, supported by positive T-cell assays and major histocompatibility complex (MHC) ligand assays. The second, a proteomics-based level, aggregates disease-associated peptides and proteins identified through mass spectrometry from human patient samples, annotated with differential expression patterns, predicted immunogenicity scores, and functional features. This two-tier architecture allows researchers to distinguish between candidates backed by direct immunological experimentation and those flagged through high-throughput proteomic discovery that await laboratory validation.</p>
<p>The technical ambition behind SPAID is considerable. To capture non-canonical proteins, the team assembled candidate sequences from more than 660,000 non-coding RNA entries in RNAcentral and over 332,000 intronic sequences from the IntroVerse database. Each candidate was evaluated for coding potential using two independent algorithms, CPAT and CNCI, and only sequences passing both thresholds were retained. Open reading frames were then predicted with NCBI&#8217;s ORFfinder and translated into amino acid sequences, which were de-duplicated against the UniProt reference set. The result is a unified protein sequence space of 576,516 sequences, combining 42,444 canonical UniProt proteins with 534,072 non-canonical proteins, including 447,445 intron-derived and 86,627 ncRNA-derived candidates.</p>
<p>Onto this reference framework, the researchers mapped a wealth of experimental data. From 292 publications, they integrated T-cell and MHC ligand assay records, ultimately identifying 1,141 unique validated epitopes from 10 autoimmune diseases supported by 2,750 positive T-cell assay records, alongside 20,424 distinct epitopes from 21,566 positive MHC ligand assays across five diseases. In total, these experimentally supported epitopes mapped to 21,349 unique proteins, spanning 16,966 canonical, 1,681 intron-derived, and 2,702 ncRNA-derived proteins. The inclusion of non-canonical proteins at this level is particularly striking, as it suggests that proteins translated from non-coding RNAs and introns can serve as genuine immune targets in human autoimmunity.</p>
<p>The proteomics-based level is equally extensive. Drawing on public repositories including PRIDE, MassIVE.quant, jPOST, PeptideAtlas, and iProX, the team curated 675 human proteomic samples spanning 14 autoimmune diseases, stratified into 51 disease- and tissue-specific cohorts. Peptides were identified by searching tandem mass spectra against the unified sequence space using DIA-NN for data-independent acquisition datasets and MaxQuant for data-dependent acquisition, with stringent false discovery rate control of 1 percent at the peptide-spectrum match, peptide, and protein-group levels. This rigorous filtering was essential because non-canonical peptides carry a heightened risk of false-positive identification. The search yielded 176,363 mass spectrometry-identified peptides assigned to 26,085 disease-associated proteins, including 927 ncRNA-derived and 531 intron-derived proteins.</p>
<p>To transform raw protein identifications into biologically meaningful signals, SPAID performs differential expression analysis for each cohort, comparing disease samples against matched controls. Proteins were classified as disease-only detected, upregulated, downregulated, or other, and results across multiple cohorts were integrated using Robust Rank Aggregation to assess cross-study consistency. Across the 14 diseases, the platform identified 4,577 disease-only detected proteins, 2,571 significantly upregulated proteins, and 757 significantly downregulated proteins. Notably, the disease-only category included 193 ncRNA-derived and 28 intron-derived proteins, demonstrating that non-canonical translation products participate in disease-specific proteomic signatures rather than representing background noise.</p>
<p>One of the most consequential questions the team addressed was whether these non-canonical proteins are reproducible. In diseases supported by at least three independent proteomic cohorts, more than 40 percent of non-canonical proteins were repeatedly detected: 61.51 percent in psoriasis, 54.58 percent in systemic lupus erythematosus, 51.41 percent in Crohn&#8217;s disease, and 40.32 percent in rheumatoid arthritis. Even more striking, among the repeatedly detected proteins, expression patterns showed remarkable concordance, with 99.33 percent consistency in rheumatoid arthritis, 93.93 percent in lupus, and 87.72 percent in psoriasis. Cross-referencing with published literature revealed that only 1.10 percent of the 1,458 non-canonical proteins identified across the diseases had prior experimental support, meaning the overwhelming majority represent previously unrecognized translation products now documented at scale for the first time.</p>
<p>To pinpoint which of these proteins might actually provoke immune responses, SPAID incorporates an immunogenicity prediction pipeline. Every mass spectrometry-detected peptide was segmented into overlapping 8- to 14-mer fragments and evaluated for HLA class I presentation across 12 functional supertypes, integrating MHC binding affinity, peptide-MHC stability, and T-cell recognition probability. Candidates were then refined using PanPep, a machine-learning tool that estimates T-cell receptor interaction probabilities against a panel of 419 CDR3 sequences. Overall, 14.70 percent of the 176,363 detected peptides were classified as putatively immunogenic, mapping to 15,558 immunogenic proteins, or 59.64 percent of all disease-associated proteins in the database. Immunogenic candidates were strongly enriched among disease-only detected proteins, representing 69.24 percent of that subset, consistent with the idea that proteins elevated under inflammatory conditions feed the antigen-processing machinery that can expose sequestered self-determinants and cryptic epitopes.</p>
<p>By intersecting three features, disease-specific proteomic detection, predicted immunogenicity, and experimental epitope support, the team defined a high-confidence core set of 2,023 candidate autoantigens. The platform&#8217;s practical utility was then demonstrated in an independent rheumatoid arthritis cohort, where serum proteomics of five patients and five healthy controls identified 1,559 proteins, 132 of which were RA-associated. A two-step validation pipeline using SPAID recovered clinically established biomarkers such as gamma-interferon-inducible protein 16 (IFI16) and immunoglobulin mu heavy chain, while also flagging novel candidates. Nine proteins, including myosin-9 (MYH9), glutathione S-transferase P (GSTP1), and hemoglobin subunit gamma-1/2 (HBG1/2), harbored experimentally validated epitopes. APOA4, apolipoprotein A-IV, emerged as an entirely novel candidate with highly specific enrichment in RA samples and strong predicted immunogenicity but no prior literature link to the disease, illustrating how the database can surface unexpected therapeutic leads.</p>
<p>Beyond its scientific content, SPAID offers a polished web interface built on a MySQL backend with a Java-based server and interactive ECharts visualizations. Users can search by disease, tissue, or protein attributes, run BLAST searches against transcript, protein, and peptide datasets, and explore hierarchical gene, protein, and peptide pages featuring expression boxplots, volcano plots, 3D structural models from the Protein Data Bank or ColabFold, predicted post-translational modification sites generated with PTM-Mamba, and an MS/MS spectrum annotator. The authors are candid about limitations: immunogenicity predictions currently cover only HLA class I, omitting CD4 T-cell biology tied to HLA class II, and the underlying proteomic data skew toward accessible tissues such as blood and skin. Mass spectrometry, however stringent, cannot on its own prove functional translation or physiological epitope presentation. Positioned as a candidate discovery resource rather than a definitive catalog, SPAID nonetheless represents a foundational shift in how autoantigen research can be conducted, and its developers plan future expansions to include HLA class II data and broader tissue proteomics, potentially accelerating diagnostics and targeted therapies for millions of autoimmune patients.</p>
<p><strong>Subject of Research:</strong> A comprehensive database for disease-specific autoantigen discovery in autoimmune disorders</p>
<p><strong>Article Title:</strong> SPAID: a comprehensive database for disease-specific autoantigens in autoimmune disorders</p>
<p><strong>Article References:</strong> Deng, S., Wei, F., Pang, Y., Zhang, L., Zhi, S., Chen, T., Zuo, Z., Ren, J., Xie, Y., &amp; Luo, X. (2026). SPAID: a comprehensive database for disease-specific autoantigens in autoimmune disorders. <em>Advanced Biotechnology, 4</em>(2), Article 23. <a href="https://doi.org/10.1007/s44307-026-00117-8" rel="noopener noreferrer">https://doi.org/10.1007/s44307-026-00117-8</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s44307-026-00117-8" rel="noopener noreferrer">10.1007/s44307-026-00117-8</a></p>
<p><strong>Keywords:</strong> autoimmune diseases, autoantigens, SPAID database, proteomics, non-canonical proteins, mass spectrometry, epitopes, HLA class I, immunogenicity prediction, rheumatoid arthritis, biomarker discovery, protein databases</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">204364</post-id>	</item>
	</channel>
</rss>
