<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>bioinformatics education &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/bioinformatics-education/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Sat, 03 Oct 2026 23:45:19 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>bioinformatics education &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>Large Language Models Are Learning to Read the Language of Life</title>
		<link>https://scienmag.com/large-language-models-are-learning-to-read-the-language-of-life/</link>
		
		<dc:creator><![CDATA[Courtney Benton]]></dc:creator>
		<pubDate>Sat, 03 Oct 2026 23:45:19 +0000</pubDate>
				<category><![CDATA[Social Science]]></category>
		<category><![CDATA[AI in genomics research]]></category>
		<category><![CDATA[AlphaFold]]></category>
		<category><![CDATA[bioinformatics and deep learning]]></category>
		<category><![CDATA[bioinformatics education]]></category>
		<category><![CDATA[deep learning]]></category>
		<category><![CDATA[digital transformation in molecular biology]]></category>
		<category><![CDATA[DNA and protein sequence interpretation]]></category>
		<category><![CDATA[DNA language models]]></category>
		<category><![CDATA[DNA sequence analysis]]></category>
		<category><![CDATA[foundation models]]></category>
		<category><![CDATA[genomic language models]]></category>
		<category><![CDATA[genomics]]></category>
		<category><![CDATA[large language models]]></category>
		<category><![CDATA[large language models in molecular biology]]></category>
		<category><![CDATA[machine learning for life sciences]]></category>
		<category><![CDATA[natural language processing for biological data]]></category>
		<category><![CDATA[protein structure prediction]]></category>
		<category><![CDATA[RNA and protein decoding]]></category>
		<category><![CDATA[RNA secondary structure]]></category>
		<category><![CDATA[sequence-to-function prediction in biology]]></category>
		<category><![CDATA[single-cell analysis]]></category>
		<category><![CDATA[transformer architecture]]></category>
		<category><![CDATA[transformer models in genomics]]></category>
		<category><![CDATA[variant effect prediction]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=232502</guid>

					<description><![CDATA[A new review from Wuhan University details how large language models are rapidly transforming protein, DNA, and RNA research, from structure prediction to single-cell analysis and genomic education.]]></description>
										<content:encoded><![CDATA[<p>When ChatGPT stunned the world with its command of human language, few predicted that the same underlying technology would soon be trained on DNA, RNA, and proteins. Yet a comprehensive review published in Frontiers of Digital Education by Shaopeng Li, Weiliang Fan, and Yu Zhou of Wuhan University maps out exactly how large language models, or LLMs, have swept into genomics in just four years, transforming how scientists interpret the molecular instructions that govern life. The review, published on 28 May 2025, argues that the sequential nature of biological data makes it strikingly similar to human text, and that the architectures built to master one can be repurposed to decode the other.</p>
<p>The analogy is more than a metaphor. DNA is written in an alphabet of four nucleotide letters, proteins in an alphabet of twenty amino acids, and RNA in a four-letter code of its own. Just as words derive meaning from their context in a sentence, the function of a nucleotide or amino acid depends on its position within a sequence and on the sequences that surround it. The transformer architecture, introduced in 2017 in the landmark paper Attention Is All You Need, excels precisely at capturing such long-range contextual dependencies. By applying self-attention mechanisms across millions of biological sequences, models can learn statistical patterns that reflect evolutionary constraints, structural motifs, and functional sites without any explicit programming.</p>
<p>The review organizes the field into biological foundation models, general-purpose systems pretrained on massive unlabeled sequence databases, and specialized models tailored to particular problems. On the protein side, models such as ProtTrans and the Evolutionary Scale Modeling family demonstrated that unsupervised learning on hundreds of millions of sequences yields internal representations from which structure and function can be read off. This line of work culminated in AlphaFold and its successors, which achieved atomic-level accuracy in protein structure prediction and, with AlphaFold 3, extended to predicting interactions between proteins, nucleic acids, and small molecules. Generative protein LLMs such as ProGen and ProGen2 have gone further, producing entirely novel amino acid sequences that fold into functional proteins, a capability the review highlights as one of the most consequential advances of the past decade.</p>
<p>DNA language models have followed a parallel trajectory. The Nucleotide Transformer, built and evaluated on human and multi-species genomes, showed that robust foundation models could be trained at nucleotide resolution. Subsequent work demonstrated that such models are powerful predictors of genome-wide variant effects, meaning they can estimate whether a mutation in a patient&#8217;s genome is likely to be harmful, a task central to clinical genetics. The review also points to Evo and its successor Evo 2, genome-scale models capable of modeling and designing sequences across all domains of life, from bacteria to humans. These systems operate at a scale that was unthinkable for genomics only a few years ago, processing contexts long enough to encompass entire genes and their regulatory neighborhoods.</p>
<p>RNA has proven a particularly fertile ground because its structure and function are notoriously difficult to predict from sequence alone. Classical approaches relied on thermodynamic folding algorithms, but deep learning models such as UFold and RNAformer have shown that learned representations can match or exceed physics-based methods for secondary structure prediction. More recently, language model-based deep learning approaches have achieved accurate RNA 3D structure prediction, and motif-aware pretraining strategies have produced multi-purpose RNA models that handle splicing prediction, binding site identification, and function annotation within a single framework. The review notes that self-supervised learning on millions of primary RNA sequences from dozens of vertebrate species has substantially improved sequence-based splicing prediction, with direct implications for interpreting disease-causing mutations in untranslated regions and splice sites.</p>
<p>Beyond the three canonical sequence types, the review surveys a rapidly growing ecosystem of specialized applications. Interaction prediction models forecast which proteins bind which RNAs, which RNAs pair with other RNAs, and how mutations perturb these contacts. Structure prediction systems now cover proteins, RNA, and biomolecular complexes at near-experimental accuracy. Perhaps most striking is the single-cell revolution: foundation models such as scGPT, scBERT, and Geneformer treat gene expression profiles across thousands of individual cells as a kind of language, learning universal representations that support cell type annotation, perturbation response prediction, and the integration of massive transcriptomic datasets. Studies assessing GPT-4 for cell type annotation suggest that general-purpose chatbots can also contribute, though specialized models currently hold the edge on most benchmarks.</p>
<p>The authors are candid about the obstacles. Hallucination, the tendency of generative models to produce plausible but false outputs, is a documented failure mode of natural language generation that carries serious risks when outputs inform drug design or clinical interpretation. Benchmarking remains immature: evaluation suites such as DART-Eval for regulatory DNA and PertEval-scFM for perturbation prediction reveal that performance gains do not always transfer across datasets and cell types. Computational cost is another barrier, since genome-scale attention models demand enormous memory, motivating architectural innovations such as the Hyena hierarchy and Mamba-based state space models that the review identifies as promising directions for long-context biological modeling. Data quality, batch effects in single-cell experiments, and the interpretability of learned representations round out the open challenges.</p>
<p>What distinguishes this review from many surveys of artificial intelligence in biology is its educational dimension. Because it appears in a journal focused on digital education, the authors devote substantial attention to integrating LLMs into genomics teaching and learning. They propose practical projects in which students fine-tune pretrained models on small datasets, use conversational agents to simplify bioinformatics workflows, and learn to critically evaluate model outputs. Tools such as BioMANIA, which allows researchers to interrogate bioinformatics pipelines through natural language conversation, and BioCoder, a benchmark for bioinformatics code generation, illustrate how LLMs can lower the technical barriers that traditionally separate biologists from computational analysis. The authors argue that fluency in these tools is becoming as essential to the modern life scientist as statistical literacy.</p>
<p>The trajectory the review describes is remarkably compressed. In roughly four years, language models have progressed from proof-of-concept embeddings to systems that simulate 500 million years of evolutionary history, design functional proteins, predict the pathogenicity of genetic variants, and annotate cell atlases containing millions of cells. The Wuhan University team frames the current moment as an inflection point: the same scaling laws and reasoning enhancements that drive progress in general-purpose AI, including reinforcement learning approaches that sharpen chain-of-thought reasoning, are poised to accelerate biological discovery further. If the challenges of reliability, evaluation, and interpretability can be met, the authors suggest, the fusion of language models and genomics may redefine both how biology is researched and how the next generation of scientists is trained, turning the genome from an inscrutable code into a text that machines can read, annotate, and even rewrite.</p>
<p><strong>Subject of Research:</strong> Applications of large language models in genomics, including biological foundation models for protein, DNA, and RNA sequence analysis</p>
<p><strong>Article Title:</strong> AI-Empowered Genome Decoding: Applications of Large Language Models in Genomics</p>
<p><strong>Article References:</strong> Li, S., Fan, W., &amp; Zhou, Y. (2025). AI-Empowered Genome Decoding: Applications of Large Language Models in Genomics. <em>Frontiers of Digital Education, 2</em>(1), Article 14. <a href="https://doi.org/10.1007/s44366-025-0051-1" rel="noopener noreferrer">https://doi.org/10.1007/s44366-025-0051-1</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s44366-025-0051-1" rel="noopener noreferrer">10.1007/s44366-025-0051-1</a></p>
<p><strong>Keywords:</strong> large language models, genomics, deep learning, protein structure prediction, DNA language models, RNA secondary structure, single-cell analysis, foundation models, AlphaFold, transformer architecture, bioinformatics education, variant effect prediction</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">232502</post-id>	</item>
	</channel>
</rss>
