<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>model fine-tuning &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/model-fine-tuning/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Fri, 02 Oct 2026 06:32:11 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.3</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>model fine-tuning &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>Open-Weight AI Models Like DeepSeek Could Reshape Global Education Access</title>
		<link>https://scienmag.com/open-weight-ai-models-like-deepseek-could-reshape-global-education-access/</link>
		
		<dc:creator><![CDATA[Courtney Benton]]></dc:creator>
		<pubDate>Fri, 02 Oct 2026 06:32:11 +0000</pubDate>
				<category><![CDATA[Social Science]]></category>
		<category><![CDATA[AI in classrooms]]></category>
		<category><![CDATA[AI-driven educational empowerment]]></category>
		<category><![CDATA[Chinese AI development in education]]></category>
		<category><![CDATA[DeepSeek]]></category>
		<category><![CDATA[DeepSeek large language model]]></category>
		<category><![CDATA[democratization of AI technology]]></category>
		<category><![CDATA[digital education]]></category>
		<category><![CDATA[digital education innovations 2025]]></category>
		<category><![CDATA[Educational Equity]]></category>
		<category><![CDATA[global education access]]></category>
		<category><![CDATA[intelligent tutoring systems]]></category>
		<category><![CDATA[large language models]]></category>
		<category><![CDATA[Mixture of Experts]]></category>
		<category><![CDATA[model fine-tuning]]></category>
		<category><![CDATA[multilingual education]]></category>
		<category><![CDATA[open versus proprietary AI systems]]></category>
		<category><![CDATA[open-source AI in education]]></category>
		<category><![CDATA[open-weight AI]]></category>
		<category><![CDATA[Open-Weight AI Models]]></category>
		<category><![CDATA[personalized tutoring]]></category>
		<category><![CDATA[societal benefits of open AI models]]></category>
		<category><![CDATA[transformative impact of large language models]]></category>
		<category><![CDATA[transformer-based neural networks for learning]]></category>
		<category><![CDATA[Zhejiang University]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=226174</guid>

					<description><![CDATA[A new commentary in Frontiers of Digital Education argues that open-weight large language models such as DeepSeek could extend personalized, multilingual AI tutoring to learners and teachers across whole societies.]]></description>
										<content:encoded><![CDATA[<p>When a research commentary appears in a journal devoted to digital education and its title contains the name of an artificial intelligence system rather than a teaching method, it signals a shift in how scholars are thinking about the future of learning. A commentary published in Frontiers of Digital Education by Fei Wu of the College of Computer Science and Technology at Zhejiang University does exactly that, arguing that DeepSeek, the large language model family developed in China, points toward a form of global education empowerment that could extend across whole societies rather than remaining confined to well-resourced institutions. The piece, published on 8 May 2025 as article 26 in the journal&#8217;s second volume, frames the rapid maturation of open large language models as an educational event as much as a technological one.</p>
<p>To understand why an education journal would devote a commentary to a single model family, it helps to look at what DeepSeek actually is and how it differs from the closed, proprietary systems that dominated public attention in the years before its release. Large language models are neural networks, typically built on the transformer architecture, that are trained on enormous text corpora to predict the next token in a sequence. From that seemingly simple objective, and with sufficient scale of parameters and data, such models acquire the ability to answer questions, summarize documents, write and debug code, translate between languages, and carry out multi-step reasoning. The quality of these abilities depends on the scale of the model, the curation of the training data, and the alignment techniques applied after pre-training, such as supervised fine-tuning and reinforcement learning from human feedback.</p>
<p>What distinguished DeepSeek in the eyes of many observers was the combination of strong reported performance with an open-weight release strategy. Open-weight models publish their trained parameters so that anyone with the hardware and expertise can download, run, fine-tune, and build upon them, in contrast to application-programming-interface access to closed models where the weights remain hidden. This distinction matters enormously for education, because the economics of access change completely. A school district, a university laboratory, or a ministry of education that deploys an open-weight model on its own servers pays for computation rather than per-token licensing, and gains the ability to inspect, adapt, and localize the system in ways that closed platforms do not permit.</p>
<p>The technical machinery behind such efficiency gains is worth spelling out, because it explains why commentary authors see educational empowerment as realistic rather than aspirational. Modern efficient large models increasingly rely on mixture-of-experts architectures, in which only a subset of the network&#8217;s parameters is activated for any given token, allowing total parameter counts to grow while the computational cost of each forward pass stays bounded. They also employ techniques such as multi-head latent attention to compress key-value caches and reduce memory traffic during inference, and quantization methods that shrink the numerical precision of weights so that models can run on consumer-grade graphics cards or even laptops. When a model family achieves competitive reasoning performance at a fraction of the training cost reported for frontier closed models, the barrier to entry for institutions in lower-income countries drops by orders of magnitude.</p>
<p>Wu&#8217;s commentary situates these developments within the long-standing problem of educational inequity. Access to high-quality instruction, tutoring, and learning materials has always been distributed unevenly, both between countries and within them. A skilled human tutor can adapt explanations to a learner&#8217;s misconceptions in real time, but such tutoring is expensive and scarce, a constraint that has historically limited the reach of personalized education. Intelligent tutoring systems have pursued this goal for decades, yet earlier generations of software were brittle, requiring hand-authored rules or domain models that could not generalize beyond narrow curricula. Large language models changed the calculus because their knowledge and their instructional flexibility emerge from general pre-training rather than from laborious manual encoding of subject matter.</p>
<p>The commentary&#8217;s vision of empowerment for the whole society rests on several concrete affordances that open models bring to learners and teachers. A student in a remote region with a smartphone and intermittent connectivity can, in principle, query a locally deployed model about algebra, grammar, or science concepts in her own language, receiving explanations tailored to her level. A teacher can use the model to draft lesson plans, generate practice problems at graded difficulty levels, produce multiple explanations of the same concept for different learning styles, and automate the first pass of feedback on written work. Administrators can analyze learning data to identify where curricula are failing. None of these applications is hypothetical in kind; each has been demonstrated in research prototypes, and the open-weight availability of capable models makes them deployable without dependence on foreign cloud providers or unaffordable subscription fees.</p>
<p>Language is a central part of this argument. The most capable proprietary models have historically performed best in English, leaving learners in the majority of the world&#8217;s languages at a disadvantage. Open-weight models can be fine-tuned on corpora in underrepresented languages, a process that requires far less data and compute than training from scratch. Continued pre-training on domain-specific and language-specific text, followed by instruction tuning with locally authored examples, can produce educational assistants that understand regional curricula, national examination formats, and culturally situated examples. This capacity for localization is precisely what a global empowerment agenda requires, and it is a capability that closed platforms, whatever their quality, do not offer to the communities that need it most.</p>
<p>At the same time, the commentary&#8217;s optimistic framing invites scrutiny of the risks that accompany any large-scale deployment of generative models in education. Language models can produce fluent but incorrect statements, a phenomenon usually called hallucination, and learners who lack domain knowledge are the least equipped to detect such errors. Uncritical reliance on generated answers could undermine the productive struggle through which students actually learn. There are also questions of data privacy when student interactions are logged, of algorithmic bias when training corpora encode social stereotypes, and of academic integrity when the same tool that explains a concept can also complete the homework. Responsible deployment therefore demands pedagogical design that positions the model as a tutor and scaffold rather than an answer engine, alongside transparency about model limitations and human oversight of high-stakes assessments.</p>
<p>The publication details of the commentary itself illustrate how the academic ecosystem is adapting. Wu is affiliated with Zhejiang University in Hangzhou, and the piece appeared in Frontiers of Digital Education, a journal published by Higher Education Press through Springer Nature that focuses on how digital technologies transform teaching and learning. According to the journal&#8217;s disclosure, Wu serves on its editorial board and was excluded from the peer-review process and all editorial decisions concerning his own article, with independent editors handling review to minimize bias. The journal reports that the article has already accumulated citations within months of publication, an indication of how quickly the research community is engaging with the questions it raises. The author states that all data analysed in the study are contained within the published article itself.</p>
<p>Whether open-weight models like DeepSeek fulfill the promise of global educational empowerment will depend on choices that extend well beyond model architecture. Hardware access, electricity reliability, internet infrastructure, teacher training, and government policy all shape whether a technically available capability becomes a socially realized one. But the direction of travel is clear: the marginal cost of providing a competent, patient, multilingual explanation of almost any school subject is falling toward the cost of computation alone. If the educational community builds the safeguards, curricula, and local adaptations needed to deploy these systems wisely, the commentary&#8217;s title may come to read less like a slogan and more like a description of what actually happened, as a technology developed for general purposes found its most consequential application in the classrooms of the whole society.</p>
<p><strong>Subject of Research:</strong> The role of open-weight large language models such as DeepSeek in advancing global educational equity and empowerment</p>
<p><strong>Article Title:</strong> DeepSeek: Toward Global Education Empowerment for the Whole Society</p>
<p><strong>Article References:</strong> Wu, F. (2025). DeepSeek: Toward Global Education Empowerment for the Whole Society. <em>Frontiers of Digital Education, 2</em>(2), Article 26. <a href="https://doi.org/10.1007/s44366-025-0062-y" rel="noopener noreferrer">https://doi.org/10.1007/s44366-025-0062-y</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s44366-025-0062-y" rel="noopener noreferrer">10.1007/s44366-025-0062-y</a></p>
<p><strong>Keywords:</strong> DeepSeek, large language models, open-weight AI, digital education, educational equity, personalized tutoring, mixture-of-experts, model fine-tuning, multilingual education, intelligent tutoring systems, AI in classrooms, Zhejiang University</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">226174</post-id>	</item>
		<item>
		<title>When Knowledge Graphs Meet Large Language Models: A New Roadmap for Trustworthy AI Reasoning</title>
		<link>https://scienmag.com/when-knowledge-graphs-meet-large-language-models-a-new-roadmap-for-trustworthy-ai-reasoning/</link>
		
		<dc:creator><![CDATA[Denise Maddox]]></dc:creator>
		<pubDate>Sun, 20 Sep 2026 19:38:34 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[AI agents]]></category>
		<category><![CDATA[AI fact verification]]></category>
		<category><![CDATA[AI representation gap]]></category>
		<category><![CDATA[cognitive science in AI development]]></category>
		<category><![CDATA[cognitive synergy]]></category>
		<category><![CDATA[explainability in AI]]></category>
		<category><![CDATA[explainable AI]]></category>
		<category><![CDATA[factual hallucination]]></category>
		<category><![CDATA[hierarchical theoretical frameworks for AI]]></category>
		<category><![CDATA[knowledge agents]]></category>
		<category><![CDATA[knowledge graph and language model fusion]]></category>
		<category><![CDATA[Knowledge graph integration with large language models]]></category>
		<category><![CDATA[knowledge graphs]]></category>
		<category><![CDATA[knowledge reasoning]]></category>
		<category><![CDATA[large language models]]></category>
		<category><![CDATA[model fine-tuning]]></category>
		<category><![CDATA[multi-step reasoning challenges]]></category>
		<category><![CDATA[neuro-symbolic integration]]></category>
		<category><![CDATA[paradigm conflicts in AI]]></category>
		<category><![CDATA[prompt engineering]]></category>
		<category><![CDATA[retrieval-augmented generation]]></category>
		<category><![CDATA[synergy bottleneck in AI systems]]></category>
		<category><![CDATA[trustworthy AI reasoning]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=201848</guid>

					<description><![CDATA[A new survey maps how knowledge graphs can fix the hallucinations, weak reasoning, and opacity of large language models through five technical pathways.]]></description>
										<content:encoded><![CDATA[<p>Large language models have dazzled the world with their fluency, yet they remain haunted by a familiar trio of flaws: they invent facts, they stumble through multi-step reasoning, and they cannot explain why they said what they said. A new survey published in Knowledge and Information Systems argues that the antidote may already exist in one of artificial intelligence&#8217;s oldest and most reliable inventions, the knowledge graph, and it offers the most systematic map yet of how these two very different kinds of machine intelligence can be fused into something greater than either alone.</p>
<p>The review, authored by Jiale Wu, Yijiang Zhao, Zhuhua Liao, and Min Liu of Hunan University of Science and Technology, goes beyond the usual catalog of techniques that dominates the survey literature. Instead, the team builds a hierarchical theoretical framework, described as a root-phenomenon-consequence structure, grounded in the cognitive science of neuro-symbolic integration. The central claim is that the difficulties of combining knowledge graphs with large language models are not a scattered collection of engineering annoyances but the surface expressions of three deep, interlocking challenges: a representation gap, a synergy bottleneck, and a paradigm conflict.</p>
<p>The representation gap refers to the fundamental mismatch between how knowledge graphs and language models encode meaning. A knowledge graph stores the world as discrete triples, subject, relation, object, arranged in a symbolic network that a machine can traverse with perfect fidelity. A large language model, by contrast, compresses statistical regularities of language into billions of continuous neural parameters. One system reasons with symbols it can inspect; the other reasons with patterns it cannot. Bridging these two encodings without losing the strengths of either is the first root problem the survey identifies.</p>
<p>The synergy bottleneck concerns what happens when the two systems are actually coupled. Simply retrieving a subgraph and pasting it into a prompt does not guarantee that the model will use the evidence correctly, and fine-tuning a model on graph data can degrade the very linguistic competence that made it useful. The paradigm conflict, meanwhile, is more philosophical: symbolic systems are built for exact, verifiable inference, while neural systems are built for tolerant, probabilistic generalization. The survey argues that progress depends on recognizing these tensions explicitly rather than papering over them with ever-larger models.</p>
<p>To organize the technical landscape, the authors propose a five-dimensional taxonomy of integration approaches. The first dimension is prompt engineering, where knowledge graph content is translated into text and placed in the model&#8217;s context window. Methods in this family include chain-of-thought prompting over graphs, frameworks such as KG-GPT and KG-CoT, and systems like Think-on-Graph that guide a language model step by step along relevant knowledge paths. Prompting is attractive because it requires no retraining, but it consumes context space and depends heavily on the quality of the retrieved evidence.</p>
<p>The second dimension, retrieval augmented generation, has become one of the fastest-moving areas in the field. Rather than relying on a model&#8217;s frozen internal memory, these systems fetch structured evidence at question time. The survey traces an evolution from classic dense retrieval to graph-aware pipelines such as GNN-RAG, HyKGE for medical question answering, and the Think-on-Graph series, whose later versions employ multi-agent, dual-evolving context retrieval over heterogeneous graphs. Neurobiologically inspired architectures such as HippoRAG, which model long-term memory as a graph, illustrate how far the retrieval paradigm has drifted from simple keyword search toward something resembling structured recall.</p>
<p>The third dimension is model fine-tuning, in which knowledge graph information is baked into the model&#8217;s parameters. The survey highlights parameter-efficient techniques such as KG-Adapter, infuser-guided knowledge integration, and knowledge-graph-enhanced model editing, along with distillation approaches that transfer reasoning ability from large teachers to smaller students. Fine-tuning promises deeper integration than prompting, but it raises cost, rigidity, and knowledge-timeliness problems: a model trained on yesterday&#8217;s graph cannot easily learn today&#8217;s facts.</p>
<p>The fourth and fifth dimensions mark what the authors see as the field&#8217;s evolutionary frontier. Large reasoning model collaboration pairs knowledge graphs with the new generation of reasoning-heavy models, exemplified by reinforcement-learning-trained systems such as DeepSeek-R1, and by frameworks like KG-o1 and Search-o1 that let a reasoning model consult a graph during extended chains of deliberation. Knowledge agents, the final dimension, go further still: autonomous systems such as KG-Agent, AriGraph with its episodic memory, and Generate-on-Graph treat the language model as an agent that can plan, query, and even extend an incomplete knowledge graph on its own. The survey characterizes the overall trajectory of these five dimensions as a progression from external guidance toward autonomous cognition, a shift with profound implications for how much trust such systems can eventually earn.</p>
<p>The practical stakes are already visible in vertical domains. In health care, graph-augmented frameworks such as Medical Graph RAG, KoSEL, and MedReason ground clinical question answering in curated medical knowledge, reducing the risk of confidently wrong answers in settings where errors can harm patients. In law, systems like ChatLaw combine knowledge graphs with mixture-of-experts architectures to anchor legal reasoning in statutes and precedent, while researchers have explored how graph-based prompting can clarify the legal implications of model outputs. In scientific research, knowledge graphs are being coupled with language models for tasks ranging from biomedical literature mining in Alzheimer&#8217;s studies to automated retrosynthesis planning of macromolecules, where a model proposes reaction routes that a structured chemical knowledge base validates.</p>
<p>None of this amounts to a solved problem, and the survey is candid about the open challenges. Neuro-symbolic alignment remains immature: there is still no principled theory for when symbolic structure should override neural intuition or vice versa. Knowledge timeliness is equally pressing, since real-world facts change far faster than models or graphs can be updated, and stale evidence can be worse than none. The authors point toward lightweight reasoning engines that could make graph-guided inference affordable at scale, and toward high-trust agent systems in which every reasoning step is auditable against an explicit knowledge structure. If those directions mature, the hybrid of neural fluency and symbolic rigor that this survey maps out could become the architecture on which genuinely reliable AI reasoning is built, transforming language models from persuasive improvisers into accountable thinkers that show their work.</p>
<p><strong>Subject of Research:</strong> The integration of knowledge graphs with large language models to improve factual accuracy, reasoning, and explainability in AI systems.</p>
<p><strong>Article Title:</strong> A survey on knowledge graph-augmented large language model reasoning: theoretical challenges, technical pathways, and evolutionary logic</p>
<p><strong>Article References:</strong> Wu, J., Zhao, Y., Liao, Z., &amp; Liu, M. (2026). A survey on knowledge graph-augmented large language model reasoning: theoretical challenges, technical pathways, and evolutionary logic. <em>Knowledge and Information Systems, 68</em>(1), Article 261. <a href="https://doi.org/10.1007/s10115-026-02880-5" rel="noopener noreferrer">https://doi.org/10.1007/s10115-026-02880-5</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s10115-026-02880-5" rel="noopener noreferrer">10.1007/s10115-026-02880-5</a></p>
<p><strong>Keywords:</strong> large language models, knowledge graphs, neuro-symbolic integration, retrieval augmented generation, prompt engineering, model fine-tuning, knowledge agents, factual hallucination, knowledge reasoning, cognitive synergy, AI agents, explainable AI</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">201848</post-id>	</item>
	</channel>
</rss>
