<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>open-source educational AI tools &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/open-source-educational-ai-tools/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Fri, 02 Oct 2026 23:54:52 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>open-source educational AI tools &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>Chinese Education Gets Its Own Open-Source AI Language Model</title>
		<link>https://scienmag.com/chinese-education-gets-its-own-open-source-ai-language-model/</link>
		
		<dc:creator><![CDATA[Courtney Benton]]></dc:creator>
		<pubDate>Fri, 02 Oct 2026 23:54:52 +0000</pubDate>
				<category><![CDATA[Social Science]]></category>
		<category><![CDATA[AI in Chinese education]]></category>
		<category><![CDATA[AI research in Chinese education]]></category>
		<category><![CDATA[C-Eval]]></category>
		<category><![CDATA[CELLM]]></category>
		<category><![CDATA[CELLM Chinese education language model]]></category>
		<category><![CDATA[Chinese education]]></category>
		<category><![CDATA[Chinese education AI]]></category>
		<category><![CDATA[Chinese educational data training]]></category>
		<category><![CDATA[Chinese language tokenization challenges]]></category>
		<category><![CDATA[Chinese-specific NLP models]]></category>
		<category><![CDATA[CMMLU]]></category>
		<category><![CDATA[DeepSpeed]]></category>
		<category><![CDATA[education-focused AI model development]]></category>
		<category><![CDATA[educational AI]]></category>
		<category><![CDATA[instruction fine-tuning]]></category>
		<category><![CDATA[large language models]]></category>
		<category><![CDATA[large language models for Chinese language]]></category>
		<category><![CDATA[multilingual AI for education]]></category>
		<category><![CDATA[open-source]]></category>
		<category><![CDATA[open-source educational AI tools]]></category>
		<category><![CDATA[open-source language models for Chinese]]></category>
		<category><![CDATA[pre-training]]></category>
		<category><![CDATA[rotary position embeddings]]></category>
		<category><![CDATA[transformer architecture]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=229691</guid>

					<description><![CDATA[Researchers in Shanghai have open-sourced a 1.5-billion-parameter language model trained from scratch for Chinese education, complete with a 258,000-entry instruction dataset and public benchmarks.]]></description>
										<content:encoded><![CDATA[<p>A team of researchers in Shanghai has built and released a large language model designed specifically for Chinese education, and they have given away every piece of it: the model weights, the training data, and the code. The system, called the Chinese Education Large Language Model, or CELLM, was developed by Wentao Liu of East China Normal University&#8217;s Shanghai Institute for AI Education, Hao Hao of Shanghai Jiao Tong University, and Aimin Zhou of East China Normal University&#8217;s School of Computer Science and Technology. Their study, published in Frontiers of Digital Education, describes a compact 1.5-billion-parameter model trained from scratch on education-focused Chinese text, along with a newly open-sourced instruction dataset of more than 258,000 entries.</p>
<p>The motivation behind the project is a persistent gap in the artificial intelligence landscape. Although open-source large language models have advanced rapidly in recent years, most of the research community&#8217;s effort has poured into general-purpose models trained predominantly on English data. That imbalance creates real problems for anyone studying or deploying AI in Chinese education. Chinese is a language with distinctive tokenization challenges, rich morphological structure, and a vast educational literature that English-centric models simply do not capture well. Educational applications add another layer of specificity: the vocabulary of pedagogy, curriculum standards, exam questions, and subject-specific reasoning in Chinese differs substantially from the web text that dominates most training corpora.</p>
<p>Rather than fine-tuning an existing model, the team took the more demanding route of training CELLM from the ground up, a decision that gave them full control over the data pipeline and the architecture. The process unfolded in two stages. The first was pre-training, in which the model learned the statistical structure of language from an open-source dataset drawn from the Chinese education domain. During this phase, the model absorbed the patterns of educational Chinese text, building the foundational knowledge that later stages would refine. The second stage was instruction fine-tuning, which teaches a base model to follow directions, answer questions, and behave like an assistant rather than a text predictor.</p>
<p>For that second stage, the researchers faced a familiar obstacle: high-quality Chinese instruction data for education was scarce. Their response was to build their own. They constructed a Chinese instruction dataset comprising over 258,000 data entries and, in the spirit of the entire project, released it openly. Instruction datasets of this kind typically pair prompts with high-quality responses, spanning formats such as question answering, explanations, and multi-turn dialogue. By curating the data themselves, the team could steer the model toward the kinds of interactions that matter in educational contexts, from explaining a mathematics problem to discussing language teaching strategies.</p>
<p>The architecture underlying CELLM draws on the core technologies that have defined modern open-source language models. The team reviewed and synthesized the design choices of representative open-source systems before settling on their own configuration. Among the techniques reflected in the model&#8217;s lineage are rotary position embeddings, an approach that encodes word order by rotating vector representations and has become a staple of contemporary transformer designs. The model also builds on advances in attention mechanisms, including grouped query attention, which reduces computational cost by sharing key-value projections across multiple query heads, a strategy popularized by efficient inference research.</p>
<p>Other components of the design reflect a decade of accumulated transformer engineering. The feed-forward layers employ variants of gated linear units, activation functions shown to improve transformer performance over standard alternatives. Training at scale was supported by DeepSpeed, the distributed training framework developed to make models with hundreds of billions of parameters feasible on real hardware. The team also drew on research into scaling laws and overtraining, which examines how model performance grows with parameters and data, informing how a relatively small model can be trained to punch above its weight class.</p>
<p>The choice of a 1.5-billion-parameter size is itself significant. In an era when frontier models boast hundreds of billions of parameters, a compact model might seem modest. But smaller models are dramatically cheaper to train, run, and deploy, which matters enormously for educational research groups, schools, and developers working with limited computing budgets. A 1.5-billion-parameter model can run on consumer-grade hardware, making it practical for experimentation and classroom applications alike. The trade-off is capability, and the evaluation results provide a transparent accounting of where the model stands.</p>
<p>That transparency is one of the study&#8217;s most valuable contributions. The researchers benchmarked CELLM across multiple evaluation datasets, including suites designed to measure massive multitask language understanding in Chinese, such as C-Eval and CMMLU, which test knowledge across disciplines and difficulty levels, alongside mathematical problem-solving benchmarks. The published results establish a reference baseline for future research, giving subsequent teams a clear point of comparison. In a field where claims often outpace evidence, a fully documented baseline for an education-specific Chinese model fills a genuine need.</p>
<p>The open-source release strategy amplifies the work&#8217;s potential impact. Everything generated during the study, including the models, the data, and the code, has been made publicly available. This stands in deliberate contrast to the closed ecosystems of the largest commercial AI systems, whose training data and internal workings remain opaque. Open release allows other researchers to scrutinize the model, reproduce the results, extend the training, or adapt the instruction dataset to adjacent domains. It also enables studies of the models themselves, including investigations of bias, generalization, and how a model&#8217;s capabilities trace back to its pre-training data, questions that are difficult or impossible to answer with proprietary systems.</p>
<p>For the field of Chinese education research, CELLM represents both a tool and a template. As a tool, it offers researchers a capable, domain-tuned language model they can study, fine-tune, and deploy without licensing barriers. As a template, it demonstrates that building a specialized model from scratch, with a purpose-built instruction dataset and rigorous public evaluation, is achievable by a small academic team. As open-source language models continue to proliferate across languages and domains, the Shanghai team&#8217;s work signals a shift toward AI research that treats education not as an afterthought of general-purpose systems, but as a domain deserving models, data, and benchmarks of its own.</p>
<p><strong>Subject of Research:</strong> An open-source large language model trained from scratch for Chinese education research</p>
<p><strong>Article Title:</strong> An Open-Source Large Language Model for Chinese Education Research</p>
<p><strong>Article References:</strong> Liu, W., Hao, H., &amp; Zhou, A. (2025). An Open-Source Large Language Model for Chinese Education Research. <em>Frontiers of Digital Education, 2</em>(2), Article 23. <a href="https://doi.org/10.1007/s44366-025-0060-0" rel="noopener noreferrer">https://doi.org/10.1007/s44366-025-0060-0</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s44366-025-0060-0" rel="noopener noreferrer">10.1007/s44366-025-0060-0</a></p>
<p><strong>Keywords:</strong> large language models, open source, Chinese education, CELLM, instruction fine-tuning, pre-training, transformer architecture, rotary position embeddings, DeepSpeed, C-Eval, CMMLU, educational AI</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">229691</post-id>	</item>
	</channel>
</rss>
