<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>Mixture of Experts &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/mixture-of-experts/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Fri, 02 Oct 2026 06:32:11 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>Mixture of Experts &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>Open-Weight AI Models Like DeepSeek Could Reshape Global Education Access</title>
		<link>https://scienmag.com/open-weight-ai-models-like-deepseek-could-reshape-global-education-access/</link>
		
		<dc:creator><![CDATA[Courtney Benton]]></dc:creator>
		<pubDate>Fri, 02 Oct 2026 06:32:11 +0000</pubDate>
				<category><![CDATA[Social Science]]></category>
		<category><![CDATA[AI in classrooms]]></category>
		<category><![CDATA[AI-driven educational empowerment]]></category>
		<category><![CDATA[Chinese AI development in education]]></category>
		<category><![CDATA[DeepSeek]]></category>
		<category><![CDATA[DeepSeek large language model]]></category>
		<category><![CDATA[democratization of AI technology]]></category>
		<category><![CDATA[digital education]]></category>
		<category><![CDATA[digital education innovations 2025]]></category>
		<category><![CDATA[Educational Equity]]></category>
		<category><![CDATA[global education access]]></category>
		<category><![CDATA[intelligent tutoring systems]]></category>
		<category><![CDATA[large language models]]></category>
		<category><![CDATA[Mixture of Experts]]></category>
		<category><![CDATA[model fine-tuning]]></category>
		<category><![CDATA[multilingual education]]></category>
		<category><![CDATA[open versus proprietary AI systems]]></category>
		<category><![CDATA[open-source AI in education]]></category>
		<category><![CDATA[open-weight AI]]></category>
		<category><![CDATA[Open-Weight AI Models]]></category>
		<category><![CDATA[personalized tutoring]]></category>
		<category><![CDATA[societal benefits of open AI models]]></category>
		<category><![CDATA[transformative impact of large language models]]></category>
		<category><![CDATA[transformer-based neural networks for learning]]></category>
		<category><![CDATA[Zhejiang University]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=226174</guid>

					<description><![CDATA[A new commentary in Frontiers of Digital Education argues that open-weight large language models such as DeepSeek could extend personalized, multilingual AI tutoring to learners and teachers across whole societies.]]></description>
										<content:encoded><![CDATA[<p>When a research commentary appears in a journal devoted to digital education and its title contains the name of an artificial intelligence system rather than a teaching method, it signals a shift in how scholars are thinking about the future of learning. A commentary published in Frontiers of Digital Education by Fei Wu of the College of Computer Science and Technology at Zhejiang University does exactly that, arguing that DeepSeek, the large language model family developed in China, points toward a form of global education empowerment that could extend across whole societies rather than remaining confined to well-resourced institutions. The piece, published on 8 May 2025 as article 26 in the journal&#8217;s second volume, frames the rapid maturation of open large language models as an educational event as much as a technological one.</p>
<p>To understand why an education journal would devote a commentary to a single model family, it helps to look at what DeepSeek actually is and how it differs from the closed, proprietary systems that dominated public attention in the years before its release. Large language models are neural networks, typically built on the transformer architecture, that are trained on enormous text corpora to predict the next token in a sequence. From that seemingly simple objective, and with sufficient scale of parameters and data, such models acquire the ability to answer questions, summarize documents, write and debug code, translate between languages, and carry out multi-step reasoning. The quality of these abilities depends on the scale of the model, the curation of the training data, and the alignment techniques applied after pre-training, such as supervised fine-tuning and reinforcement learning from human feedback.</p>
<p>What distinguished DeepSeek in the eyes of many observers was the combination of strong reported performance with an open-weight release strategy. Open-weight models publish their trained parameters so that anyone with the hardware and expertise can download, run, fine-tune, and build upon them, in contrast to application-programming-interface access to closed models where the weights remain hidden. This distinction matters enormously for education, because the economics of access change completely. A school district, a university laboratory, or a ministry of education that deploys an open-weight model on its own servers pays for computation rather than per-token licensing, and gains the ability to inspect, adapt, and localize the system in ways that closed platforms do not permit.</p>
<p>The technical machinery behind such efficiency gains is worth spelling out, because it explains why commentary authors see educational empowerment as realistic rather than aspirational. Modern efficient large models increasingly rely on mixture-of-experts architectures, in which only a subset of the network&#8217;s parameters is activated for any given token, allowing total parameter counts to grow while the computational cost of each forward pass stays bounded. They also employ techniques such as multi-head latent attention to compress key-value caches and reduce memory traffic during inference, and quantization methods that shrink the numerical precision of weights so that models can run on consumer-grade graphics cards or even laptops. When a model family achieves competitive reasoning performance at a fraction of the training cost reported for frontier closed models, the barrier to entry for institutions in lower-income countries drops by orders of magnitude.</p>
<p>Wu&#8217;s commentary situates these developments within the long-standing problem of educational inequity. Access to high-quality instruction, tutoring, and learning materials has always been distributed unevenly, both between countries and within them. A skilled human tutor can adapt explanations to a learner&#8217;s misconceptions in real time, but such tutoring is expensive and scarce, a constraint that has historically limited the reach of personalized education. Intelligent tutoring systems have pursued this goal for decades, yet earlier generations of software were brittle, requiring hand-authored rules or domain models that could not generalize beyond narrow curricula. Large language models changed the calculus because their knowledge and their instructional flexibility emerge from general pre-training rather than from laborious manual encoding of subject matter.</p>
<p>The commentary&#8217;s vision of empowerment for the whole society rests on several concrete affordances that open models bring to learners and teachers. A student in a remote region with a smartphone and intermittent connectivity can, in principle, query a locally deployed model about algebra, grammar, or science concepts in her own language, receiving explanations tailored to her level. A teacher can use the model to draft lesson plans, generate practice problems at graded difficulty levels, produce multiple explanations of the same concept for different learning styles, and automate the first pass of feedback on written work. Administrators can analyze learning data to identify where curricula are failing. None of these applications is hypothetical in kind; each has been demonstrated in research prototypes, and the open-weight availability of capable models makes them deployable without dependence on foreign cloud providers or unaffordable subscription fees.</p>
<p>Language is a central part of this argument. The most capable proprietary models have historically performed best in English, leaving learners in the majority of the world&#8217;s languages at a disadvantage. Open-weight models can be fine-tuned on corpora in underrepresented languages, a process that requires far less data and compute than training from scratch. Continued pre-training on domain-specific and language-specific text, followed by instruction tuning with locally authored examples, can produce educational assistants that understand regional curricula, national examination formats, and culturally situated examples. This capacity for localization is precisely what a global empowerment agenda requires, and it is a capability that closed platforms, whatever their quality, do not offer to the communities that need it most.</p>
<p>At the same time, the commentary&#8217;s optimistic framing invites scrutiny of the risks that accompany any large-scale deployment of generative models in education. Language models can produce fluent but incorrect statements, a phenomenon usually called hallucination, and learners who lack domain knowledge are the least equipped to detect such errors. Uncritical reliance on generated answers could undermine the productive struggle through which students actually learn. There are also questions of data privacy when student interactions are logged, of algorithmic bias when training corpora encode social stereotypes, and of academic integrity when the same tool that explains a concept can also complete the homework. Responsible deployment therefore demands pedagogical design that positions the model as a tutor and scaffold rather than an answer engine, alongside transparency about model limitations and human oversight of high-stakes assessments.</p>
<p>The publication details of the commentary itself illustrate how the academic ecosystem is adapting. Wu is affiliated with Zhejiang University in Hangzhou, and the piece appeared in Frontiers of Digital Education, a journal published by Higher Education Press through Springer Nature that focuses on how digital technologies transform teaching and learning. According to the journal&#8217;s disclosure, Wu serves on its editorial board and was excluded from the peer-review process and all editorial decisions concerning his own article, with independent editors handling review to minimize bias. The journal reports that the article has already accumulated citations within months of publication, an indication of how quickly the research community is engaging with the questions it raises. The author states that all data analysed in the study are contained within the published article itself.</p>
<p>Whether open-weight models like DeepSeek fulfill the promise of global educational empowerment will depend on choices that extend well beyond model architecture. Hardware access, electricity reliability, internet infrastructure, teacher training, and government policy all shape whether a technically available capability becomes a socially realized one. But the direction of travel is clear: the marginal cost of providing a competent, patient, multilingual explanation of almost any school subject is falling toward the cost of computation alone. If the educational community builds the safeguards, curricula, and local adaptations needed to deploy these systems wisely, the commentary&#8217;s title may come to read less like a slogan and more like a description of what actually happened, as a technology developed for general purposes found its most consequential application in the classrooms of the whole society.</p>
<p><strong>Subject of Research:</strong> The role of open-weight large language models such as DeepSeek in advancing global educational equity and empowerment</p>
<p><strong>Article Title:</strong> DeepSeek: Toward Global Education Empowerment for the Whole Society</p>
<p><strong>Article References:</strong> Wu, F. (2025). DeepSeek: Toward Global Education Empowerment for the Whole Society. <em>Frontiers of Digital Education, 2</em>(2), Article 26. <a href="https://doi.org/10.1007/s44366-025-0062-y" rel="noopener noreferrer">https://doi.org/10.1007/s44366-025-0062-y</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s44366-025-0062-y" rel="noopener noreferrer">10.1007/s44366-025-0062-y</a></p>
<p><strong>Keywords:</strong> DeepSeek, large language models, open-weight AI, digital education, educational equity, personalized tutoring, mixture-of-experts, model fine-tuning, multilingual education, intelligent tutoring systems, AI in classrooms, Zhejiang University</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">226174</post-id>	</item>
		<item>
		<title>AI Learns to Read a Power Plant&#8217;s Mind: Transparent Neural Network Models Coal-Fired Boiler-Turbine Dynamics</title>
		<link>https://scienmag.com/ai-learns-to-read-a-power-plants-mind-transparent-neural-network-models-coal-fired-boiler-turbine-dynamics/</link>
		
		<dc:creator><![CDATA[Blake Davidson]]></dc:creator>
		<pubDate>Fri, 02 Oct 2026 05:53:06 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[AI applications in renewable energy integration]]></category>
		<category><![CDATA[AI safety in industrial control]]></category>
		<category><![CDATA[attention mechanism]]></category>
		<category><![CDATA[boiler-turbine system]]></category>
		<category><![CDATA[coal-fired power plant]]></category>
		<category><![CDATA[coal-fired power plant automation]]></category>
		<category><![CDATA[control systems]]></category>
		<category><![CDATA[deep learning]]></category>
		<category><![CDATA[dynamic behavior prediction of power plant equipment]]></category>
		<category><![CDATA[dynamic modeling]]></category>
		<category><![CDATA[explainable AI for industrial safety]]></category>
		<category><![CDATA[industrial AI]]></category>
		<category><![CDATA[industrial AI for power plant efficiency]]></category>
		<category><![CDATA[interpretability]]></category>
		<category><![CDATA[machine learning interpretability in power grid management]]></category>
		<category><![CDATA[Mixture of Experts]]></category>
		<category><![CDATA[multi-task learning]]></category>
		<category><![CDATA[neural network explainability]]></category>
		<category><![CDATA[neural network modeling of boiler-turbine dynamics]]></category>
		<category><![CDATA[neural network trust in energy systems]]></category>
		<category><![CDATA[physics-aligned deep learning for power plant control]]></category>
		<category><![CDATA[renewable energy integration]]></category>
		<category><![CDATA[time series prediction]]></category>
		<category><![CDATA[transparent AI models for boiler-turbine systems]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=226006</guid>

					<description><![CDATA[Researchers in China have developed a transparent deep learning framework called MDANet that accurately predicts the dynamic behavior of a 600 MW coal-fired boiler-turbine unit while revealing internal attention patterns consistent with the plant's known physics.]]></description>
										<content:encoded><![CDATA[<p>Deep learning has conquered everything from protein folding to language translation, but inside the control rooms of coal-fired power plants, a stubborn problem has kept artificial intelligence at arm&#8217;s length: the best-performing neural networks are black boxes. Engineers who must keep a 600 megawatt boiler-turbine unit running safely cannot simply trust a model that spits out predictions without explaining itself. A new study published in Complex &amp; Intelligent Systems tackles this trust problem head-on, presenting a physics-aligned deep learning framework that not only predicts the behavior of a boiler-turbine system with high accuracy but also shows its reasoning in a form that engineers can verify against physical intuition.</p>
<p>The research, led by Yufei Wang, Nan Li, Hui Shi, Gang Xie, and Xiaoyin Nie of the Shanxi Key Laboratory of Advanced Control and Industrial Intelligence at Taiyuan University of Science and Technology, addresses a challenge that has grown more urgent as renewable energy floods modern power grids. When wind and solar output fluctuates, coal-fired units are increasingly asked to ramp their load up and down, cycling through operating conditions far more varied than the steady baseload duty for which they were originally designed. That flexibility demand places new stress on the dynamic models used to predict and control how the plant responds, and inaccurate predictions can translate directly into safety risks and efficiency losses.</p>
<p>At the heart of the difficulty is the sheer complexity of a boiler-turbine system. It is a strongly nonlinear, multivariable machine in which fuel flow, feedwater, turbine valve positions, and air supply interact in tangled feedback loops. Worse still, different variables respond on different time scales: main steam pressure may react to a disturbance within seconds, while main steam temperature drifts slowly over minutes because of the thermal inertia of massive metal components. A single, one-size-fits-all neural network architecture struggles to capture this variable-specific, multi-scale character, and conventional data-driven models that do capture it often hide their internal logic behind millions of opaque parameters.</p>
<p>The team&#8217;s answer is a framework they call the Multi-dimensional Dynamic Attention Network, or MDANet. Rather than treating all inputs and all time scales identically, the architecture is built to mirror the physical structure of the plant it models. The first pillar of the design is adaptive receptive-field temporal feature extraction, which allows the network to learn, for each measured variable, how far back in time it needs to look. Fast-moving signals get short memory windows; sluggish, high-inertia signals get long ones. This directly encodes the multi-scale dynamic behavior that engineers know characterizes boiler-turbine systems, and it does so without hard-coding specific time constants, letting the data reveal the appropriate scales.</p>
<p>The second pillar, adaptive feature re-weighting, tackles the problem of shifting operating conditions. A coal-fired unit swinging between low-load and high-load operation does not obey a single fixed set of relationships; the relative importance of one input variable to an output can change as the plant moves through its operating envelope. MDANet&#8217;s re-weighting mechanism dynamically adjusts how much attention each variable receives, so the model remains faithful to the plant&#8217;s behavior across a wide range of conditions rather than overfitting to one narrow regime. This is the kind of condition-dependent sensitivity that classical linear models cannot represent and that naive deep networks learn only implicitly, buried where no engineer can inspect it.</p>
<p>The third pillar addresses the fact that a boiler-turbine unit must be modeled as a coupled multi-output system. Unit load, main steam pressure, and main steam temperature are not independent quantities; they are linked through shared energy and mass flows, and predicting them in isolation discards valuable information. To capture this coupling, the researchers employ a temporal context-aware multi-gate mixture-of-experts structure. In this design, multiple specialized expert sub-networks process the input, and learned gating networks decide, moment by moment and task by task, which experts should contribute to each output prediction. The result is a form of structured knowledge sharing: related outputs draw on overlapping expertise while still retaining the flexibility to specialize where their dynamics diverge.</p>
<p>Validation was carried out on real distributed control system data from a 600 MW coal-fired boiler-turbine unit, giving the evaluation an industrial realism that laboratory benchmarks often lack. The authors report that comparative experiments, ablation studies, and out-of-sample robustness analysis all demonstrate that MDANet achieves accurate and stable predictions for the three key outputs: unit load, main steam pressure, and main steam temperature. The ablation studies are particularly telling, because they show that removing any one of the three architectural pillars degrades performance, confirming that each component earns its place rather than adding decorative complexity.</p>
<p>What sets the work apart, however, is its insistence on interpretability as a first-class requirement rather than an afterthought. The researchers used attention-based visualization to expose which input variables and which time windows the network relied on when forming its predictions, and they supplemented these visual maps with quantitative interpretability analysis. The crucial finding is that the learned relevance patterns inside the model are consistent with the expected dynamic behavior of the boiler-turbine system. In other words, when the network predicts a change in main steam temperature, the attention patterns show it watching the inputs that thermodynamics says should matter, over the time horizons that the plant&#8217;s thermal inertia implies. The model&#8217;s internal logic and the engineer&#8217;s physical understanding converge.</p>
<p>This alignment between learned representations and physical expectation is precisely what the term physics-aligned is meant to convey. The framework does not embed explicit physical equations as hard constraints, a strategy used in some physics-informed neural networks. Instead, it shapes the architecture so that the network can naturally discover physically meaningful structure from data, and then verifies that the discovered structure matches reality. For plant engineers, this means a prediction can be audited: if the model&#8217;s attention points to a plausible causal chain, confidence rises; if it points somewhere nonsensical, the anomaly itself becomes a diagnostic signal. That auditability is the difference between a research curiosity and a tool that can be trusted in safety-critical operation.</p>
<p>The broader implications extend well beyond a single power station. As grids worldwide absorb ever-higher shares of variable renewable generation, the flexible operation of remaining thermal plants has become a linchpin of energy security, and transparent dynamic models are a prerequisite for advanced monitoring, fault detection, and control in that context. The Taiyuan team&#8217;s results suggest a practical middle path between the accuracy of black-box deep learning and the transparency of first-principles modeling: architectures whose inductive biases guide them toward physically sensible solutions, paired with interpretability tools that make the learned behavior visible. If such physics-aligned approaches generalize to other industrial processes, from chemical reactors to cement kilns, they could help dissolve the trust barrier that has long separated powerful machine learning from the engineers who operate the machines. The study, published open access with funding support from the National Natural Science Foundation of China and several Shanxi provincial programs, offers a concrete demonstration that in industrial artificial intelligence, seeing what the model sees may matter as much as how well it predicts.</p>
<p><strong>Subject of Research:</strong> Physics-aligned deep learning for interpretable dynamic modeling of coal-fired boiler-turbine systems</p>
<p><strong>Article Title:</strong> Interpretable dynamic modeling of coal-fired boiler–turbine systems: a physics-aligned deep learning approach</p>
<p><strong>Article References:</strong> Wang, Y., Li, N., Shi, H., Xie, G., &amp; Nie, X. (2026). Interpretable dynamic modeling of coal-fired boiler–turbine systems: a physics-aligned deep learning approach. <em>Complex &amp;amp; Intelligent Systems</em>. <a href="https://doi.org/10.1007/s40747-026-02479-x" rel="noopener noreferrer">https://doi.org/10.1007/s40747-026-02479-x</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s40747-026-02479-x" rel="noopener noreferrer">10.1007/s40747-026-02479-x</a></p>
<p><strong>Keywords:</strong> boiler-turbine system, dynamic modeling, deep learning, attention mechanism, interpretability, multi-task learning, mixture-of-experts, coal-fired power plant, renewable energy integration, industrial AI, time series prediction, control systems</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">226006</post-id>	</item>
		<item>
		<title>From CLIP to Llama 4: A sweeping review maps the rise of pretrained multimodal AI</title>
		<link>https://scienmag.com/from-clip-to-llama-4-a-sweeping-review-maps-the-rise-of-pretrained-multimodal-ai/</link>
		
		<dc:creator><![CDATA[Denise Maddox]]></dc:creator>
		<pubDate>Thu, 01 Oct 2026 20:07:58 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[advancements from CLIP to Llama 4]]></category>
		<category><![CDATA[AI applications in wildlife monitoring]]></category>
		<category><![CDATA[AI in autonomous transportation]]></category>
		<category><![CDATA[AI in industrial manufacturing]]></category>
		<category><![CDATA[BLIP]]></category>
		<category><![CDATA[CLIP]]></category>
		<category><![CDATA[cross-modal data interpretation]]></category>
		<category><![CDATA[deep learning models in healthcare]]></category>
		<category><![CDATA[Flamingo]]></category>
		<category><![CDATA[GPT-4o]]></category>
		<category><![CDATA[healthcare AI]]></category>
		<category><![CDATA[interdisciplinary AI model analysis]]></category>
		<category><![CDATA[Llama 4]]></category>
		<category><![CDATA[Mixture of Experts]]></category>
		<category><![CDATA[model hallucination]]></category>
		<category><![CDATA[multimodal deep learning]]></category>
		<category><![CDATA[multimodal model architectures]]></category>
		<category><![CDATA[multimodal training datasets and objectives]]></category>
		<category><![CDATA[open-access AI research review]]></category>
		<category><![CDATA[PaliGemma]]></category>
		<category><![CDATA[Pretrained multimodal AI]]></category>
		<category><![CDATA[systematic review of multimodal models]]></category>
		<category><![CDATA[Vision Transformers]]></category>
		<category><![CDATA[vision-language models]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=223590</guid>

					<description><![CDATA[A new comprehensive review analyzes 13 pretrained multimodal AI models from 2020 to 2025, comparing their architectures, datasets, and applications across 11 domains while exposing persistent challenges in data quality, computational cost, and reliability.]]></description>
										<content:encoded><![CDATA[<p>Artificial intelligence has quietly crossed a threshold. The most capable systems no longer see the world through a single sense: they read text, interpret images, parse audio, and watch video, all within one unified model. A new open-access review published in Discover Informatics by Azhar A. Hadi and K. P. Supreethi of Jawaharlal Nehru Technological University Hyderabad offers the most systematic map yet of this transformation, dissecting thirteen pretrained multimodal deep learning models released between 2020 and 2025 and tracing how they are reshaping fields from radiology to wildlife monitoring.</p>
<p>The review, which screened more than 400 candidate papers down to 120 rigorously assessed studies, arrives at a moment when the field is accelerating faster than most surveys can track. Earlier reviews, the authors argue, tended to focus narrowly on single domains such as medicine or language processing, leaving researchers without a comparative view of the architectures, datasets, and training objectives that define the current generation of models. By covering eleven application domains, including healthcare, education, autonomous transportation, e-commerce, and industrial manufacturing, the paper aims to give both newcomers and specialists a coherent picture of what these systems can actually do.</p>
<p>At the technical heart of the review lies a family of architectures built on the transformer, the attention-based design that revolutionized natural language processing before conquering vision. The Vision Transformer, or ViT, broke with convolutional tradition by slicing images into fixed 16-by-16-pixel patches, flattening them into vectors, and processing them exactly like words in a sentence. This simple reframing allowed vision models to inherit the scaling behavior of language models, and it became the visual backbone for much of what followed. CLIP, developed by OpenAI, went further by training an image encoder and a text encoder jointly on 400 million image-text pairs scraped from the internet, learning to pull matching pairs together in a shared latent space while pushing mismatched ones apart. That contrastive objective gave machines a rudimentary form of grounded understanding, and CLIP now serves as the visual front end for many multimodal large language models.</p>
<p>Subsequent generations refined this recipe in strikingly different directions. BLIP introduced a flexible encoder-decoder design and a bootstrapping mechanism called CapFilt, which generates synthetic captions and filters noisy training pairs to improve data quality. Its successor, BLIP-2, took a radically efficient path: rather than retraining everything, it froze both a pretrained image encoder and a large language model, connecting them through a lightweight Querying Transformer that translates visual features into a form the language model can consume. DeepMind&#8217;s Flamingo achieved few-shot multimodal learning by threading visual inputs into a frozen language model through a Perceiver Resampler and gated cross-attention layers, allowing it to answer visual questions from just a handful of examples. Google&#8217;s PaLI unified multilingual text generation via mT5 with high-capacity ViT encoders, while LLaVA pioneered visual instruction tuning, using GPT-4 to synthesize training dialogues and then aligning a CLIP vision encoder with a Vicuna language decoder.</p>
<p>The review also charts the frontier models that have captured public attention. GPT-4V extended OpenAI&#8217;s flagship language model to accept images, achieving human-level performance on many professional benchmarks while remaining weaker than humans on certain real-world tasks. GPT-4o, released in May 2024, went further still, processing text, images, audio, and video in a single network and becoming the first large language model to perform real-time emotion recognition from video. Florence-2, built on a Dual Attention Vision Transformer and trained on the FLD-5B dataset with 5.4 billion annotations across 126 million images, handles captioning, detection, segmentation, and grounding through one prompt-based sequence-to-sequence framework. SigLIP-2 replaced the conventional softmax with a sigmoid loss, decoupling batch size from training efficacy and supporting batches of up to one million examples while maintaining strong performance in more than 100 languages.</p>
<p>Efficiency has emerged as the defining engineering challenge, and the review documents how the newest models answer it. PaliGemma 2 pairs a SigLIP vision encoder with Gemma language models in sizes from 3 billion to 28 billion parameters, trained in three stages that progressively raise image resolution. Gemma 3 interleaves a single global attention layer with every five local layers using a restricted 1024-token sliding window, taming the memory demands of its 128,000-token context. Llama 4 adopts a sparse Mixture-of-Experts design in which a routing mechanism activates only a fraction of expert networks per token, expanding effective capacity while keeping inference cheap, and stretches context to roughly 10 million tokens for document-scale reasoning. These techniques, alongside parameter-efficient fine-tuning methods such as Low-Rank Adaptation, form what the authors call a crucial pathway to deployable multimodal AI.</p>
<p>The application evidence assembled in the review is striking in its breadth. In healthcare, fine-tuned combinations of vision transformers and language models reach up to 98.5 percent accuracy on lung disease diagnosis, while the PaliGemma-CXR system interprets tuberculosis chest X-rays with 90.32 percent accuracy. ChatIOS, which couples 3D point-cloud encoders with GPT-4V, achieved 93 percent intersection-over-union in automatic tooth segmentation from intraoral scans. Yet the same section catalogues sobering failures: GPT-4V identified anatomy correctly in 87.1 percent of radiology cases but pathology in only 35.2 percent, and diagnostic hallucination rates exceeding 40 percent were reported in some evaluations. In transportation, a fine-tuned PaliGemma model read license plates with 97.66 percent character accuracy, while GPT-4o managed 67 percent on pedestrian behavior prediction. In manufacturing, a CLIP-based defect classifier hit 99.9 percent AUROC with 6.6-millisecond inference, fast enough for real-time production lines.</p>
<p>Against these successes, the review is candid about the field&#8217;s structural weaknesses. Data remains the first bottleneck: many studies rely on small, single-center, or single-vendor datasets, annotations from lone experts, and corpora so noisy that models learn spurious correlations, the authors&#8217; memorable example being polar bears on ice. Computational cost is the second, with models scaling to 120 billion parameters and demanding specialized hardware that restricts access for smaller research groups. Reliability is the third: models hallucinate plausible but incorrect findings, struggle with composite figures and micro-expressions, misclassify small pedestrians at low pixel density, and remain vulnerable to adversarial attacks such as projected gradient descent. Ethical concerns compound the technical ones, since multimodal datasets encode societal biases that models can perpetuate in sensitive domains like healthcare and education, and training on medical images raises privacy stakes that demand techniques such as federated learning and differential privacy.</p>
<p>The authors close with a research agenda that reads as a diagnosis of the field&#8217;s growing pains. They call for multi-center, human-annotated datasets that pair imaging with genetic and clinical data; for encoder refinements that can handle long audio and raw high-resolution images directly; for larger but memory-efficient transformers; and for explainability to be treated as a core evaluation requirement rather than an afterthought. Techniques like model distillation, which compresses large multimodal systems into deployable smaller versions, and Mixture-of-Experts scaling are highlighted as the most promising routes to accessible systems. What emerges from the 120 studies is a field in transition: the architectural foundations are largely settled, the benchmarks are impressive, but trust, efficiency, and generalization to the messy diversity of real-world data remain the unfinished work. For researchers deciding where to invest the next five years, this review functions as both a map and a warning.</p>
<p><strong>Subject of Research:</strong> Pretrained multimodal deep learning models, their architectures, applications, and future research directions</p>
<p><strong>Article Title:</strong> A comprehensive review on recent pretrained multimodal deep learning models from architectures to future directions</p>
<p><strong>Article References:</strong> Hadi, A. A., &amp; Supreethi, K. P. (2026). A comprehensive review on recent pretrained multimodal deep learning models from architectures to future directions. <em>Discover Informatics, 1</em>(1), Article 21. <a href="https://doi.org/10.1007/s44564-026-00022-1" rel="noopener noreferrer">https://doi.org/10.1007/s44564-026-00022-1</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s44564-026-00022-1" rel="noopener noreferrer">10.1007/s44564-026-00022-1</a></p>
<p><strong>Keywords:</strong> multimodal deep learning, CLIP, vision transformers, GPT-4o, BLIP, Flamingo, PaliGemma, Llama 4, vision-language models, healthcare AI, Mixture-of-Experts, model hallucination</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">223590</post-id>	</item>
		<item>
		<title>New Map of AI Giants Sorts the World&#8217;s Language Models Into Six Powers</title>
		<link>https://scienmag.com/new-map-of-ai-giants-sorts-the-worlds-language-models-into-six-powers/</link>
		
		<dc:creator><![CDATA[Denise Maddox]]></dc:creator>
		<pubDate>Wed, 30 Sep 2026 22:53:03 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[AI capabilities classification]]></category>
		<category><![CDATA[AI ethics]]></category>
		<category><![CDATA[AI model capability mapping]]></category>
		<category><![CDATA[AI model performance evaluation]]></category>
		<category><![CDATA[comparison of proprietary and open-source language models]]></category>
		<category><![CDATA[comprehension abilities of language models]]></category>
		<category><![CDATA[emergence]]></category>
		<category><![CDATA[generative AI systems analysis]]></category>
		<category><![CDATA[hallucination]]></category>
		<category><![CDATA[information extraction in AI]]></category>
		<category><![CDATA[language model transformation skills]]></category>
		<category><![CDATA[large language models]]></category>
		<category><![CDATA[Large language models taxonomy]]></category>
		<category><![CDATA[Mixture of Experts]]></category>
		<category><![CDATA[multimodal AI]]></category>
		<category><![CDATA[NLP system functionality categories]]></category>
		<category><![CDATA[problem-solving in AI models]]></category>
		<category><![CDATA[reinforcement learning from human feedback]]></category>
		<category><![CDATA[retrieval-augmented generation]]></category>
		<category><![CDATA[scaling laws]]></category>
		<category><![CDATA[small language models]]></category>
		<category><![CDATA[systematic review of AI giants]]></category>
		<category><![CDATA[taxonomy]]></category>
		<category><![CDATA[transformers]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=219846</guid>

					<description><![CDATA[A new systematic review in Neural Computing and Applications maps the capabilities of modern large language models into six categories and proposes testable mathematical frameworks for how their abilities emerge.]]></description>
										<content:encoded><![CDATA[<p>Few technologies have climbed as fast as large language models, the systems behind ChatGPT, Claude, Gemini and their rivals. Yet for all the headlines, the field has lacked something basic: a shared map of what these machines can actually do. A new systematic review published in Neural Computing and Applications by Mohamed Nazih Omri of the University of Sousse and Borhen Louhichi of Imam Mohammad Ibn Saud Islamic University attempts to draw that map, sorting the sprawling LLM landscape into a rigorous taxonomy and backing it with comparative analysis of the leading proprietary and open-source systems, from GPT-4 and Gemini to Claude, LLaMA and DeepSeek.</p>
<p>The paper&#8217;s central organizational contribution is a six-part taxonomy of language model capabilities: Generation, Extraction, Classification, Transformation, Problem-solving and Comprehension. The categories are deliberately functional rather than technical. Generation covers the production of new text, code and media; Extraction pulls structured information out of unstructured documents; Classification assigns labels to inputs; Transformation rewrites content between formats, styles or languages; Problem-solving encompasses reasoning, planning and mathematical work; and Comprehension describes the models&#8217; ability to interpret meaning across long and complex contexts. The authors argue that this framework lets researchers compare models across vendors and scales on equal footing, instead of relying on marketing benchmarks that shift from release to release.</p>
<p>Beneath the taxonomy sits a more ambitious theoretical claim. The review introduces two formal constructs, an Emergence Function and an Architectural Efficiency Coefficient, intended to describe how capabilities appear as models scale and how efficiently a given architecture converts parameters and energy into performance. The Architectural Efficiency Coefficient is defined as performance divided by the product of parameters and energy, multiplied by the logarithm of throughput. In an illustrative calculation, the authors estimate a coefficient of roughly 2.30 for mixture-of-experts architectures, using the fact that Mixtral 8x7B reaches about 92 percent of GPT-4&#8217;s MMLU score while activating only 13 billion of its 47 billion parameters per token. Crucially, the authors are candid that these formulations are presented as testable hypotheses, not validated laws, and an appendix spells out exactly what would be needed to confirm them: controlled training runs, identical hardware, direct energy measurements and published replication code.</p>
<p>The survey traces the technical lineage that produced today&#8217;s giants. The transformer architecture, introduced in 2017 with the paper &#8216;Attention Is All You Need&#8217;, replaced the recurrent networks of the previous decade with self-attention mechanisms that process entire sequences in parallel. Scaling laws published in 2020 showed that model performance improves predictably with parameters, data and compute, setting off the race toward ever-larger systems. The review follows that arc through BERT and its bidirectional pretraining, the GPT series&#8217; few-shot learning, instruction tuning with human feedback, and the recent shift toward inference-time scaling, in which models such as OpenAI&#8217;s o1 and o3 spend more computation at answer time to chain together reasoning steps.</p>
<p>Training methodology receives its own systematic treatment. The authors describe the now-standard multi-stage pipeline: massive self-supervised pretraining on trillion-token corpora, followed by supervised fine-tuning on curated instruction data, followed by reinforcement learning from human feedback to align outputs with human preferences. They highlight curriculum learning, in which training data is ordered from simple to complex, and efficient architectural paradigms such as mixture-of-experts sparsity, which lets models grow in total size without a proportional rise in per-token cost. Techniques like low-rank adaptation and quantized fine-tuning, which allow large models to be customized on modest hardware, feature prominently as the field&#8217;s answer to the crushing economics of full retraining.</p>
<p>The empirical comparisons reveal a striking convergence. Proprietary flagships from OpenAI, Google DeepMind and Anthropic still lead many benchmarks, but open-weight models have closed much of the gap. DeepSeek&#8217;s R1 model, which uses reinforcement learning to incentivize reasoning, demonstrated that open systems can match frontier reasoning performance at a fraction of the training cost. Meta&#8217;s Llama family, Alibaba&#8217;s Qwen series, Mistral&#8217;s compact 7-billion-parameter model and the Falcon models from the Technology Innovation Institute show that the open ecosystem now spans everything from edge-deployable assistants to near-frontier generalists with 128,000-token context windows. The review frames this as a structural shift: capability is no longer the sole property of closed labs.</p>
<p>Applications receive equally broad coverage. In healthcare, models such as BioGPT and Med-PaLM have shown the ability to encode clinical knowledge, while synthetic medical text generation has improved diagnostic code classification accuracy by up to 17.8 percent in reported experiments. In education, AI tutors built on these models promise personalized instruction, though the authors note open questions about learning outcomes. In law, GPT-4-based systems are already used for document analysis. Creative domains, software development through Code Llama-style models, and autonomous agents that combine language models with tools, memory and planning round out the survey of real-world deployment.</p>
<p>The review is notably unsparing about the field&#8217;s problems. Hallucination, the confident generation of false statements, remains a fundamental weakness tied to the models&#8217; statistical nature. Bias embedded in training data propagates into outputs, and mitigation techniques are still immature. Computational cost and the carbon footprint of training large models raise sustainability concerns, even as analysts project that emissions may plateau and shrink with efficiency gains. Privacy risks persist, with research demonstrating that training data can be extracted from deployed models. The authors also cite the &#8216;stochastic parrots&#8217; critique, which questions whether scale alone can deliver genuine understanding, and they catalog jailbreaking and adversarial attacks that undermine safety guardrails. Regulatory frameworks, including the EU AI Act, are presented as an emerging constraint on deployment.</p>
<p>Looking forward, the survey identifies several trends it expects to shape the next phase. Small language models, distilled and compressed versions of their giant cousins, are positioned as the pragmatic future for many applications, offering most of the utility at a fraction of the cost. Multimodal integration, exemplified by Gemini&#8217;s native handling of text, images and audio and by embodied models such as PaLM-E, is dissolving the boundary between language and perception. Retrieval-augmented generation, which grounds model outputs in external documents at query time, offers a partial remedy for hallucination and stale knowledge. The authors also point to interpretability research, from mechanistic analyses of transformer circuits to attribution methods, as essential for building systems whose decisions can be trusted and audited.</p>
<p>What makes the paper unusual among surveys is its insistence on intellectual honesty about its own contributions. The mathematical framework for emergence and efficiency is offered as a hypothesis-generating scaffold, with the authors explicitly listing what they did not do: no model was trained from scratch, no energy consumption was measured directly, and no statistical significance testing was performed. That transparency, combined with the six-category taxonomy and the synthesis of architectural trade-offs, positions the review as both a reference work for practitioners navigating a crowded market and a starting point for the empirical studies that will determine whether the field&#8217;s scaling obsession gives way to something smarter: architectures that squeeze more capability out of every parameter and every joule.</p>
<p><strong>Subject of Research:</strong> A systematic taxonomy and empirical analysis of large language model architectures, training methods and applications</p>
<p><strong>Article Title:</strong> Decoding the Giants: a systematic taxonomy and empirical analysis of large language model architectures, training, and applications</p>
<p><strong>Article References:</strong> Omri, M. N., &amp; Louhichi, B. (2026). Decoding the Giants: a systematic taxonomy and empirical analysis of large language model architectures, training, and applications. <em>Neural Computing and Applications, 38</em>(19), Article 759. <a href="https://doi.org/10.1007/s00521-026-12365-9" rel="noopener noreferrer">https://doi.org/10.1007/s00521-026-12365-9</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s00521-026-12365-9" rel="noopener noreferrer">10.1007/s00521-026-12365-9</a></p>
<p><strong>Keywords:</strong> large language models, transformers, taxonomy, emergence, scaling laws, mixture of experts, reinforcement learning from human feedback, small language models, multimodal AI, retrieval-augmented generation, hallucination, AI ethics</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">219846</post-id>	</item>
		<item>
		<title>AI Meets Classic Search: New Two-Stage Algorithm Keeps Drone Testing on Schedule When Equipment Fails</title>
		<link>https://scienmag.com/ai-meets-classic-search-new-two-stage-algorithm-keeps-drone-testing-on-schedule-when-equipment-fails/</link>
		
		<dc:creator><![CDATA[Denise Maddox]]></dc:creator>
		<pubDate>Wed, 30 Sep 2026 19:12:17 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[AI-based industrial scheduling]]></category>
		<category><![CDATA[AI-driven test schedule optimization]]></category>
		<category><![CDATA[Artificial Intelligence]]></category>
		<category><![CDATA[combinatorial optimization]]></category>
		<category><![CDATA[dynamic rescheduling]]></category>
		<category><![CDATA[dynamic rescheduling in manufacturing]]></category>
		<category><![CDATA[equipment failure]]></category>
		<category><![CDATA[equipment failure mitigation]]></category>
		<category><![CDATA[group relative policy optimization]]></category>
		<category><![CDATA[hybrid reinforcement learning algorithms]]></category>
		<category><![CDATA[integration of classical search techniques with AI]]></category>
		<category><![CDATA[maintaining testing schedules during equipment failures]]></category>
		<category><![CDATA[makespan optimization]]></category>
		<category><![CDATA[metaheuristics]]></category>
		<category><![CDATA[Mixture of Experts]]></category>
		<category><![CDATA[real-time repair time estimation]]></category>
		<category><![CDATA[reinforcement learning]]></category>
		<category><![CDATA[resource-constrained multi-project scheduling]]></category>
		<category><![CDATA[resource-constrained project scheduling]]></category>
		<category><![CDATA[Tabu search]]></category>
		<category><![CDATA[time-critical industrial processes]]></category>
		<category><![CDATA[UAV test bed disruption management]]></category>
		<category><![CDATA[UAV testing]]></category>
		<category><![CDATA[Unmanned aerial vehicle testing]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=218426</guid>

					<description><![CDATA[Researchers in Beijing have developed a two-stage hybrid algorithm combining group relative policy optimization with Tabu search that reschedules disrupted UAV batch testing in under two seconds while keeping schedule overruns below 9.14 percent even under large repair-time uncertainties.]]></description>
										<content:encoded><![CDATA[<p>When a test bench fails in the middle of a batch of unmanned aerial vehicle trials, the consequences ripple far beyond the broken machine. Downstream validation tasks stall, technicians sit idle, and the entire testing calendar can slip by days or weeks. A research team in Beijing has now unveiled a hybrid artificial intelligence framework that promises to keep such disruptions from cascading, and the results suggest that a marriage between modern reinforcement learning and a decades-old search technique may be exactly what time-critical industrial scheduling has been waiting for.</p>
<p>The study, published in Applied Intelligence by Zhibin Mao, Yan Gao, Qian Pu, Minghui Wang and Haikuo Shen of Beijing Jiaotong University and the China Academy of Launch Vehicle Technology, tackles a problem that has long frustrated test facility managers: dynamic rescheduling under equipment failure. In their formulation, when a test item breaks down, an estimated repair time becomes available at the very moment of the disruption. The scheduler must then decide, almost instantly, how to reassign remaining tasks across constrained resources so that the overall completion time, known as the makespan, suffers as little as possible.</p>
<p>Mathematically, the researchers reformulated this recovery problem as a static resource-constrained multi-project scheduling problem, or RCMPSP, a notoriously hard combinatorial optimization challenge. In an RCMPSP, multiple projects compete for a limited pool of shared resources, and each task must respect precedence relations, meaning certain activities cannot begin until their predecessors finish. Finding an optimal schedule is computationally intractable for realistic instance sizes, which is why practitioners have historically relied on simple priority rules that assign tasks in a fixed order of importance, accepting suboptimal outcomes in exchange for speed.</p>
<p>The new method, dubbed GRPO-TS, departs from that tradition with a two-stage architecture. In the first stage, a policy network trained with Group Relative Policy Optimization, a reinforcement learning algorithm that has attracted wide attention for its efficiency in large language model training, generates an initial feasible schedule. Unlike conventional proximal policy optimization, GRPO evaluates groups of candidate actions relative to one another, which can stabilize learning and reduce the variance of policy updates. The authors adapted this idea to the scheduling domain, letting the network learn how to sequence competing test tasks under resource contention.</p>
<p>A key technical innovation lies in the architecture of the policy network itself. The researchers equipped it with a mixture-of-experts module, a design in which specialized subnetworks activate selectively depending on the input, allowing different experts to specialize in different scheduling regimes, such as periods of heavy resource competition versus periods dominated by precedence constraints. They also incorporated historical feature fusion, feeding the network information about past states so that it can capture how resource conflicts and task dependencies evolve over the course of a testing campaign. This gives the learned policy a form of temporal awareness that static priority rules fundamentally lack.</p>
<p>Yet a learned policy alone rarely produces a truly polished schedule. That is where the second stage comes in: Tabu Search, a classical metaheuristic introduced in the late 1980s, refines the initial solution through local neighborhood moves, systematically swapping and repositioning tasks while maintaining a tabu list that forbids recently revisited solutions to escape local optima. The division of labor is elegant. The neural policy supplies a high-quality starting point in milliseconds, and the metaheuristic polishes it with targeted local improvements, avoiding the wasteful random exploration that often makes pure metaheuristics slow on large instances.</p>
<p>The experimental evidence is striking. On extended RCMPSP benchmark instances, GRPO-TS achieved the lowest normalized average makespan among all tested baselines, including both classical priority-rule dispatching and modern deep reinforcement learning approaches. It also outperformed classical metaheuristics on medium and large instances, suggesting that the hybrid strategy scales better than either pure learning or pure search. Perhaps most importantly for real-world deployment, the online rescheduling time ranged from just 0.376 to 1.837 seconds, fast enough to re-plan a disrupted testing campaign before technicians have even finished diagnosing the failed equipment.</p>
<p>Robustness to uncertainty was another focus of the evaluation. Estimated repair times are, by definition, estimates, and real maintenance operations routinely deviate from predictions. The team stress-tested their framework by perturbing the estimated repair times by up to plus or minus thirty percent on mixed-scale instances. Even under these perturbations, the makespan increase did not exceed 9.14 percent, indicating that the rescheduled plans remain near-optimal even when the underlying assumptions about repair duration turn out to be substantially wrong. For test facilities where a single day of delay can cost significant sums, that kind of resilience is a meaningful guarantee.</p>
<p>The work sits within a broader and rapidly growing research movement that hybridizes reinforcement learning with evolutionary and local search methods. Recent surveys have documented a surge of such algorithms across domains from satellite scheduling to electric vehicle routing, on the logic that learned heuristics can guide classical optimizers toward promising regions of the search space while the optimizers supply the fine-grained refinement that neural networks struggle to achieve alone. The GRPO-TS framework is a particularly clean instantiation of this philosophy, and its application to UAV batch testing gives it a concrete industrial anchor rather than a purely academic benchmark.</p>
<p>The implications extend beyond drone testing. Resource-constrained multi-project scheduling arises in aircraft maintenance, construction, semiconductor fabrication, cloud computing and any setting where multiple concurrent workloads compete for scarce machines and personnel. The authors note that their data generator, parameter settings and source code are available from the corresponding author upon reasonable request, subject to institutional data-sharing policies, which should facilitate replication and adaptation by other groups. Supported by China&#8217;s National Key Research and Development Program, the research signals a future in which the moment a test rig fails, an intelligent scheduler quietly rebuilds the entire plan in under two seconds, and the production line barely notices. For an industry racing to certify fleets of autonomous aircraft, that future may arrive sooner than expected.</p>
<p><strong>Subject of Research:</strong> Dynamic rescheduling of resource-constrained UAV batch testing using a hybrid reinforcement learning and Tabu search framework</p>
<p><strong>Article Title:</strong> A two-stage framework integrating group relative policy optimization with Tabu search for dynamic rescheduling in UAV testing</p>
<p><strong>Article References:</strong> Mao, Z., Gao, Y., Pu, Q., Wang, M., &amp; Shen, H. (2026). A two-stage framework integrating group relative policy optimization with Tabu search for dynamic rescheduling in UAV testing. <em>Applied Intelligence, 56</em>(15), Article 457. <a href="https://doi.org/10.1007/s10489-026-07449-x" rel="noopener noreferrer">https://doi.org/10.1007/s10489-026-07449-x</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s10489-026-07449-x" rel="noopener noreferrer">10.1007/s10489-026-07449-x</a></p>
<p><strong>Keywords:</strong> UAV testing, dynamic rescheduling, group relative policy optimization, Tabu search, resource-constrained multi-project scheduling, reinforcement learning, metaheuristics, makespan optimization, mixture-of-experts, equipment failure, combinatorial optimization, artificial intelligence</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">218426</post-id>	</item>
		<item>
		<title>Helicopter Engine Digital Model Hits Near-Perfect Accuracy Despite Scarce Flight Data</title>
		<link>https://scienmag.com/helicopter-engine-digital-model-hits-near-perfect-accuracy-despite-scarce-flight-data/</link>
		
		<dc:creator><![CDATA[Grant Pearson]]></dc:creator>
		<pubDate>Sat, 26 Sep 2026 00:32:43 +0000</pubDate>
				<category><![CDATA[Space]]></category>
		<category><![CDATA[aero-engine control system optimization]]></category>
		<category><![CDATA[aerospace digital twin technology]]></category>
		<category><![CDATA[aerospace engineering]]></category>
		<category><![CDATA[aircraft engine fault detection]]></category>
		<category><![CDATA[data scarcity challenges in engine diagnostics]]></category>
		<category><![CDATA[digital twin]]></category>
		<category><![CDATA[fault diagnosis]]></category>
		<category><![CDATA[flight data calibration]]></category>
		<category><![CDATA[flight envelope adaptation in helicopter engines]]></category>
		<category><![CDATA[health monitoring]]></category>
		<category><![CDATA[Helicopter engine health monitoring]]></category>
		<category><![CDATA[hybrid modeling]]></category>
		<category><![CDATA[hybrid modeling strategies for aero-engines]]></category>
		<category><![CDATA[limited flight data in helicopter engine modeling]]></category>
		<category><![CDATA[LSTM]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[machine learning in aerospace safety]]></category>
		<category><![CDATA[Mixture of Experts]]></category>
		<category><![CDATA[onboard modeling]]></category>
		<category><![CDATA[predictive maintenance for rotorcraft engines]]></category>
		<category><![CDATA[real-time onboard engine modeling]]></category>
		<category><![CDATA[sensor faults]]></category>
		<category><![CDATA[thermodynamic behavior of turboshaft engines]]></category>
		<category><![CDATA[turboshaft engine]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=215651</guid>

					<description><![CDATA[Researchers have created a hybrid onboard model for turboshaft engines that combines component-level physics simulations with sparse flight data calibration, achieving near-perfect accuracy and sharp fault sensitivity across the full flight envelope.]]></description>
										<content:encoded><![CDATA[<p>Helicopter engines live complicated lives. Unlike the relatively steady cruise of an airliner, a turboshaft engine powering a rotorcraft must swing, within seconds, from a hover to a dash, from a autorotative descent to a combat climb, across altitudes and temperatures that reshape its thermodynamic behavior at every turn. Keeping an accurate mathematical portrait of such an engine running in real time aboard the aircraft is one of the enduring challenges of aero-engine health management, because that portrait, known as an onboard model, is what allows the engine&#8217;s control and monitoring systems to detect a degraded compressor, a drifting sensor, or an incipient fault before it becomes a catastrophe. A new study published in the International Journal of Aeronautical and Space Sciences by Weihong Huang, Qiangang Zheng, Cheng Chen, and Haibo Zhang of Nanjing University of Aeronautics and Astronautics now reports a hybrid modeling strategy that delivers near-perfect predictive accuracy across the entire flight envelope while consuming only a tiny fraction of the flight data such models were previously assumed to require.</p>
<p>The core dilemma the researchers confront is a data scarcity problem that will be familiar to anyone working at the intersection of machine learning and safety-critical engineering. Data-driven models are hungry for examples, yet genuine steady-state flight data from a helicopter turboshaft are extremely scarce and unevenly distributed in practice. Most of a flight is spent in transient maneuvers, and the calm, stabilized segments from which a steady-state model can learn represent only a small and biased slice of the operating envelope. Train a conventional neural network on what little is available and it will excel at predicting the conditions it has seen while silently failing everywhere else, a dangerous property for a model that is supposed to serve as the reference against which engine health is judged. Physics-based component-level simulations, meanwhile, cover the whole envelope but inevitably deviate from the specific engine actually hanging beneath the airframe, because manufacturing tolerances, installed accessories, inlet distortions, and sensor calibration all introduce systematic offsets that no generic simulation can anticipate.</p>
<p>The team&#8217;s solution is to let each kind of knowledge do the job it is best at. First, they construct a baseline steady-state model using the mixture-of-experts architecture, a machine learning design that traces back to a seminal 1991 paper by Jacobs, Jordan, Nowlan, and Hinton and has recently become famous as a scaling strategy for large language models. In a mixture-of-experts system, multiple specialist networks each handle a different region of the input space, and a gating network learns to blend or select their outputs. The researchers trained their expert networks on component-level steady-state grid data, which are cheap to generate in bulk from the physics simulation. The result is a global prior: a complete mapping from flight condition to engine outputs covering the entire envelope, capturing the broad physics correctly but carrying the systematic deviations inherent in any simulation of a real installed engine.</p>
<p>Then comes the calibration step, and this is where the engineering judgment of the approach shows. Rather than retraining the model wholesale on the scarce flight data, which would risk destroying the generalization bought with the simulation data, the authors introduce a staged fine-tuning strategy. The limited real flight steady-state measurements are used only to correct the systematic offset between the baseline model and the actual onboard environment. The training proceeds in two stages with progressively smaller learning rates, first at one tenth and then at one hundredth of the original pre-training rate, a discipline that nudges the model toward the true engine without letting it overfit the handful of calibration points. The hyperparameter details, published in the paper&#8217;s appendix, show careful use of the AdamW optimizer with weight decay and gradient clipping, plus early stopping to halt training the moment validation performance stops improving.</p>
<p>The quantitative gains reported are striking. After calibration, the root-mean-square errors of three key engine parameters, gas generator speed, compressor discharge pressure, and power turbine inlet temperature, were reduced by 98.18 percent, 98.30 percent, and 95.23 percent respectively. The coefficient of determination for these outputs, which had actually been negative before calibration, meaning the uncorrected baseline predicted worse than simply guessing the mean, rose to above 0.99 after fine-tuning. Crucially, the authors demonstrate that this dramatic improvement on flight data did not come at the price of full-envelope generalization; the calibrated model retained its accurate coverage of operating conditions far from the sparse calibration points. That combination, extreme accuracy on real data plus preserved breadth, is precisely the trade-off that has stymied previous attempts at onboard modeling.</p>
<p>Steady-state accuracy, however, is only half the problem. Helicopter engines spend much of their lives in transient operation, where thermal inertia, rotor dynamics, and fuel system lag mean the engine&#8217;s instantaneous state depends on its recent history, not just its current operating point. To capture this, the team built a dynamic model they call MEL, a temporally gated mixture-of-experts in which each expert pairs a steady-state multilayer perceptron with a stacked long short-term memory network. LSTM units are recurrent neurons with internal gates that let them carry information across time, making them well suited to learning how an engine relaxes toward its steady state after a fuel change. An independent LSTM-based gating network watches the temporal evolution of the inputs and dynamically assigns weights to the experts, so the model effectively learns which specialist to trust as the engine sweeps through a maneuver.</p>
<p>On a dynamic flight test set, MEL achieved a coefficient of determination of 0.9923, outperforming both a global LSTM model trained without the expert structure, which scored 0.9899, and the pure steady-state mixture-of-experts baseline, which managed 0.9764. The margins look small in raw percentage points, but in the world of engine monitoring, where residuals of a fraction of a percent can distinguish a healthy sensor from a failing one, they are meaningful. The architecture also inherits a virtue of modular design: because each expert specializes, the model can represent operating regimes with genuinely different dynamics, such as a rapid spool-up versus a slow power trim, without one network compromising its accuracy across all of them.</p>
<p>The most consequential test, for practical health monitoring, is whether the model notices when something goes wrong. The authors injected artificial sensor bias faults into the data, a 5 percent bias on compressor discharge pressure and a 2 percent bias on power turbine inlet temperature, and watched the residuals, the differences between measured and predicted values. The MEL model&#8217;s residual distribution exhibited clear and quantifiable shifts under both fault conditions, providing exactly the statistical signature an onboard fault detection system needs to flag an anomaly. The conventional steady-state mixture-of-experts model, by contrast, showed almost no response to the injected faults, a sobering demonstration that a model which fits the steady data poorly in the first place cannot serve as a sensitive diagnostic reference.</p>
<p>The broader significance of the work lies in what it suggests about the future of engine health management and digital twins. Physics-informed and hybrid modeling has been a growing theme across aerospace research, with recent efforts spanning physics-based analysis combined with machine learning for real-time performance modeling, physics-informed neural networks for digital twin condition monitoring, and adaptive transfer learning to cope with domain shifts between simulated and operational data. The Nanjing team&#8217;s contribution is a clean, well-quantified recipe for a specific and difficult case: the turboshaft engine, with its brutal transients and its chronic shortage of steady flight data. By anchoring the model in component-level physics and using flight data surgically, the method sidesteps the false choice between models that are accurate but narrow and models that are broad but biased.</p>
<p>There are, of course, the usual caveats that accompany any single study. The results rest on the quality of the component-level simulation used to generate the prior, on the representativeness of the available flight calibration data, and on fault injections rather than naturally occurring failures, and the paper&#8217;s training details, while thoroughly documented, describe a workflow that other operators would need to replicate on their own fleets and engine marks. Still, the headline numbers speak plainly: determination coefficients above 0.99 across the envelope, error reductions exceeding 95 percent on every key parameter, and a dynamic model that visibly flinches when a sensor lies. For helicopter operators, engine manufacturers, and the engineers building the next generation of onboard monitoring systems, the study offers a credible path to high-fidelity engine models that fit aboard the aircraft, run in real time, and learn from the data flights actually produce, which is to say, very little of it.</p>
<p><strong>Subject of Research:</strong> Hybrid onboard modeling of turboshaft engines using mixture-of-experts networks and flight data calibration for health monitoring</p>
<p><strong>Article Title:</strong> Hybrid Onboard Modeling of Turboshaft Engines Integrating Component-Level Priors with Flight Data Calibration</p>
<p><strong>Article References:</strong> Huang, W., Zheng, Q., Chen, C., &amp; Zhang, H. (2026). Hybrid Onboard Modeling of Turboshaft Engines Integrating Component-Level Priors with Flight Data Calibration. <em>International Journal of Aeronautical and Space Sciences</em>. <a href="https://doi.org/10.1007/s42405-026-01264-x" rel="noopener noreferrer">https://doi.org/10.1007/s42405-026-01264-x</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s42405-026-01264-x" rel="noopener noreferrer">10.1007/s42405-026-01264-x</a></p>
<p><strong>Keywords:</strong> turboshaft engine, onboard modeling, mixture-of-experts, LSTM, flight data calibration, fault diagnosis, health monitoring, digital twin, machine learning, aerospace engineering, sensor faults, hybrid modeling</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">215651</post-id>	</item>
		<item>
		<title>AI Learns to Read the Skies: Language Models Predict Air Traffic Complexity</title>
		<link>https://scienmag.com/ai-learns-to-read-the-skies-language-models-predict-air-traffic-complexity/</link>
		
		<dc:creator><![CDATA[Denise Maddox]]></dc:creator>
		<pubDate>Thu, 24 Sep 2026 01:29:23 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[AI in air traffic management]]></category>
		<category><![CDATA[air traffic complexity]]></category>
		<category><![CDATA[air traffic complexity prediction]]></category>
		<category><![CDATA[air traffic conflict anticipation]]></category>
		<category><![CDATA[air traffic control workload]]></category>
		<category><![CDATA[air traffic flow optimization]]></category>
		<category><![CDATA[air traffic management]]></category>
		<category><![CDATA[airspace management]]></category>
		<category><![CDATA[airspace sector interactions]]></category>
		<category><![CDATA[airspace sectors]]></category>
		<category><![CDATA[Applied Intelligence]]></category>
		<category><![CDATA[attention mechanism]]></category>
		<category><![CDATA[complex airspace route network]]></category>
		<category><![CDATA[flight path geometry analysis]]></category>
		<category><![CDATA[Graph Neural Networks]]></category>
		<category><![CDATA[large language models]]></category>
		<category><![CDATA[large language models for aviation]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[macro F1-score]]></category>
		<category><![CDATA[Mixture of Experts]]></category>
		<category><![CDATA[predictive modeling in aviation]]></category>
		<category><![CDATA[spatio-temporal prediction]]></category>
		<category><![CDATA[time-series forecasting]]></category>
		<category><![CDATA[weather impact on air traffic]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=211966</guid>

					<description><![CDATA[Researchers at Shanxi University have developed MAST-LLM, a framework that repurposes large language models with mixture-of-experts alignment and bidirectional spatio-temporal attention to predict airspace complexity, improving accuracy by up to 9.5 percentage points.]]></description>
										<content:encoded><![CDATA[<p>Air traffic controllers face one of the most demanding cognitive workloads of any profession: tracking dozens of aircraft converging through a shared volume of sky, each moving at hundreds of kilometers per hour, while anticipating conflicts minutes before they materialize. How difficult a given slice of airspace will be to manage, a quantity researchers call airspace complexity, is not simply a matter of counting planes. It emerges from the geometry of flight paths, the structure of the underlying route network, weather, flow restrictions, and the shifting interactions among all of these. Predicting that complexity ahead of time is a central goal of modern air traffic management, because accurate forecasts allow controllers and flow managers to rebalance traffic before sectors become overloaded. A new study published in Applied Intelligence by Rui Cheng, Jianping Fan, Chao Zhang, Meiqin Wu, Ruixin Chen, Mingxuan Chai, and Anna Wang of Shanxi University introduces a framework called MAST-LLM that attacks this prediction problem with an unusual tool: a large language model, repurposed to reason about the dynamics of the sky.</p>
<p>The challenge that motivated the work is structural. Airspace is divided into sectors, and each sector&#8217;s complexity depends on two kinds of relationships that are difficult to model together. Spatially, sectors are connected both by simple geographical adjacency, meaning they share a border, and by dynamic traffic flows, since aircraft entering one sector typically exited another, creating dependencies that shift with the daily rhythm of departures and arrivals. Temporally, complexity evolves across multiple time scales at once: fast fluctuations driven by individual climbing and descending aircraft, medium-term waves tied to airport scheduling, and long-range dependencies that link morning traffic management decisions to afternoon congestion. Existing spatio-temporal prediction frameworks, including the graph neural networks that have dominated traffic forecasting in recent years, tend to rely on static graphs that freeze these relationships in place and on shallow temporal alignment that struggles to connect patterns separated by long gaps in time. The result is a systematic weakness precisely where forecasters need the most help: high-complexity situations and long prediction horizons.</p>
<p>MAST-LLM, which stands for a Multimodal Adaptive Spatio-Temporal framework enhanced by Large Language Models, addresses these weaknesses in three coordinated stages. The first stage is a temporal alignment phase built on a Mixture-of-Experts architecture. A Mixture-of-Experts model is a neural network design in which multiple specialized subnetworks, the experts, each process the input, and a gating mechanism learns to weight their contributions depending on the character of the data at hand. In MAST-LLM, this design lets the framework capture heterogeneous temporal patterns, so that the distinct rhythms of different air traffic variables can each be handled by appropriately specialized experts. Critically, this phase also aligns domain-specific air traffic sequences with the representations that a pre-trained language model has already learned. Rather than forcing raw aviation data into a model built for text, the framework translates traffic dynamics into a form that the language model&#8217;s internal representations can meaningfully encode, bridging the gap between two very different data worlds.</p>
<p>The second stage is a full spatio-temporal fine-tuning phase, in which the framework integrates adaptive multimodal spatial representations with multi-scale temporal features. The word multimodal here refers to the combination of different kinds of information about the airspace, such as structural spatial relationships and dynamic traffic measurements, into a shared representation that can adapt as conditions change rather than remaining fixed. The centerpiece of this stage is a mechanism the authors call Bidirectional Spatio-Temporal Attention, or BSTA. Attention mechanisms, first popularized by the transformer architecture underlying modern language models, allow a network to weigh the relevance of every element of its input against every other element. BSTA extends this idea so that spatial and temporal information interact in both directions: spatial structure informs how temporal patterns are interpreted, and temporal evolution informs how spatial relationships are weighted. This joint interaction modeling is what allows the framework to reason about the airspace as a single coupled system rather than as separate spatial and temporal problems stitched together.</p>
<p>The third stage exploits the property that makes large language models attractive for this task in the first place: their capacity for global reasoning and long-range dependency modeling. Language models are trained on sequences in which meaning can depend on context established thousands of tokens earlier, and their architectures are built to preserve and use that distant context. By establishing a unified representation that captures both local dynamics, the minute-to-minute behavior of individual sectors, and global contextual dependencies, the patterns that propagate across an entire region&#8217;s airspace over hours, MAST-LLM inherits this long-context strength. The authors argue that this is precisely what earlier frameworks lacked: a way for a prediction about one sector at one moment to draw on evidence from distant sectors and distant times within a single coherent computation.</p>
<p>The empirical results reported in the paper are striking. In extensive experiments, MAST-LLM achieved superior performance at short- and medium-term forecasting horizons and remained competitive at longer horizons, with improvements of up to 9.5 percentage points in accuracy and 12.9 percentage points in macro F1-score over the strongest baselines. The macro F1-score is a particularly meaningful metric here because it averages the F1-score, which balances precision and recall, across all complexity classes, ensuring that improvements on rare but dangerous high-complexity conditions count as much as improvements on routine ones. The authors supplemented the headline numbers with comprehensive ablation studies, which remove individual components to verify that each one contributes, along with factor importance analyses, sensitivity analyses, and visualization analyses that together confirm the effectiveness and robustness of every core element of the design. Statistical significance was assessed with paired two-sided t-tests against the strongest baseline, with exact p-values reported in the paper&#8217;s appendix.</p>
<p>The work builds on a substantial lineage of research into both airspace complexity and spatio-temporal machine learning. Measures of air traffic complexity stretch back decades, from early workload prediction studies by Chatterji and Sridhar to probabilistic complexity measures in three-dimensional airspace developed by Prandini and colleagues, and to spatiotemporal graph indicators proposed by Isufaj and collaborators. On the machine learning side, the framework draws on the spatio-temporal graph convolutional networks introduced by Yu, Yin, and Zhu in 2018 and the diffusion convolutional recurrent networks of Li and colleagues from the same year, as well as more recent attention-based architectures such as GMAN and PDFormer. Notably, the same research group had previously developed MAST-GNN, a multimodal adaptive spatio-temporal graph neural network for the same prediction task, and MAST-LLM can be seen as an evolution of that line of work, replacing static graph reasoning with the adaptive, language-model-driven approach.</p>
<p>The study also sits within a rapidly growing movement to apply large language models to time-series and traffic problems. Recent research has shown that pre-trained language models can be reprogrammed for general time-series analysis, as in the One Fits All work of Zhou and colleagues, and for dedicated forecasting frameworks such as Time-LLM by Jin and colleagues and LLM4TS by Chang and colleagues. Parallel efforts have applied these ideas to wind power and wind speed forecasting with BERT4ST, STELLM, and STCA-LLM, to spatio-temporal imputation with GATGPT, and to urban traffic prediction with ST-LLM plus. A 2026 survey by Long and colleagues in IEEE Transactions on Big Data catalogues the accelerating adoption of language models across traffic forecasting applications. MAST-LLM distinguishes itself within this crowded field by combining the Mixture-of-Experts alignment strategy with bidirectional spatio-temporal attention, a pairing the authors present as tailored to the specific structure of airspace dynamics rather than borrowed wholesale from other domains.</p>
<p>The practical implications could be considerable. Accurate complexity forecasts feed directly into demand-capacity balancing, the process by which air navigation service providers decide how much traffic each sector can safely absorb and where flow restrictions should be imposed. Better predictions, especially at medium and long horizons, give managers more lead time to reroute flights, adjust sector configurations, and staff control positions appropriately, potentially reducing both delays and controller overload. The authors have made the airspace complexity dataset used in the study publicly available through a project repository, lowering the barrier for other groups to build on the approach. The research was supported by funders including the National Natural Science Foundation of China and the Ministry of Education of China, and the authors note that generative AI tools were used to assist with language refinement of the manuscript, with all content reviewed and verified by the team.</p>
<p>For the field of air traffic management, the study offers a proof of concept that the reasoning machinery of large language models, originally built for human language, can be redirected toward the physical dynamics of the sky. For the broader machine learning community, it adds to mounting evidence that pre-trained language models serve as powerful general-purpose sequence models whose learned representations transfer far beyond text. Whether such frameworks can be deployed in the safety-critical, certification-heavy environment of real air traffic control remains an open question, and the authors&#8217; results, while strong, come from benchmark evaluation rather than live operations. Still, as global air traffic continues to grow toward and beyond pre-pandemic levels, tools that can anticipate when a sector is about to become unmanageable, minutes or hours in advance, address one of the most consequential prediction problems in transportation, and MAST-LLM demonstrates that the newest generation of AI models may be up to the task.</p>
<p><strong>Subject of Research:</strong> Large language model-driven multimodal spatio-temporal prediction of air traffic complexity</p>
<p><strong>Article Title:</strong> MAST-LLM: A large language model-driven multimodal adaptive spatio-temporal framework for air traffic complexity prediction</p>
<p><strong>Article References:</strong> Cheng, R., Fan, J., Zhang, C., Wu, M., Chen, R., Chai, M., &amp; Wang, A. (2026). MAST-LLM: A large language model-driven multimodal adaptive spatio-temporal framework for air traffic complexity prediction. <em>Applied Intelligence, 56</em>(15), Article 442. <a href="https://doi.org/10.1007/s10489-026-07469-7" rel="noopener noreferrer">https://doi.org/10.1007/s10489-026-07469-7</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s10489-026-07469-7" rel="noopener noreferrer">10.1007/s10489-026-07469-7</a></p>
<p><strong>Keywords:</strong> air traffic complexity, large language models, spatio-temporal prediction, mixture of experts, attention mechanism, graph neural networks, air traffic management, time series forecasting, Applied Intelligence, machine learning, airspace sectors, macro F1-score</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">211966</post-id>	</item>
		<item>
		<title>AI Learns to Choreograph Emotion: New Model Matches Dance to Music&#8217;s Mood</title>
		<link>https://scienmag.com/ai-learns-to-choreograph-emotion-new-model-matches-dance-to-musics-mood/</link>
		
		<dc:creator><![CDATA[Blake Davidson]]></dc:creator>
		<pubDate>Tue, 22 Sep 2026 14:28:02 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[3D dance generation]]></category>
		<category><![CDATA[affective computing]]></category>
		<category><![CDATA[affective computing in dance synthesis]]></category>
		<category><![CDATA[AI-driven dance choreography]]></category>
		<category><![CDATA[benchmarks for emotion-synchronized dance]]></category>
		<category><![CDATA[continuous emotion dimensions in artificial intelligence]]></category>
		<category><![CDATA[continuous emotion modeling]]></category>
		<category><![CDATA[cross-modal alignment]]></category>
		<category><![CDATA[diffusion models]]></category>
		<category><![CDATA[diffusion-based emotion modeling in AI]]></category>
		<category><![CDATA[digital choreography]]></category>
		<category><![CDATA[emotion consistency]]></category>
		<category><![CDATA[emotion-aware 3D dance animation]]></category>
		<category><![CDATA[generative AI]]></category>
		<category><![CDATA[human-like emotional response in AI dance]]></category>
		<category><![CDATA[innovative AI frameworks for artistic expression]]></category>
		<category><![CDATA[limitations of traditional music-driven dance systems]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[machine learning for emotionally responsive choreography]]></category>
		<category><![CDATA[Mixture of Experts]]></category>
		<category><![CDATA[music and emotion integration in dance generation]]></category>
		<category><![CDATA[music-driven motion synthesis]]></category>
		<category><![CDATA[perceptual testing of AI dance realism]]></category>
		<category><![CDATA[valence–arousal routing]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=205875</guid>

					<description><![CDATA[A new diffusion-based framework called AffectMoE uses continuous valence–arousal soft routing over emotion-specific experts to generate 3D dance movements that stay emotionally consistent with the driving music.]]></description>
										<content:encoded><![CDATA[<p>When a piece of music swells with joy or sinks into melancholy, human dancers respond instinctively, translating emotional nuance into the curve of an arm or the weight of a step. Teaching a machine to do the same has proven far harder than simply making a digital figure move in time with a beat. Now, a new study published in Complex &amp; Intelligent Systems introduces AffectMoE, a diffusion-based artificial intelligence framework that treats emotion not as a superficial label attached to generated movement, but as a continuous, physically consequential dimension woven through the entire motion synthesis process. The result is 3D dance animation that researchers say stays emotionally faithful to the music driving it, with measurable gains over existing methods on standard benchmarks and in human perceptual testing.</p>
<p>The work, authored by Rui Zhang of the Conservatory of Music and Huai Opera Academy at Yancheng Teachers University in China, addresses a persistent weakness in music-driven 3D dance generation. Prevailing systems, the paper argues, treat emotion as a shallow auxiliary condition, bolted onto a pipeline that is fundamentally optimized for rhythm and beat alignment. The consequence is a familiar failure mode: a digital dancer that hits every downbeat yet moves with an emotional flatness that contradicts the music, or worse, produces gestures whose affective character clashes with what listeners hear. For applications ranging from game development and virtual performance to rehabilitation and entertainment, that mismatch undermines the believability of the entire output.</p>
<p>AffectMoE&#8217;s central innovation is an architectural one: a continuous valence–arousal soft routing mechanism applied over a mixture of experts. The valence–arousal model, a cornerstone of affective science, represents emotion as a point in a two-dimensional space, where valence tracks the pleasantness of a feeling from negative to positive and arousal tracks its intensity from calm to excited. Rather than forcing the system to choose among discrete emotional categories such as happy, sad, or angry, AffectMoE distributes a set of learnable emotion prototype experts across this continuous two-dimensional space. Each expert specializes in motion dynamics characteristic of its region of the emotional landscape, while an additional pool of emotion-agnostic experts handles general movement quality that does not depend on mood.</p>
<p>The routing itself is where the framework departs most sharply from prior work. A learned router observes the emotional content of the input music, mapped onto the valence–arousal plane, and produces convex combination weights over the experts. Instead of a hard assignment, in which one expert would be selected and the others ignored, the soft routing scheme blends the contributions of many experts simultaneously, with weights that shift smoothly as the music&#8217;s emotional character drifts. This enables interpolation of motion dynamics along continuous affective dimensions: a piece of music sitting between serene and exuberant produces movement that blends the motor vocabulary of both regions rather than snapping abruptly from one style to another. In practical terms, the generated dancer glides through emotional transitions the way a human performer would, rather than lurching between canned moods.</p>
<p>Keeping such a mixture-of-experts system healthy requires care. A well-known pathology in expert-based architectures is expert collapse, in which the router learns to funnel nearly all inputs to a small subset of specialists, leaving the rest untrained and wasted. AffectMoE counters this with a weak load balancing strategy, gentle enough to preserve the natural specialization of experts but firm enough to keep every prototype in productive use. On top of this, the framework integrates two additional training objectives: cross-modal contrastive alignment, which pulls the representations of music and matching motion closer together in a shared embedding space, and emotion consistency regularization, which penalizes generated movements whose emotional character diverges from the music that prompted them.</p>
<p>The architecture sits atop a diffusion model, a class of generative systems that has transformed machine learning in recent years by learning to reverse a gradual noising process, producing coherent samples from structured noise. In this context, the diffusion backbone generates sequences of 3D skeletal motion conditioned on the input music, while the emotion-aware routing shapes how the denoising trajectory unfolds, steering intermediate representations toward experts whose specialties match the prevailing affective state. The combination means that emotional control operates at every step of generation rather than being imposed as a post-hoc filter, which the author identifies as essential to achieving genuine emotional coherence in the final animation.</p>
<p>Quantitatively, the framework delivers strong results on two widely used datasets: FineDance and AIST++, both containing music paired with human 3D motion capture. AffectMoE achieves a FIDk score of 47.8, a Fréchet distance-style metric that measures how closely the distribution of generated motions matches that of real human dancing, with improvements the study reports as statistically significant. More novel is the proposed Emotion Consistency Score, an evaluation designed to quantify whether generated movement actually reflects the intended affect. Under this metric, the system reaches a quadrant accuracy of 73 percent, meaning the emotional character of generated dances lands in the correct quadrant of the valence–arousal plane nearly three-quarters of the time, a substantial feat given how subtle the mapping between sound and gesture can be.</p>
<p>One of the study&#8217;s most striking demonstrations involves controllability. When the system is run in override mode, where an operator manually specifies a target point in the valence–arousal space, continuous soft routing produces monotonic controllability curves: as the specified emotion is pushed progressively further along a dimension, the generated movement shifts correspondingly and consistently in the same direction. The paper notes that single-module emotional alignment approaches do not exhibit this behavior under the same setting, suggesting that the distributed expert architecture is not merely producing plausible-looking motion but genuinely responds to affective inputs in an orderly, interpretable way. For creators of interactive media, that kind of smooth, predictable control could prove as valuable as raw generation quality.</p>
<p>Human judgment, however, remains the ultimate test of emotional expressiveness, and the study includes a perceptual experiment with 30 participants. Conducted with informed consent and reported in aggregated, anonymized form, the user study indicated that viewers perceived higher emotional congruence between music and movement for AffectMoE&#8217;s outputs compared with alternatives. The finding complements the quantitative benchmarks: metrics can capture distributional fidelity and classification accuracy, but only human observers can judge whether a digital dancer genuinely feels the music, and on that criterion the participants sided with the new approach. The author declares no relevant financial or non-financial conflicts of interest, and the work was supported by the Research Fund Filing Project of the Heilongjiang Provincial Department of Education under a project exploring the inheritance of Oroqen facial expression art, an interesting cultural thread connecting computational emotion modeling with the preservation of traditional expressive forms.</p>
<p>The implications stretch beyond the laboratory. Music-driven dance generation is becoming a genuine industrial concern, powering virtual idols, game choreography, short-form video content, and digital doubles of human performers. A system that respects the emotional arc of a soundtrack, rather than merely its tempo, opens the door to generated performances that feel authored rather than assembled. Equally significant is the methodological lesson: emotion, long treated as a categorical afterthought in generative modeling, may be better handled as a continuous control dimension with dedicated capacity distributed across a model&#8217;s architecture. If the mixture-of-experts pattern proves transferable, the soft routing of affect could influence a broader family of generative systems, from expressive virtual agents to emotion-aware animation tools. As AI-generated performance matures, the gap between technically correct motion and emotionally convincing motion is exactly where the next frontier lies, and AffectMoE offers a concrete, empirically validated map of how to cross it, one smoothly interpolated feeling at a time.</p>
<p><strong>Subject of Research:</strong> Music-driven 3D dance generation with continuous emotion modeling using mixture-of-experts routing</p>
<p><strong>Article Title:</strong> AffectMoE: continuous valence–arousal soft routing of mixture of experts for emotion-consistent 3D dance generation</p>
<p><strong>Article References:</strong> Zhang, R. (2026). AffectMoE: continuous valence–arousal soft routing of mixture of experts for emotion-consistent 3D dance generation. <em>Complex &amp;amp; Intelligent Systems</em>. <a href="https://doi.org/10.1007/s40747-026-02490-2" rel="noopener noreferrer">https://doi.org/10.1007/s40747-026-02490-2</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s40747-026-02490-2" rel="noopener noreferrer">10.1007/s40747-026-02490-2</a></p>
<p><strong>Keywords:</strong> 3D dance generation, mixture of experts, continuous emotion modeling, valence–arousal routing, diffusion models, affective computing, music-driven motion synthesis, emotion consistency, cross-modal alignment, machine learning, digital choreography, generative AI</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">205875</post-id>	</item>
		<item>
		<title>Mixture-of-Experts AI Model Brings Fine-Grained Vision to High-Resolution Satellite Imagery</title>
		<link>https://scienmag.com/mixture-of-experts-ai-model-brings-fine-grained-vision-to-high-resolution-satellite-imagery/</link>
		
		<dc:creator><![CDATA[Blake Davidson]]></dc:creator>
		<pubDate>Sun, 20 Sep 2026 20:58:20 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[advanced remote sensing AI techniques]]></category>
		<category><![CDATA[Artificial Intelligence]]></category>
		<category><![CDATA[automated aerial image interpretation]]></category>
		<category><![CDATA[conditional computation]]></category>
		<category><![CDATA[deep learning]]></category>
		<category><![CDATA[Earth observation]]></category>
		<category><![CDATA[efficient AI routing for remote sensing]]></category>
		<category><![CDATA[fine-grained interpretation]]></category>
		<category><![CDATA[fine-grained remote sensing image understanding]]></category>
		<category><![CDATA[high-resolution imagery]]></category>
		<category><![CDATA[high-resolution satellite imagery analysis]]></category>
		<category><![CDATA[image captioning]]></category>
		<category><![CDATA[intelligent routing in machine learning]]></category>
		<category><![CDATA[Mixture of Experts]]></category>
		<category><![CDATA[mixture-of-experts neural network architecture]]></category>
		<category><![CDATA[remote sensing]]></category>
		<category><![CDATA[remote sensing interpretation]]></category>
		<category><![CDATA[satellite imagery]]></category>
		<category><![CDATA[satellite imagery scene classification]]></category>
		<category><![CDATA[scalable neural networks for high-res images]]></category>
		<category><![CDATA[specialized subnetworks in AI models]]></category>
		<category><![CDATA[vision-language model]]></category>
		<category><![CDATA[vision-language models for satellite data]]></category>
		<category><![CDATA[visual question answering]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=202264</guid>

					<description><![CDATA[A new mixture-of-experts vision-language model called RSFG-MoE is designed to deliver fine-grained, high-resolution interpretation of remote sensing imagery.]]></description>
										<content:encoded><![CDATA[<p>Remote sensing has always promised a god&#8217;s-eye view of the planet, but turning that torrent of pixels into usable knowledge has remained stubbornly difficult. Satellite and aerial imagery now arrives at resolutions fine enough to distinguish individual rooftops, road markings, tree crowns and ships at berth, yet most automated interpretation systems still describe scenes in coarse, generic terms. A new study published in Scientific Reports introduces RSFG-MoE, a vision-language model built around a mixture-of-experts architecture and designed specifically for fine-grained interpretation of high-resolution remote sensing images, and it argues that the missing ingredient is not raw computing power but a smarter way of routing specialized knowledge to the right kind of visual question.</p>
<p>The core idea behind RSFG-MoE is borrowed from a technique that has reshaped modern artificial intelligence: the mixture of experts. Instead of forcing one enormous neural network to learn everything at once, a mixture-of-experts model contains many specialized subnetworks, the experts, alongside a learned dispatcher called a router or gate. For every input, the gate examines the data and activates only a small subset of experts, allowing the model to grow its total capacity without a proportional rise in the computational cost of each prediction. In practical terms, the system can hold far more knowledge than an equally expensive dense model, because only the relevant specialists fire for any given image-text pair.</p>
<p>Why does this matter for satellite imagery? Because fine-grained remote sensing interpretation is an unusually heterogeneous task. A single scene might require recognizing the difference between similar aircraft models on an airfield, distinguishing crop types by subtle texture and phenology, counting ships in a crowded harbor, describing building footprints, or answering open-ended questions about land use. A dense model tends to average across all these demands, blurring the fine distinctions that specialists would catch. By contrast, the mixture-of-experts design lets different regions of the network quietly specialize, so the experts that excel at aircraft recognition are not the same ones burdened with crop classification, and the router learns to send each query down the most informative path.</p>
<p>RSFG-MoE pairs this architecture with a vision-language framework, meaning it processes images and natural language together. On one side, a visual encoder ingests high-resolution remote sensing scenes, typically by breaking them into patches and converting them into rich visual features. On the other side, a language component handles instructions, questions, captions and answers. The two streams are fused so the model can perform tasks such as visual question answering, image captioning and fine-grained scene description, all framed around aerial and orbital imagery rather than the ground-level photographs that dominate mainstream AI research.</p>
<p>That distinction is more consequential than it may sound. Most vision-language models are trained on web-scale datasets of everyday photos, where objects are viewed from human eye level at modest resolutions. Remote sensing images obey different rules: scale varies dramatically with sensor altitude, objects appear in arbitrary orientations, shadows and illumination shift with the sun angle, and the same location can look entirely different across seasons. A car photographed from the street and the same car seen from six hundred kilometers overhead share a name but almost nothing in visual appearance. RSFG-MoE is oriented to that aerial perspective from the ground up, which is precisely what fine-grained interpretation demands.</p>
<p>The fine-grained angle is the model&#8217;s defining ambition. Where conventional remote sensing captioning might say an image shows an airport, a fine-grained system should be able to specify how many aircraft are parked at gates, what types they appear to be, and how runways and taxiways are arranged. Where a generic model answers that a coastal image contains water and land, a fine-grained one can distinguish marinas, breakwaters, aquaculture pens and sandy beaches. Achieving this requires the model to preserve and exploit subtle local details, which is exactly where the specialized experts and the high-resolution visual encoding are meant to pay off.</p>
<p>The training strategy behind a model like RSFG-MoE typically combines several streams of supervision. Large-scale image-text alignment teaches the network to connect visual features with language at a general level. Task-specific datasets sharpen performance on remote sensing benchmarks such as captioning and visual question answering. Instruction tuning, a technique popularized in the era of large language models, teaches the system to follow varied natural-language commands, so that a single trained model can be steered toward counting, describing, comparing or reasoning without being retrained from scratch. The mixture-of-experts structure adds its own training considerations, including load-balancing objectives that prevent the router from collapsing onto a few favorite experts and leaving the rest unused.</p>
<p>For readers wondering why efficiency is mentioned so often in connection with this work, the arithmetic of high-resolution imagery explains it. A scene captured at sub-meter resolution can contain thousands of times more pixels than a typical internet photo, and processing every patch with a full dense model quickly becomes prohibitively expensive, whether the analysis serves disaster response, urban planning, agriculture or environmental monitoring. Conditional computation, of which mixture-of-experts is the leading example, offers a way out: total model capacity can scale up while the cost of each individual inference stays manageable, because only a fraction of the parameters are consulted for any given input. That trade-off between capacity and cost is arguably the central engineering question of contemporary AI, and RSFG-MoE applies it directly to Earth observation.</p>
<p>The potential applications extend well beyond academic benchmarks. Fine-grained automated interpretation could help emergency teams assess earthquake or flood damage scene by scene, guide precision agriculture by tracking crop health at the field level, support customs and port authorities in monitoring vessel traffic, and give city planners up-to-date inventories of buildings and infrastructure. It could also aid climate science, where tracking deforestation, glacier retreat and land-cover change requires consistent, detailed description of imagery collected over years and decades. A model that can be asked, in plain language, targeted questions about what a satellite image contains lowers the barrier between raw data and the people who need answers.</p>
<p>Challenges remain, as they always do in this fast-moving field. Mixture-of-experts models are notoriously demanding to train stably, and their gains depend on datasets that actually reward specialization. Remote sensing benchmarks are still smaller and less diverse than their ground-level counterparts, and questions of domain shift, sensor variation and evaluation rigor are actively debated. Yet the direction signaled by RSFG-MoE is clear: the next generation of Earth-observation AI will not simply be bigger, it will be better organized, with specialized knowledge held in reserve and summoned only when a particular patch of the planet, or a particular question about it, demands that expertise. As constellations of imaging satellites multiply and the resolution of commercial imagery continues to climb, models that can read the fine print of the planet&#8217;s surface are likely to move from research papers into daily operational use.</p>
<p><strong>Subject of Research:</strong> A mixture-of-experts vision-language model for fine-grained high-resolution remote sensing image interpretation</p>
<p><strong>Article Title:</strong> RSFG-MoE: a mixture-of-experts vision-language model oriented to fine-grained high-resolution remote sensing image interpretation</p>
<p><strong>Article References:</strong> Li, L., Wang, T., Zhang, Y., &amp; Zhang, N. (2026). RSFG-MoE: a mixture-of-experts vision-language model oriented to fine-grained high-resolution remote sensing image interpretation. <em>Scientific Reports</em>. <a href="https://doi.org/10.1038/s41598-026-70081-9" rel="noopener noreferrer">https://doi.org/10.1038/s41598-026-70081-9</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1038/s41598-026-70081-9" rel="noopener noreferrer">10.1038/s41598-026-70081-9</a></p>
<p><strong>Keywords:</strong> remote sensing, mixture-of-experts, vision-language model, high-resolution imagery, artificial intelligence, fine-grained interpretation, Earth observation, deep learning, image captioning, visual question answering, satellite imagery, conditional computation</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">202264</post-id>	</item>
		<item>
		<title>Blood Test Algorithm Cuts Unnecessary Alzheimer&#8217;s PET Scans in Clinical Trial Screening</title>
		<link>https://scienmag.com/blood-test-algorithm-cuts-unnecessary-alzheimers-pet-scans-in-clinical-trial-screening/</link>
		
		<dc:creator><![CDATA[Ophelia Keating]]></dc:creator>
		<pubDate>Sun, 13 Sep 2026 01:36:50 +0000</pubDate>
				<category><![CDATA[Mathematics]]></category>
		<category><![CDATA[AHEAD 3-45]]></category>
		<category><![CDATA[AHEAD 3-45 Alzheimer's trial]]></category>
		<category><![CDATA[Alzheimer's clinical trial participant recruitment]]></category>
		<category><![CDATA[Alzheimer's disease]]></category>
		<category><![CDATA[Alzheimer's disease blood screening]]></category>
		<category><![CDATA[amyloid beta]]></category>
		<category><![CDATA[amyloid-beta protein in Alzheimer's diagnosis]]></category>
		<category><![CDATA[blood biomarkers]]></category>
		<category><![CDATA[blood-based Alzheimer's diagnostic algorithm]]></category>
		<category><![CDATA[clinical trial recruitment]]></category>
		<category><![CDATA[cost-effective Alzheimer’s screening methods]]></category>
		<category><![CDATA[early detection]]></category>
		<category><![CDATA[early detection of Alzheimer's disease]]></category>
		<category><![CDATA[Keck School of Medicine of USC]]></category>
		<category><![CDATA[lecanemab]]></category>
		<category><![CDATA[Mixture of Experts]]></category>
		<category><![CDATA[P-tau217]]></category>
		<category><![CDATA[PET imaging]]></category>
		<category><![CDATA[PET imaging in Alzheimer's research]]></category>
		<category><![CDATA[PET scan reduction in clinical trials]]></category>
		<category><![CDATA[plasma screening]]></category>
		<category><![CDATA[potential for earlier Alzheimer's treatment]]></category>
		<category><![CDATA[USC Alzheimer's research advancements]]></category>
		<category><![CDATA[use of statistical modeling in Alzheimer's diagnostics]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=200512</guid>

					<description><![CDATA[A USC-developed blood-based screening algorithm cut the rate of unnecessary PET scans among Alzheimer's clinical trial candidates from over 70 percent to 31 percent during recruitment for the AHEAD 3-45 study.]]></description>
										<content:encoded><![CDATA[<p>A new blood-based screening algorithm developed at the University of Southern California has sharply reduced the number of costly, unnecessary PET imaging scans required to identify candidates for a major Alzheimer&#8217;s disease clinical trial, according to research from scientists at the Keck School of Medicine of USC published in Alzheimer&#8217;s &amp; Dementia: The Journal of the Alzheimer&#8217;s Association. The algorithm, deployed during recruitment for the international phase 3 AHEAD 3-45 study, cut the proportion of candidates who underwent a positron-emission tomography scan but ultimately did not qualify for the trial from more than 70 percent to 31 percent. The finding represents one of the clearest demonstrations yet that simple blood tests, combined with thoughtful statistical modeling, can transform the logistics of recruiting participants for studies of a disease that begins silently in the brain decades before memory fades.</p>
<p>The AHEAD 3-45 trial is testing a deceptively simple question: whether treating people with an approved Alzheimer&#8217;s drug earlier, before any outward sign of decline, can produce better outcomes than waiting for symptoms to appear. The drug in question, lecanemab, works by clearing sticky aggregates of amyloid-beta protein that accumulate in the brain and are closely associated with the progression of Alzheimer&#8217;s disease. In patients who already show cognitive symptoms, the drug can slow the clinical progression of the disease by roughly 30 percent. But it is currently prescribed only for people who have already begun to decline, leaving open the tantalizing possibility that intervening earlier could do far more.</p>
<p>That possibility creates an enormous recruitment challenge. The AHEAD 3-45 trial specifically targets people with early amyloid-beta buildup, a process that can begin as much as twenty or thirty years before the first noticeable problems with memory or thinking. Yet only about 30 percent of cognitively healthy adults over the age of 65 carry amyloid levels high enough to qualify for a trial like AHEAD. Finding those individuals means screening large numbers of older adults who feel perfectly well and have no obvious symptoms, then confirming with certainty that the biological hallmarks of the disease are present before enrolling them in a study that will involve repeated assessments and, ultimately, treatment with an amyloid-clearing antibody.</p>
<p>Before blood plasma screening entered the picture, that confirmation depended almost entirely on PET imaging. Positron-emission tomography is considered the gold standard for visualizing amyloid in the living brain, but it is expensive, requires access to specialized scanners and radiotracers, and adds months to the enrollment timeline. Candidates for the AHEAD 3-45 trial once faced a screening process stretching up to three months from their first visit to enrollment, and after undergoing a PET scan, more than 70 percent would learn they were ineligible. In other words, for every participant enrolled, the trial was paying for and performing multiple scans that revealed nothing more than the absence of the very biology the study was hunting for.</p>
<p>&#8220;The screening part of an Alzheimer&#8217;s trial is usually one of the costliest parts of the clinical trial for study sites,&#8221; said corresponding author Oliver Langford, MS, modeling and simulation director at the USC Epstein Family Alzheimer&#8217;s Therapeutic Research Institute at the Keck School of Medicine. &#8220;We definitely helped reduce the burden for participating clinical sites and for patients by reducing the number of individuals having to undergo PET scans.&#8221; The savings cascade beyond trial budgets. Fewer scans mean fewer clinic visits, less exposure to radiation and needles for volunteers, and a faster path from first contact to enrollment for the people the trial most needs to reach.</p>
<p>The algorithmic tool at the heart of the study incorporated two blood plasma biomarkers that have risen rapidly to prominence in Alzheimer&#8217;s research. The first is the amyloid-beta ratio, an early signal of amyloid accumulation that reflects the gradual sequestration of the protein into plaques in the brain. The second is p-tau217, a recently discovered marker of a phosphorylated form of the tau protein that more reliably reflects the actual amyloid burden carried in an individual&#8217;s brain. Blood tests measuring these proteins have shown steadily improving accuracy in recent years, and their performance has fueled a broader shift toward making Alzheimer&#8217;s detection accessible in ordinary clinical settings rather than confined to specialized imaging centers.</p>
<p>Building the model required data on a remarkable scale. The USC team developed the algorithm using information from 1,080 participants in the AHEAD study and then validated it in an independent dataset drawn from the Wisconsin Registry for Alzheimer&#8217;s Prevention, a long-running observational study of dementia risk. Beyond the two plasma biomarkers, the model incorporated age and APOE4 carrier status, two well-established factors that increase the likelihood of amyloid accumulation. Crucially, the algorithm was not frozen at birth. It was refined across three successive versions during active recruitment, which ran from 2020 to 2024, allowing the researchers to fold in lessons from real-world screening performance as the trial progressed.</p>
<p>The staged rollout demonstrated how quickly each refinement translated into efficiency. When first introduced in February 2022, the version of the algorithm built on the amyloid-beta ratio alone reduced the proportion of participants who received a PET scan but proved ineligible from 71 percent to 50 percent. In May 2023, a second version integrated p-tau217, and that figure fell further to 31 percent. The improvement enabled the trial to enroll participants with both intermediate and elevated amyloid levels far more efficiently, sparing more than half of the would-be scanned candidates an unnecessary imaging procedure. For a phase 3 study recruiting across dozens of sites internationally, the cumulative effect on cost, time and volunteer burden is substantial.</p>
<p>Central to the algorithm&#8217;s design is a statistical technique known as Mixture of Experts, or MoE, which allows the model to grapple with a biological reality that simpler tools ignore: amyloid does not accumulate uniformly across the population. Rather than forcing each person into a binary positive or negative category, the team built a model that estimates where an individual falls along a continuous spectrum of amyloid buildup. &#8220;When we look at amyloid levels at a population level, we see a peak where people are amyloid-negative and another peak for those with elevated amyloid plaque buildup,&#8221; Langford explained. &#8220;There is a region in the intermediate range that isn&#8217;t fully captured by the plasma marker alone. The Mixture of Experts approach helps model that uncertainty more effectively and is better suited to the type of data collected.&#8221; By explicitly modeling the murky middle of the distribution, the algorithm can flag borderline cases for confirmatory imaging while confidently routing clear-cut cases either into the trial or out of it.</p>
<p>The implications extend well beyond a single clinical trial. The results add to a growing body of evidence that blood tests can serve as a practical first-pass filter before more intensive and expensive diagnostics, a reordering of the diagnostic pipeline that could reshape both research and routine care. &#8220;It&#8217;s going to allow more people to access testing that can help determine whether they have Alzheimer&#8217;s disease pathology,&#8221; Langford said. &#8220;If you are able to go to your doctor and get a blood test done, you&#8217;ll be able to understand whether you have the disease earlier.&#8221; More efficient screening for trials like AHEAD 3-45 may also bring the field a step closer to its ultimate ambition: primary prevention of Alzheimer&#8217;s disease. &#8220;If we can intervene earlier,&#8221; Langford said, &#8220;we&#8217;ll have a larger effect and hopefully prevent people from having symptoms.&#8221; In addition to Langford, the study&#8217;s authors include Rema Raman, Paul Aisen, Doris Molina-Henry, Gustavo A. Jimenez-Maggiora, Robert A. Rissman and Michael C. Donohue of USC; Reisa Sperling, Keith Johnson, Colin Birkenbihl, Madison Cuppels and Rachel F. Buckley of Harvard Medical School, Brigham and Women&#8217;s Hospital and Massachusetts General Hospital; Pallavi Sachdev and David Li of Eisai Inc.; and Sterling C. Johnson of C2N Diagnostics. The AHEAD Study is conducted through the Alzheimer&#8217;s Clinical Trials Consortium, funded by the National Institute on Aging of the National Institutes of Health under award U24AG057437, and is a public-private partnership supported by the NIA, Eisai, the GHR Foundation, the Alzheimer&#8217;s Association and other philanthropic organizations.</p>
<p><strong>Subject of Research:</strong> Development and validation of a blood-based biomarker screening algorithm to improve recruitment efficiency for a preclinical Alzheimer&#x27;s disease clinical trial</p>
<p><strong>Article Title:</strong> USC develops algorithmic tool to improve screening of patients for Alzheimer&#x27;s clinical trial</p>
<p><strong>Article References:</strong> USC develops algorithmic tool to improve screening of patients for Alzheimer&#x27;s clinical trial. (n.d.). <a href="https://www.eurekalert.org/news-releases/1142829" rel="noopener noreferrer">Original publication</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> Not provided</p>
<p><strong>Keywords:</strong> Alzheimer&#x27;s disease, blood biomarkers, PET imaging, lecanemab, amyloid-beta, p-tau217, clinical trial recruitment, Mixture of Experts, AHEAD 3-45, Keck School of Medicine of USC, plasma screening, early detection</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">200512</post-id>	</item>
	</channel>
</rss>
