<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>distribution shift &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/distribution-shift/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Fri, 02 Oct 2026 01:54:29 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>distribution shift &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>When Machines Meet the Unknown: Why AI Fault Diagnosis Breaks Down and How Scientists Map the Fix</title>
		<link>https://scienmag.com/when-machines-meet-the-unknown-why-ai-fault-diagnosis-breaks-down-and-how-scientists-map-the-fix/</link>
		
		<dc:creator><![CDATA[Ophelia Keating]]></dc:creator>
		<pubDate>Fri, 02 Oct 2026 01:54:29 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[acoustic emission analysis for machinery health]]></category>
		<category><![CDATA[AI fault diagnosis]]></category>
		<category><![CDATA[AI model robustness in industrial environments]]></category>
		<category><![CDATA[challenges of deploying AI for predictive maintenance]]></category>
		<category><![CDATA[deep learning]]></category>
		<category><![CDATA[deep learning for machinery maintenance]]></category>
		<category><![CDATA[distribution shift]]></category>
		<category><![CDATA[domain adaptation]]></category>
		<category><![CDATA[domain generalization]]></category>
		<category><![CDATA[intelligent fault diagnosis]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[machine learning fragility in real-world conditions]]></category>
		<category><![CDATA[mechanical systems]]></category>
		<category><![CDATA[model calibration]]></category>
		<category><![CDATA[neural networks in industrial monitoring]]></category>
		<category><![CDATA[open-set recognition]]></category>
		<category><![CDATA[out-of-distribution data challenges in AI]]></category>
		<category><![CDATA[out-of-distribution detection]]></category>
		<category><![CDATA[predictive maintenance]]></category>
		<category><![CDATA[scientific approaches to improving AI reliability in machinery]]></category>
		<category><![CDATA[sensor data variability in fault detection]]></category>
		<category><![CDATA[survey]]></category>
		<category><![CDATA[systematic mapping of AI fault diagnosis failures]]></category>
		<category><![CDATA[vibration signal analysis in AI diagnostics]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=224998</guid>

					<description><![CDATA[A new survey of 161 studies maps how out-of-distribution data undermines AI-based mechanical fault diagnosis and argues that generalization, calibrated novelty detection, and human escalation must be engineered and evaluated as separate system components.]]></description>
										<content:encoded><![CDATA[<p>Deep learning has transformed the way engineers watch over rotating machinery. Neural networks trained on vibration signals, acoustic emissions, and current waveforms can now spot a bearing defect or an imbalance long before a human inspector would notice anything wrong. Yet a sweeping new survey published in Artificial Intelligence Review argues that the very success of these systems conceals a dangerous fragility: the moment real-world conditions drift away from the data a model saw during training, its confident diagnoses can quietly turn into confident mistakes. The study, led by Fanwei Lin and colleagues at Xi&#8217;an Jiaotong University, is one of the most systematic attempts to map exactly where intelligent fault diagnosis fails when it leaves the laboratory.</p>
<p>The core problem is known in machine learning as out-of-distribution, or OOD, data. A diagnostic network is, at heart, a pattern matcher that learns statistical regularities from a fixed training set. If a model is trained on vibration recordings collected at one motor speed, with one sensor placement, on one fleet of machines, it implicitly assumes those regularities will hold forever. In industrial reality they never do. Loads change, speeds change, sensors age and are replaced, entire machine models are swapped in, and genuinely new fault modes appear that no engineer anticipated when the labels were written. Each of these changes shifts the input distribution, and the survey&#8217;s central claim is that this shift splits into two distinct technical demands that the field has too often blurred together.</p>
<p>The first demand is generalization: the model must keep classifying known fault categories accurately even when conditions differ from training. The second demand is detection: the model must recognize when an input belongs to no known category at all and refuse to force it into one of its trained labels. These goals pull in opposite directions. A model tuned to be maximally confident on known faults will often assign absurdly high certainty to unfamiliar inputs, because softmax-based classifiers were never designed to say &#8220;I don&#8217;t know.&#8221; Conversely, a detector tuned to be suspicious of everything novel may start rejecting legitimate fault signals that merely look unusual. The survey treats this tension as the defining operational boundary of the field.</p>
<p>To build the review on solid ground, the authors ran a reproducible literature search across IEEE Xplore, the Web of Science Core Collection, and Scopus, covering publications from January 1, 2020, to July 16, 2026. The search returned 280 database records, which the team reduced to 181 after removing duplicates and then distilled to 161 primary studies specifically concerned with out-of-distribution issues in mechanical fault diagnosis. Each included study was coded according to its deployment role and the type of distribution shift it addressed. That systematic coding, the authors argue, is what allows the field&#8217;s scattered results to be compared on equal terms rather than celebrated anecdotally.</p>
<p>A key conceptual contribution of the paper is a task-relative operational definition of what &#8220;out-of-distribution&#8221; actually means in this domain. Rather than treating OOD as a vague property of data, the survey anchors it in three elements: the effective training support, meaning the data the model actually learned from; the training label space, meaning the fault categories the model was taught; and a declared deployment envelope, meaning the range of conditions the system is promised to handle in service. This framing lets the authors cleanly separate condition-related shifts, which they call C-shifts, from fault-semantic shifts, or F-shifts, and from combined C+F settings where both occur at once. A change in rotational speed is a C-shift; the sudden appearance of a compound gear-mesh defect never present in the training labels is an F-shift; a new machine running at a new speed with a new fault is the combined case.</p>
<p>This taxonomy matters because the two shift types demand different machinery. Generalization-oriented methods, such as domain adaptation and domain generalization, aim to learn representations that are invariant to condition changes, so that a fault signature extracted from a slow-running pump still matches the features the classifier learned from fast-running ones. Detection-oriented methods, drawing on open-set recognition and OOD detection research, instead focus on measuring how far an input lies from the training support, using tools like distance metrics in feature space, energy scores, density estimates, or reconstruction errors from generative models. The survey organizes representative techniques by their deployment role and underlying mechanism, comparing their assumptions, data requirements, output types, strengths, and limitations in a single framework.</p>
<p>The comparison exposes uncomfortable gaps between what papers report and what deployment requires. Many studies evaluate generalization on benchmark datasets where the shift is mild and the label space is closed, which flatters accuracy numbers while saying little about behavior when a truly unknown fault arrives. Others demonstrate detection on synthetic anomalies that barely resemble real mechanical novelty. The survey also notes that the literature frequently conflates general anomaly detection, which flags any statistical outlier, with unknown-fault rejection, which must specifically decide whether a signal belongs outside the declared fault taxonomy. A cooling-fan hum that is statistically odd but mechanically harmless should not trigger the same alarm as an incipient bearing spall, yet generic detectors cannot make that distinction without domain-aware design.</p>
<p>Perhaps the survey&#8217;s most consequential argument is architectural in the broadest sense: reliable diagnosis is not a single model property but a system property. The authors contend that any trustworthy deployment must evaluate three components separately and then integrate them coherently. First, known-class generalization must be measured under realistic condition shifts within the declared envelope. Second, OOD detection must be calibrated, meaning that when the system reports uncertainty or novelty, those reports should correspond to actual error rates rather than arbitrary scores. Third, there must be an explicit escalation or human-referral pathway, so that detected unknowns are routed to engineers rather than silently absorbed into the nearest familiar class. A pipeline that excels at the first two but lacks the third still fails the operator, because a flagged anomaly with no defined next step is functionally indistinguishable from an unflagged one.</p>
<p>The scope of the review is deliberately bounded, and the authors are transparent about the exclusions. They do not attempt an exhaustive treatment of closed-set domain adaptation, of domain generalization studies lacking unseen mechanical-domain evaluation, of general anomaly detection unrelated to unknown-fault rejection, of non-mechanical industrial process monitoring, or of general-purpose computer-vision OOD methods, except where such work supplies necessary conceptual foundations. This restraint keeps the 161 included studies focused on the mechanical diagnosis setting, where signals are physical, faults evolve over time, and the cost of a missed or fabricated diagnosis is measured in downtime, damage, and sometimes safety.</p>
<p>For an industry racing to embed AI in turbines, gearboxes, and production lines, the message is sobering but constructive. The survey does not claim that deep-learning fault diagnosis is unreliable; it claims that reliability has been measured with the wrong yardstick. Accuracy on a held-out test set drawn from the same distribution says nothing about the day a new sensor type, a new operating regime, or a never-before-seen failure mode arrives. By giving the field a shared vocabulary of C-shifts, F-shifts, training support, and deployment envelopes, and by insisting that generalization, calibrated detection, and human escalation be judged as separate, testable commitments, the Xi&#8217;an Jiaotong team has effectively written the acceptance criteria that future diagnostic systems will have to meet before they can honestly be called intelligent. The machines, it turns out, are only as smart as their honesty about what they do not know.</p>
<p><strong>Subject of Research:</strong> Out-of-distribution generalization and unknown-fault detection in deep-learning-based intelligent fault diagnosis of mechanical systems</p>
<p><strong>Article Title:</strong> Out-of-distribution challenges in intelligent fault diagnosis: a survey of generalization and detection</p>
<p><strong>Article References:</strong> Lin, F., Zhang, L., Guo, C., Zhao, Z., Zhang, X., &amp; Chen, X. (2026). Out-of-distribution challenges in intelligent fault diagnosis: a survey of generalization and detection. <em>Artificial Intelligence Review</em>. <a href="https://doi.org/10.1007/s10462-026-11722-3" rel="noopener noreferrer">https://doi.org/10.1007/s10462-026-11722-3</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s10462-026-11722-3" rel="noopener noreferrer">10.1007/s10462-026-11722-3</a></p>
<p><strong>Keywords:</strong> out-of-distribution detection, intelligent fault diagnosis, deep learning, distribution shift, domain generalization, open-set recognition, mechanical systems, predictive maintenance, machine learning, model calibration, domain adaptation, survey</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">224998</post-id>	</item>
		<item>
		<title>When AI Cannot See the Answer: The Mathematical Limit That Optimisation Cannot Cross</title>
		<link>https://scienmag.com/when-ai-cannot-see-the-answer-the-mathematical-limit-that-optimisation-cannot-cross/</link>
		
		<dc:creator><![CDATA[Reid Dalton]]></dc:creator>
		<pubDate>Sat, 26 Sep 2026 22:25:20 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[admissibility]]></category>
		<category><![CDATA[AI limitations]]></category>
		<category><![CDATA[AI robustness and data coverage]]></category>
		<category><![CDATA[AI safety]]></category>
		<category><![CDATA[AI scaling challenges]]></category>
		<category><![CDATA[Artificial Intelligence]]></category>
		<category><![CDATA[benchmark failure]]></category>
		<category><![CDATA[closure architecture]]></category>
		<category><![CDATA[distribution shift]]></category>
		<category><![CDATA[formal verification]]></category>
		<category><![CDATA[future directions in AI reliability]]></category>
		<category><![CDATA[Goodhart's law]]></category>
		<category><![CDATA[limits of AI training and generalization]]></category>
		<category><![CDATA[neural network misclassification]]></category>
		<category><![CDATA[observation maps]]></category>
		<category><![CDATA[observational regimes in AI]]></category>
		<category><![CDATA[optimization failure vs admissibility failure]]></category>
		<category><![CDATA[quotient factorisation]]></category>
		<category><![CDATA[reinforcement learning loopholes]]></category>
		<category><![CDATA[reward hacking]]></category>
		<category><![CDATA[structural conditions in AI systems]]></category>
		<category><![CDATA[theoretical foundations of AI failures]]></category>
		<category><![CDATA[transfer learning and distribution shift]]></category>
		<category><![CDATA[warrant debt]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=216709</guid>

					<description><![CDATA[A new theoretical review argues that reward hacking, benchmark failure, and distribution shift share a single structural root: observation regimes that erase the very distinctions AI systems are asked to judge.]]></description>
										<content:encoded><![CDATA[<p>A neural network labels a panda as a gibbon after a nearly invisible perturbation, a language model grows worse on a narrow family of tasks even as it scales, a reinforcement learning agent racks up a perfect score by exploiting a loophole its designer never imagined, and a model trained on one population collapses when deployed on another that looks, by every measurable engineering criterion, identical. These are among the most-discussed failure modes in modern artificial intelligence, and each has spawned its own technical literature with its own remedies. A new theoretical review published in Discover Artificial Intelligence argues that beneath this apparent diversity lies a single structural condition, one that sits logically upstream of training, data coverage, and statistical robustness alike. The paper, authored by Duston Moore and published open access in September 2026, contends that many celebrated AI failures are not optimisation failures at all. They are failures of admissibility, and no amount of further optimisation can repair them.</p>
<p>The central idea can be stated without symbols. Any AI system acts through an observational regime: a channel that maps the world&#8217;s possible states onto what the system can actually perceive. A thermometer, for example, maps bodily states onto a single temperature reading. It can settle whether a patient has a fever, but it can never determine which infection caused that fever, because two patients with entirely different infections can register the same temperature. No physician, however brilliant, can extract that distinction from the instrument, because the instrument never preserved it. Moore formalises this intuition with an elementary condition drawn from the mathematics of quotient sets: a yes-or-no question about the world, cast as a predicate on states, is decidable by a system only when the question gives the same answer for any two states the observation map cannot distinguish. When this factorisation condition fails, the question is well-posed in the world but malformed inside the machine.</p>
<p>From this definition follows a proposition that is mathematically trivial yet strategically devastating. If a predicate is not admissible with respect to the observation map, then no decision rule operating on observations alone can agree with it, and no optimisation criterion defined over such rules can recover it. A second proposition extends the verdict to post-processing: calibration, reranking, reward shaping, decoding, and benchmark aggregation may improve performance when the required distinction survives in the representation, but none can restore a distinction the observation map has already erased. In the idiom of Edsger Dijkstra, admissibility functions as the weakest precondition for asking a question through an observation. If the question is not invariant under the identifications the observation performs, no later command operating on the same observations can make it so. The failure lies upstream of inference, and nothing downstream of the observation channel can repair it.</p>
<p>To show that this is more than an abstract truism, the paper constructs a finite example that can be checked by hand. Four regions are arranged in a ring, each overlapping only its two neighbours, with local assignments on each region and recorded discrepancies on each overlap. The question is whether the local assignments assemble into a single globally consistent object. In the worked instance, every overlap reports a mismatch of exactly one unit. The profile passes every local coherence check available to the structure, because with no triple overlaps there is no way for the residue to contradict itself. Yet summing the four mismatch equations yields zero on one side and four on the other, an outright contradiction proving that no local correction can cancel the discrepancy. The gap between local coherence and global coherence is captured by a computable residue, which Moore names warrant debt: not a measure of how much more search is needed, but a certificate that no correct global assembly exists under the current data.</p>
<p>The same triple structure, a state space, an observation map, and a predicate that fails to descend, is then written out explicitly for two of the most consequential failure modes in contemporary AI. In reward hacking, the state space consists of an agent&#8217;s environment trajectories, the observation map is the scalar reward trace, and the intended predicate is genuine task success without deception. When two trajectories carry identical reward profiles while one accomplishes the intended task and the other violates it, the reward regime is inadmissible, and the paper&#8217;s fixed-regime limit applies: no optimiser whose objective is a function of the reward alone can recover the designer&#8217;s intent. The ordinary description says the agent gamed the reward. The admissibility account gives a structurally prior diagnosis: optimisation exposed a collapse that was already present, rather than causing it. The diagnosis holds whether the reward function is hand-designed or learned from human preference comparisons.</p>
<p>Benchmark and proxy failure receive the same treatment. Here the observation map records benchmark scores and evaluation transcripts, and the intended predicate is deployment adequacy: competence, safety, and reliability on the situations a system will actually face. When two interactions look equivalent under the benchmark while one is deployment-adequate and the other is not, the benchmark induces an observational quotient that deletes exactly the distinctions that matter. Brittleness under distribution shift, on this reading, is the expected consequence of training against a regime that was never shown to be admissible for the predicate it was meant to track. The framework also yields a sharpened reading of Goodhart&#8217;s law. When a measure becomes a target, optimisation drives the system toward the extreme regions of the observation space, where the equivalence classes are most likely to contain states that differ sharply in the intended property. Goodhart&#8217;s law, in this account, is simply what admissibility failure looks like once a proxy is made operational and optimisation is allowed to run.</p>
<p>What distinguishes the paper from pure critique is its architectural proposal. Moore separates productive systems, whose repertoire is search, prediction, generation, and optimisation within a fixed regime, from closure authority, which checks whether the regime itself preserves the distinctions a judgement requires. The analogy is the proof assistant: tactics may propose derivations heuristically, but a small kernel alone admits them, and the kernel does not search, it certifies. The paper demonstrates the discipline with an executable prototype for a cyclic diagnostic system, in which a productive layer proposes a fault attribution and a closure layer constructs a residue, computes its harmonic period, and emits one of three accountable verdicts: coherence failure, global admissibility, or warrant debt. In the worked instance the warrant-debt magnitude is exactly 25/4, a finite certificate that the local diagnostic residues, though mutually coherent, do not warrant the global claim. A refinement case shows the regime itself being revised, with an added sensor distinction driving the obstruction to zero and flipping the verdict to globally admissible. The certificates are checked by an independent verifier implemented separately from the proposing process, in exact rational arithmetic, with parts of the verdict logic mechanised in the Rocq proof assistant.</p>
<p>Moore is explicit about what the proposal does not solve. The complete detection algorithm applies only to finite, declared regimes, where the observation map and its operators are available as checkable objects. In modern neural networks the effective observation map is implicit, distributed across embeddings, hidden activations, attention patterns, and reward channels, and no general method exists for extracting an auditable surrogate from those weights. The paper specifies what a successful implicit-map algorithm would need: regime exposure, predicate specification from outside the productive component, probing of observational equivalence classes, contrastive witness search, positive certification rather than mere absence of counterexamples, independent verification, and an escalation discipline that never converts a missing certificate into silent permission to proceed. It also sketches how the deterministic theory extends to stochastic observation, where admissibility becomes a sufficiency condition connecting the framework to the classical theory of statistical experiments.</p>
<p>The philosophical stakes are considerable. The paper suggests, cautiously and as motivation rather than established fact, that behaviour appearing random under one observational regime may reflect competence under distinctions that regime fails to preserve, and it supplies a falsifiability criterion: a genuine boundary-revision claim must exhibit a refined observation map and behaviour that becomes stable and repeatable under the new description, while noise remains unstable under refinement. The deeper contrast is architectural rather than carbon-versus-silicon. Contemporary systems can change their representations through training, but representational change produced by optimisation within a fixed objective is not certified repair of inadmissibility; updating within a regime and revising the regime are different obligations. The conclusion lands as a challenge to the field&#8217;s dominant metric culture. A confidence score, however carefully calibrated, is still a function of the regime already in force and cannot certify that regime&#8217;s adequacy. What consequential AI systems owe us instead, the paper argues, is an admissibility witness: a checkable indication that the regime under which an answer is offered preserves the distinctions on which the answer depends, or an honest signal that it does not.</p>
<p><strong>Subject of Research:</strong> Formal admissibility conditions defining structural limits on optimisation in artificial intelligence systems</p>
<p><strong>Article Title:</strong> Admissibility defines structural limits on optimisation in artificial intelligence</p>
<p><strong>Article References:</strong> Moore, D. (2026). Admissibility defines structural limits on optimisation in artificial intelligence. <em>Discover Artificial Intelligence, 6</em>(1), Article 1267. <a href="https://doi.org/10.1007/s44163-026-02252-6" rel="noopener noreferrer">https://doi.org/10.1007/s44163-026-02252-6</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s44163-026-02252-6" rel="noopener noreferrer">10.1007/s44163-026-02252-6</a></p>
<p><strong>Keywords:</strong> admissibility, artificial intelligence, reward hacking, Goodhart&#x27;s law, benchmark failure, warrant debt, formal verification, observation maps, AI safety, closure architecture, quotient factorisation, distribution shift</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">216709</post-id>	</item>
		<item>
		<title>AI Retrosynthesis Model RetroChimera Learns Chemistry the Way Chemists Do</title>
		<link>https://scienmag.com/ai-retrosynthesis-model-retrochimera-learns-chemistry-the-way-chemists-do/</link>
		
		<dc:creator><![CDATA[Bethany Barker]]></dc:creator>
		<pubDate>Tue, 22 Sep 2026 15:12:18 +0000</pubDate>
				<category><![CDATA[Medicine]]></category>
		<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[advanced retrosynthesis algorithms]]></category>
		<category><![CDATA[AI failures in chemical synthesis]]></category>
		<category><![CDATA[AI-assisted drug discovery]]></category>
		<category><![CDATA[Artificial Intelligence]]></category>
		<category><![CDATA[artificial intelligence in organic chemistry]]></category>
		<category><![CDATA[chemical synthesis]]></category>
		<category><![CDATA[chemical synthesis planning]]></category>
		<category><![CDATA[cheminformatics]]></category>
		<category><![CDATA[chemist-AI collaboration]]></category>
		<category><![CDATA[chemistry reaction prediction]]></category>
		<category><![CDATA[deep learning]]></category>
		<category><![CDATA[distribution shift]]></category>
		<category><![CDATA[ensembles]]></category>
		<category><![CDATA[inductive bias]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[machine learning in chemistry]]></category>
		<category><![CDATA[pharmaceutical discovery]]></category>
		<category><![CDATA[pharmaceutical industry AI tools]]></category>
		<category><![CDATA[RetroChimera]]></category>
		<category><![CDATA[RetroChimera AI system]]></category>
		<category><![CDATA[retrosynthesis]]></category>
		<category><![CDATA[Retrosynthesis AI models]]></category>
		<category><![CDATA[synthesis planning]]></category>
		<category><![CDATA[synthesis prediction accuracy]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=206211</guid>

					<description><![CDATA[A new AI retrosynthesis system called RetroChimera combines models with complementary inductive biases to outperform leading baselines and earn the preference of expert organic chemists.]]></description>
										<content:encoded><![CDATA[<p>For more than half a century, one of the most demanding intellectual exercises in chemistry has remained largely unchanged: retrosynthesis, the art of working backward from a desired molecule to the simpler building blocks and reactions that could realistically produce it. Now, a team of researchers from Microsoft Research AI for Science, Novartis Biomedical Research, the University of Cambridge, Jagiellonian University, GSK, and Bergische Universität Wuppertal reports a new artificial intelligence system, called RetroChimera, that appears to close a stubborn gap between what machine learning models can predict and what experienced chemists actually consider plausible. The study, published in Nature, describes both a detailed diagnosis of why existing AI synthesis planners fail and a novel architecture that addresses those failures head-on.</p>
<p>The motivation for the work stems from a simple but uncomfortable observation. Although AI-assisted synthesis planning has proliferated across the chemical and pharmaceutical industries in recent years, a detailed understanding of how and why these systems fail has never been achieved. According to the authors, current models still struggle with predicting less frequent yet strategically critical reactions, and they continue to produce hallucinated, incorrect predictions that are fundamentally misaligned with chemists&#8217; expectations. For a discovery chemist weighing whether to trust a model&#8217;s suggestion, such failures are not merely inconvenient; they can derail weeks of planning and undermine confidence in computational tools altogether.</p>
<p>The core of the new system rests on two newly developed model components with what the researchers describe as complementary inductive biases. Inductive biases are the built-in assumptions a model makes about the structure of the problem it is learning, and in chemistry these assumptions matter enormously. One component may excel at capturing the local, rule-like transformations that dominate reaction databases, while another is better positioned to generalize to sparse, unusual reaction classes where such rules break down. By pairing architectures whose strengths and blind spots differ in a structured way, the team ensured that the ensemble&#8217;s errors would be less correlated than those of any single model.</p>
<p>Crucially, however, simply averaging predictions from different architectures is not enough, because the outputs of retrosynthesis models are discrete chemical transformations rather than continuous scores. The researchers therefore introduced a novel, learning-based ensembling strategy that integrates the diverse components. Rather than treating the ensemble as a post-hoc投票 mechanism, the approach learns how to weigh and combine candidate predictions in a way that reflects chemical validity and likelihood. This design allows RetroChimera to draw on the complementary knowledge embedded in each constituent model without letting the weaknesses of one corrupt the strengths of the other.</p>
<p>The experimental evaluation was deliberately broad, spanning several orders of magnitude in data scale, and the results point to robustness that previous systems lacked. RetroChimera outperformed leading baselines across the benchmark settings tested, and it demonstrated the ability to learn from very small numbers of examples per reaction class. This latter capability addresses one of the most persistent pain points in machine learning for chemistry: the long tail of the reaction literature. Highly unusual transformations, such as those used to forge complex molecular frameworks in natural product synthesis or to install challenging stereochemical motifs, are chronically underrepresented in public databases, yet they are often the exact steps that determine whether an ambitious synthesis route is feasible at all.</p>
<p>The team also subjected the model to tests of generalization outside its training data, a regime in which many deep learning systems collapse into confident nonsense. RetroChimera&#8217;s predictions remained reliable under distribution shift, suggesting that the combination of diverse inductive biases and learned ensembling confers a form of robustness that single-architecture models have not matched. The researchers frame this as a demonstration of the viability of deep learning for accurate synthesis prediction in increasingly challenging regimes, where the easy, well-documented reactions that dominate benchmarks are no longer representative of the problems chemists face.</p>
<p>Perhaps the most consequential validation, however, came not from automated benchmarks but from human experts. Using both pairwise and pointwise evaluation setups, the researchers asked organic chemists to compare RetroChimera&#8217;s predictions against published reference reactions and against the outputs of other AI models. In these blind assessments, the chemists systematically preferred the suggestions generated by RetroChimera. That preference for model output over the actual published record of real reactions is a striking result, because it indicates the system can propose transformations that experts judge more sensible or more practical than the routes that chemists historically chose and journals published.</p>
<p>To test whether these gains survive contact with industrial reality, the team performed zero-shot transfer and fine-tuning experiments on internal datasets from two major pharmaceutical companies. These proprietary datasets differ from public reaction corpora in both scale and character, reflecting the molecule types, reaction conditions, and strategic priorities of real drug discovery programs. RetroChimera showed robust generalization under this distribution shift, performing well without any additional training in the zero-shot setting and improving further when fine-tuned on internal data. For pharmaceutical companies that have long been skeptical of benchmarks built on public data alone, this industrial validation may prove as important as any academic result.</p>
<p>The collaboration itself reflects the interdisciplinary demands of the problem. The author list spans machine learning researchers specializing in molecular AI, computational chemists embedded in industrial research organizations, and practicing medicinal and synthetic chemists at GSK, Novartis, and other partner organizations. Equal contributions are credited to Krzysztof Maziarz, Guoqing Liu, and Felix Pultar of Microsoft Research AI for Science, with corresponding authorship shared with Marwin H. S. Segler, who has worked on AI-driven synthesis planning for nearly a decade. Representatives from GSK&#8217;s discovery sciences group in Stevenage and Wuppertal-based chemist Mario P. Wiesenfeldt round out a team designed to ensure that evaluation criteria mirrored the judgments of working chemists rather than proxy metrics alone.</p>
<p>Looking ahead, the work suggests a broader lesson for applying AI in the natural sciences: progress may depend less on scaling a single architecture than on deliberately designing and combining models with different, complementary biases. As chemical synthesis remains a critical bottleneck in the discovery and manufacture of functional small molecules, tools that align with expert intuition while generalizing beyond the training literature could reshape how new medicines, materials, and agrochemicals are planned. If systems like RetroChimera continue to earn the trust of the chemists who use them, the long-promised collaboration between artificial intelligence and synthetic chemistry may finally be moving from demonstration to daily practice.</p>
<p><strong>Subject of Research:</strong> Chemist-aligned AI retrosynthesis planning using ensembles of diverse inductive bias models</p>
<p><strong>Article Title:</strong> Chemist-aligned retrosynthesis by ensembling diverse inductive bias models</p>
<p><strong>Article References:</strong> Maziarz, K., Liu, G., Pultar, F., Gardner, J., Gensch, T., Helie, J., Misztela, H., Tripp, A., Li, J., Kornev, A., Gaiński, P., Hoefling, H., Fortunato, M., Gupta, R., Baxter, A., Poole, D. L., Elward, J. M., Krzyzanowski, A., Pogány, P., &#8230; Segler, M. H. S. (2026). Chemist-aligned retrosynthesis by ensembling diverse inductive bias models. <em>Nature</em>. <a href="https://doi.org/10.1038/s41586-026-11160-9" rel="noopener noreferrer">https://doi.org/10.1038/s41586-026-11160-9</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1038/s41586-026-11160-9" rel="noopener noreferrer">10.1038/s41586-026-11160-9</a></p>
<p><strong>Keywords:</strong> retrosynthesis, artificial intelligence, synthesis planning, machine learning, ensembles, inductive bias, chemical synthesis, pharmaceutical discovery, cheminformatics, deep learning, distribution shift, RetroChimera</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">206211</post-id>	</item>
		<item>
		<title>AI Weather Forecasting Gets a Ruthless Audit: New Survey Exposes Hidden Flaws in How We Judge Machine-Learning Models</title>
		<link>https://scienmag.com/ai-weather-forecasting-gets-a-ruthless-audit-new-survey-exposes-hidden-flaws-in-how-we-judge-machine-learning-models/</link>
		
		<dc:creator><![CDATA[Rachel Howard]]></dc:creator>
		<pubDate>Sat, 12 Sep 2026 12:57:58 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[AI weather forecasting accuracy]]></category>
		<category><![CDATA[Artificial Intelligence]]></category>
		<category><![CDATA[benchmark design]]></category>
		<category><![CDATA[challenges of training AI models on large-scale spatiotemporal weather data]]></category>
		<category><![CDATA[climate modeling]]></category>
		<category><![CDATA[comparative analysis of AI and traditional numerical weather models]]></category>
		<category><![CDATA[critique of current AI weather forecasting evaluation methods]]></category>
		<category><![CDATA[distribution shift]]></category>
		<category><![CDATA[evaluation of machine learning models in climate science]]></category>
		<category><![CDATA[foundation models]]></category>
		<category><![CDATA[geometric deep learning]]></category>
		<category><![CDATA[impact of data complexity on AI weather forecasts]]></category>
		<category><![CDATA[importance of robust metrics for climate and weather prediction]]></category>
		<category><![CDATA[limitations of machine learning in atmospheric modeling]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[measurement challenges in AI-driven weather prediction]]></category>
		<category><![CDATA[neural operators]]></category>
		<category><![CDATA[reliability issues]]></category>
		<category><![CDATA[scientific machine learning]]></category>
		<category><![CDATA[spatiotemporal learning]]></category>
		<category><![CDATA[survey of AI applications in climate science]]></category>
		<category><![CDATA[systematic biases in weather model assessment]]></category>
		<category><![CDATA[uncertainty quantification]]></category>
		<category><![CDATA[weather forecasting]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=194571</guid>

					<description><![CDATA[A sweeping new survey in Artificial Intelligence Review maps the AI architectures transforming weather and climate prediction while exposing evaluation flaws — including a mean-squared-error bias toward blurry forecasts and ERA5 circularity — that may be inflating the apparent skill of machine-learning forecasting models.]]></description>
										<content:encoded><![CDATA[<p>Weather forecasting has quietly become one of the most visible success stories of modern artificial intelligence. In just a few years, machine-learning models have gone from experimental curiosities to systems that can rival, and in some metrics outperform, the world&#8217;s best numerical weather prediction models run on supercomputers. But according to a comprehensive new survey published in the journal Artificial Intelligence Review, the fast-moving field of AI-driven weather and climate science has a measurement problem — and the way the community currently trains and evaluates its models may be systematically rewarding the wrong kind of forecasts. The paper, authored by Andreas Holzinger of BOKU University and Graz University of Technology, together with Sandro Fiore of the University of Trento, Tullio Degiacomi of Hypermeteo, Fabrizio Antonio of the CMCC Foundation and Heimo Müller of Medical University Graz, offers both a panoramic map of the field and a pointed critique of its evaluation culture.</p>
<p>The survey&#8217;s central argument is that weather and climate represent an unusually demanding, and unusually revealing, testbed for artificial intelligence. Unlike image recognition or language modelling, atmospheric science confronts machine-learning systems with petabyte-scale spatiotemporal data spread across a rotating sphere, governed by partial differential equations and constrained by global observational and reanalysis archives. On top of that sits a challenge that most mainstream AI benchmarks simply do not have: climate non-stationarity. The statistical properties of the atmosphere are themselves shifting as the planet warms, which means models trained on the past may be silently invalidated by the very future they are asked to predict. The authors argue that this combination of scale, physics and drift makes meteorology a proving ground whose lessons generalize far beyond forecasting.</p>
<p>To organize an enormous and sometimes chaotic literature, the survey classifies the major AI architectures by their physical inductive biases — the built-in assumptions each design makes about the structure of the world. Convolutional neural networks, the workhorses of early deep-learning weather prediction, assume local spatial structure and translation invariance. Graph neural networks treat the atmosphere as an irregular mesh, naturally handling the geometry of the sphere and unstructured computational grids. Transformers bring global attention mechanisms that can capture long-range teleconnections such as the El Niño–Southern Oscillation and the Madden–Julian oscillation, which link weather patterns across entire hemispheres. Generative models, including generative adversarial networks and diffusion-based approaches, address a different problem entirely: producing realistic ensembles and downscaling coarse global fields to fine local detail.</p>
<p>Two unifying mathematical themes run through this taxonomy. The first is geometric deep learning, the program of designing networks whose internal operations respect the symmetries of the underlying space — in this case, rotation and translation on a sphere rather than on a flat plane. The second is operator learning, exemplified by neural operators such as the Fourier neural operator and its spherical variant, which learn mappings between entire function spaces rather than between individual data points. This distinction matters because weather and climate models are fundamentally functions of functions: they map one continuous field of temperature, pressure and wind onto another. Architectures that respect the spherical geometry of the planet and learn operators rather than fixed-resolution maps are, the authors argue, better positioned to generalize across resolutions and physical regimes.</p>
<p>Beyond architecture, the survey gives systematic treatment to three areas it considers underappreciated. Representation learning and foundation models — large networks pre-trained on vast atmospheric archives and then adapted to many downstream tasks — are examined as the emerging backbone of the field. Uncertainty quantification receives extensive attention, covering techniques from Bayesian approaches to ensemble generation, all aimed at answering the question operational forecasters care most about: not just what will happen, but how confident we should be. And causal discovery is framed as a complement to pure prediction, a way of using machine learning not merely to reproduce correlations in reanalysis data but to probe the physical mechanisms connecting them — a distinction that becomes critical when the climate itself is changing.</p>
<p>The paper&#8217;s most provocative contribution, however, is its naming and dissection of what the authors call evaluation pathologies in scientific machine learning. The first is the RMSE smoothness bias. Because root mean square error and mean-squared-error training objectives penalize sharp, spatially displaced features more harshly than blurry, averaged ones, models optimized on these metrics are systematically pushed toward smooth, blurred forecasts. A prediction that gets the shape of a storm exactly right but places it a few dozen kilometers off can score worse than a smeared, featureless field that is wrong everywhere but mildly. The practical consequence is that the metrics used to declare AI models superior to numerical weather prediction may be quietly selecting for aesthetically smooth mediocrity while penalizing the crisp, high-impact detail that matters most to forecasters and the public.</p>
<p>The second pathology the authors identify is ERA5 training–evaluation circularity. ERA5, the European Centre for Medium-Range Weather Forecasts&#8217; flagship reanalysis, is the de facto training ground for most AI weather models — but it is also the reference against which those models are scored. A model trained to reproduce ERA5 and then evaluated against ERA5 is, in a meaningful sense, being graded on its own homework. The circularity inflates apparent skill, obscures the reanalysis&#8217;s own biases, and makes it difficult to know how models would perform against genuinely independent observations. The third pathology, benchmark overfitting, compounds the problem: as the community iterates on a small set of standard test cases, models increasingly specialize to those cases, and leaderboard gains stop translating into real-world forecasting skill.</p>
<p>The survey does not stop at diagnosis. It frames the field&#8217;s open problems as scientific machine learning challenges that extend well beyond meteorology. Distribution shift under a non-stationary climate is the paradigm case: any AI system deployed over years must cope with input statistics that drift, potentially violating the stationarity assumptions baked into training. Physical consistency of learned operators — whether a neural network&#8217;s predictions obey conservation laws and dynamical constraints even far from its training distribution — remains unsolved. Sample efficiency in data-sparse regimes, such as the ocean interior, polar regions and the developing world&#8217;s observation networks, tests whether foundation-model approaches can transfer knowledge to places with few measurements. And the authors argue for intrinsic interpretability: not post-hoc explanations bolted onto a black box, but models whose internal reasoning is transparent enough for scientists to trust and interrogate, a theme connected to the explainable AI research program the work was partly funded to advance.</p>
<p>Why does this matter now? Because the operational stakes are rising fast. Deep-learning weather prediction systems are already being trialed by major forecasting centers, and skill on benchmarks is being cited as evidence they can replace or supplement physics-based simulation. If the benchmarks reward blur and circularity, the field risks institutionalizing models that look excellent on paper while underperforming on the rare, extreme events — hurricanes, heat waves, flash floods — where forecasts save lives. The survey&#8217;s argument is that verification against rare extremes, using metrics such as the fractions skill score and the continuous ranked probability score alongside traditional correlation measures, must become central rather than peripheral to how AI forecasters are judged.</p>
<p>The broader lesson, the authors contend, is that weather and climate offer scientific machine learning a uniquely honest mirror. The domain combines massive data, hard physics, distribution drift and unforgiving operational verification — a combination that strips away the comfortable assumptions of mainstream AI benchmarking. The survey, published open access with funding support from the Austrian Science Fund and the European Union&#8217;s Horizon Europe RI-SCALE project, is intended as both a map and a challenge: a structured account of where AI methods for the atmosphere stand today, and a warning that the path forward runs through better evaluation, not just bigger models. If the field heeds it, the same rigor that makes forecasting trustworthy in a changing climate could reshape how machine learning is validated across the sciences.</p>
<p><strong>Subject of Research:</strong> Artificial intelligence methods, benchmarking, and scientific machine learning challenges for weather and climate prediction</p>
<p><strong>Article Title:</strong> Artificial intelligence for weather and climate: a survey of methods, benchmarks, and scientific machine learning challenges</p>
<p><strong>Article References:</strong> Holzinger, A., Fiore, S., Degiacomi, T., Antonio, F., &amp; Müller, H. (2026). Artificial intelligence for weather and climate: a survey of methods, benchmarks, and scientific machine learning challenges. <em>Artificial Intelligence Review</em>. <a href="https://doi.org/10.1007/s10462-026-11690-8" rel="noopener noreferrer">https://doi.org/10.1007/s10462-026-11690-8</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s10462-026-11690-8" rel="noopener noreferrer">10.1007/s10462-026-11690-8</a></p>
<p><strong>Keywords:</strong> artificial intelligence, weather forecasting, climate modeling, machine learning, neural operators, geometric deep learning, uncertainty quantification, benchmark design, distribution shift, spatiotemporal learning, foundation models, scientific machine learning</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">194571</post-id>	</item>
	</channel>
</rss>
