<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>surrogate models &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/surrogate-models/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Thu, 08 Oct 2026 14:25:02 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.3</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>surrogate models &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>Machine Learning Takes Flight: Designing Drone Wings for Mars&#8217;s Thin Air</title>
		<link>https://scienmag.com/machine-learning-takes-flight-designing-drone-wings-for-marss-thin-air/</link>
		
		<dc:creator><![CDATA[Blake Davidson]]></dc:creator>
		<pubDate>Thu, 08 Oct 2026 14:25:02 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[aerodynamics]]></category>
		<category><![CDATA[aerodynamics of Martian drones]]></category>
		<category><![CDATA[airfoil design]]></category>
		<category><![CDATA[computational fluid dynamics]]></category>
		<category><![CDATA[computational fluid dynamics for Mars]]></category>
		<category><![CDATA[drone wing shape optimization for Mars]]></category>
		<category><![CDATA[helicopter flight on Mars]]></category>
		<category><![CDATA[hybrid AI models for aerospace]]></category>
		<category><![CDATA[low Reynolds number]]></category>
		<category><![CDATA[low Reynolds number aerodynamics]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[machine learning in aerospace engineering]]></category>
		<category><![CDATA[Mars]]></category>
		<category><![CDATA[Mars atmospheric conditions and drone technology]]></category>
		<category><![CDATA[Mars drone wing design]]></category>
		<category><![CDATA[NACA airfoils]]></category>
		<category><![CDATA[physics-guided neural network]]></category>
		<category><![CDATA[physics-guided neural networks]]></category>
		<category><![CDATA[rotorcraft]]></category>
		<category><![CDATA[space exploration]]></category>
		<category><![CDATA[surrogate models]]></category>
		<category><![CDATA[thin atmosphere drone flight]]></category>
		<category><![CDATA[UAV]]></category>
		<category><![CDATA[UAV design for extraterrestrial environments]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=248122</guid>

					<description><![CDATA[Researchers at UPES Dehradun have combined CFD simulations with machine learning, including a physics-guided neural network, to rapidly optimize drone wing airfoils for flight in Mars's extremely thin atmosphere.]]></description>
										<content:encoded><![CDATA[<p>Flying on Mars is one of the hardest problems in aerospace engineering, and a new study from researchers at UPES Dehradun, published in Aerospace Systems, shows how machine learning could make it dramatically easier. The Martian atmosphere is so thin that the air density at the surface is roughly one percent of what it is on Earth, which means any rotorcraft or fixed-wing drone sent there must generate lift from an almost empty medium. The success of NASA&#8217;s Ingenuity helicopter proved that powered flight on the Red Planet is possible, but it also revealed how fragile the aerodynamic margins are. Now, a team led by Saumya Mathur, Suruchi Gupta, Devdeep Singh and Harshit Shukla has built a hybrid framework that combines computational fluid dynamics with several families of machine learning models, including a physics-guided neural network, to rapidly identify which wing and rotor airfoil shapes perform best under Martian conditions.</p>
<p>The core challenge the researchers confronted is a numerical one. Because the Martian atmosphere is so tenuous, the chord-based Reynolds numbers experienced by small drone wings fall between roughly one thousand and one hundred thousand, far below the millions typical of terrestrial aircraft. At such low Reynolds numbers, air does not flow smoothly over a wing the way it does at full scale. Instead, the thin boundary layer of air hugging the surface tends to separate from the wing prematurely, producing laminar separation bubbles, reduced lift coefficients and stall behavior that is stubbornly nonlinear. On top of that, the speed of sound on Mars is lower than on Earth, so even a drone flying at a moderate speed can approach Mach numbers where compressibility effects distort the pressure field around the blade. Designers therefore face a double penalty: too little air to push against, and air that behaves in unfamiliar, compressible ways.</p>
<p>To map this treacherous aerodynamic landscape, the team generated a validated database of aerodynamic performance using structured-mesh computational fluid dynamics simulations. They simulated multiple NACA airfoil profiles, the classic family of wing cross-sections that has anchored a century of aeronautical design, across representative combinations of Reynolds number, Mach number and angle of attack. From each simulation they extracted the performance metrics that matter most to rotor designers: the lift-to-drag ratio, written as CL/CD, and the endurance parameter CL raised to the power of three halves divided by CD, which rewards configurations that maximize lift while minimizing drag. These two figures of essence essentially tell an engineer which airfoil will keep a drone aloft longest and most efficiently for a given power budget, which on Mars translates directly into how much science a mission can accomplish before its battery dies.</p>
<p>Running high-fidelity CFD simulations is expensive. Each case requires a carefully constructed mesh, convergence checks and hours of computation, and a full design exploration across airfoil shapes and flight conditions can consume enormous supercomputing resources. This is where the machine learning layer of the framework earns its keep. The researchers trained supervised regression models, including Random Forests, Support Vector Regression, Gradient Boosting and artificial neural networks, on the CFD-generated dataset. Once trained, these surrogate models can predict lift and drag characteristics for new combinations of airfoil geometry and flight conditions in a fraction of a second, reproducing the trends of the full physics simulations while slashing the computational cost. In effect, the neural networks and tree-based models learned the physics of low-Reynolds-number Martian aerodynamics from examples, creating a fast approximate map that designers can search far more aggressively than the original simulation grid.</p>
<p>Pure data-driven models, however, have a well-known weakness: they can interpolate beautifully within their training data but produce physically implausible predictions when pushed slightly outside it. To address this, the team implemented a Physics-Guided Neural Network, or PGNN, framework that embeds physical constraints directly into the learning process. Rather than allowing the network to fit the CFD data arbitrarily, the PGNN is penalized when its predictions violate known aerodynamic relationships, forcing the model to remain consistent with the underlying physics even in regions where training examples are sparse. This hybrid of data and physical law is part of a broader movement in computational science, exemplified by physics-informed neural networks for fluid mechanics, and it is particularly valuable in aerospace applications where a confidently wrong prediction could doom a mission design.</p>
<p>The study is careful about the boundaries of its results, which is a refreshing note of rigor in a field often prone to overclaiming. The authors state explicitly that the surrogate models are applicable only to the NACA 4-digit airfoil family they investigated and only within the operating conditions considered: a Reynolds number of 6000, a Mach number of 0.5 and a freestream turbulence intensity of 5 percent. Extending the framework to other airfoil geometries, higher Reynolds numbers or different atmospheric assumptions would require further investigation and retraining. This honesty matters because the Martian flight envelope is narrow and unforgiving; a surrogate model that quietly extrapolates beyond its validated regime could mislead an optimization loop into selecting a wing shape that fails in flight.</p>
<p>The work builds on a growing body of research into Martian rotorcraft aerodynamics. Previous studies have developed improved aerodynamic rotor models for Mars helicopters, evaluated low-Reynolds-number airfoils specifically for the Mars Helicopter rotor, and applied blade element theory coupled with CFD to optimize rotors for Mars exploration helicopters. Recent efforts have also explored machine learning for this problem, including machine learning-enhanced optimization of rotor blades for rotary-wing Mars UAVs through coupled CFD simulation and machine learning-assisted prediction of airfoil lift-to-drag characteristics for Mars helicopters. The UPES team&#8217;s contribution is to systematize this approach into a scalable CFD-ML-PGNN workflow, comparing multiple regression architectures side by side and adding physics-guided constraints, so that the pipeline can be reused and extended rather than rebuilt for each new design study.</p>
<p>Why does this matter beyond the engineering community? Aerial platforms are widely seen as the missing link in Mars exploration. Orbiters see far but cannot resolve fine detail, and rovers travel slowly across a landscape that may cover only a few kilometers over an entire mission. A drone can scout ahead of a rover, survey cliff faces, volcanic vents, polar layered deposits or candidate landing sites at centimeter scale, and reach terrain that wheels simply cannot touch. Every improvement in rotor efficiency directly extends range, endurance and payload capacity, which in turn expands the scientific return of a mission. The lift-to-drag and endurance metrics optimized in this study are not abstract numbers; they are the currency of exploration time on another planet.</p>
<p>There is also a terrestrial dividend. The ultra-low Reynolds number regime that Martian drones inhabit is the same regime occupied by small terrestrial drones, micro air vehicles and miniature surveillance platforms, all of which suffer from the same laminar separation and nonlinear stall problems. Surrogate models that predict airfoil performance cheaply and accurately at low Reynolds numbers could accelerate the design of efficient small drones on Earth, where electric multirotors and delivery UAVs face their own power-budget constraints. The methodology, generating a validated CFD database, training multiple surrogate architectures, and enforcing physical consistency through a PGNN, is a template that transfers readily to any aerodynamic design problem where simulations are expensive and the design space is large.</p>
<p>The path from this study to a flying vehicle still runs through wind tunnels, flight tests and the harsh realities of Martian atmospheric modeling, including dust, diurnal temperature swings and the CO2-dominated composition of the air. But the direction of travel is clear. As missions like Ingenuity&#8217;s successors take shape, the ability to explore thousands of airfoil and rotor configurations computationally, guided by machine learning models that respect the physics of thin, compressible, low-density flow, will compress design cycles that once took months of supercomputing into hours of surrogate-model evaluation. The researchers, working with support from the Center for Space Technology at UPES Dehradun, have demonstrated that the marriage of classical CFD and modern machine learning is not just a convenience but a genuine enabler for the next generation of aircraft designed to fly through the thin pink sky of Mars.</p>
<p><strong>Subject of Research:</strong> Machine learning-based aerodynamic optimization of UAV rotor airfoils for the Martian atmosphere</p>
<p><strong>Article Title:</strong> Machine learning based optimization of UAV wing aerodynamics in the Martian environment</p>
<p><strong>Article References:</strong> Machine learning based optimization of UAV wing aerodynamics in the Martian environment. (n.d.). <a href="https://doi.org/10.1007/s42401-026-00556-0" rel="noopener noreferrer">https://doi.org/10.1007/s42401-026-00556-0</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s42401-026-00556-0" rel="noopener noreferrer">10.1007/s42401-026-00556-0</a></p>
<p><strong>Keywords:</strong> Mars, UAV, machine learning, computational fluid dynamics, airfoil design, low Reynolds number, physics-guided neural network, rotorcraft, aerodynamics, space exploration, surrogate models, NACA airfoils</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">248122</post-id>	</item>
		<item>
		<title>New AI Explainer Finds Extreme Data Archetypes That SHAP and LIME Miss</title>
		<link>https://scienmag.com/new-ai-explainer-finds-extreme-data-archetypes-that-shap-and-lime-miss/</link>
		
		<dc:creator><![CDATA[Blake Davidson]]></dc:creator>
		<pubDate>Tue, 06 Oct 2026 02:48:15 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[AI decision-making transparency tools]]></category>
		<category><![CDATA[Applied Intelligence]]></category>
		<category><![CDATA[archetypal analysis]]></category>
		<category><![CDATA[archetypal analysis in machine learning]]></category>
		<category><![CDATA[archetype profiles in data science]]></category>
		<category><![CDATA[archetype-based data analysis]]></category>
		<category><![CDATA[black-box models]]></category>
		<category><![CDATA[convex mixture models]]></category>
		<category><![CDATA[data decomposition techniques]]></category>
		<category><![CDATA[data science]]></category>
		<category><![CDATA[decision boundary understanding in AI models]]></category>
		<category><![CDATA[explainable AI]]></category>
		<category><![CDATA[extreme data archetypes]]></category>
		<category><![CDATA[feature ranking]]></category>
		<category><![CDATA[global data architecture visualization]]></category>
		<category><![CDATA[interpretability in healthcare AI]]></category>
		<category><![CDATA[LIME]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[model interpretability]]></category>
		<category><![CDATA[SHAP]]></category>
		<category><![CDATA[SHAP and LIME limitations]]></category>
		<category><![CDATA[statistical rigor]]></category>
		<category><![CDATA[surrogate models]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=240002</guid>

					<description><![CDATA[A new framework called ARCHEX uses archetypal analysis to reveal global structural patterns in machine learning models that local explanation tools like SHAP and LIME often miss.]]></description>
										<content:encoded><![CDATA[<p>Machine learning models now make decisions in hospitals, banks, and factories, but the tools we use to peer inside them have a blind spot. The dominant explanation techniques, SHAP and LIME, work locally: they take a single prediction and trace which features pushed it one way or another. What they do not reveal is the global architecture of the data itself—the extreme, archetypal profiles that anchor the decision boundary. A new framework called ARCHEX, published in Applied Intelligence by Abraham Itzhak Weinberg, an independent researcher at AI-WEINBERG in Tel Aviv, sets out to fill that gap by borrowing a mathematical idea from the 1990s and pressing it into service for modern explainable artificial intelligence.</p>
<p>ARCHEX, short for ARCHetype-based EXplainer, is built on archetypal analysis, a decomposition technique introduced by Adele Cutler and Leo Breiman in 1994. The method represents every data point as a convex mixture of a small number of extreme profiles, or archetypes, that sit on the boundary of the data cloud. Instead of asking which cluster a point belongs to, archetypal analysis asks which archetypes it is a blend of. A patient record, for example, might be expressed as sixty percent of one extreme profile and forty percent of another. Weinberg&#8217;s insight is that these interpretable extremes can serve as a compressed coordinate system for the entire dataset, reducing thousands of raw features to a handful of meaningful dimensions.</p>
<p>The technical pipeline is deliberately simple. ARCHEX first identifies k archetypes from the training data, where k is chosen adaptively and only on the training partition to avoid information leaking into evaluation. Every observation is then projected onto the probability simplex, meaning it receives a set of non-negative membership weights across the archetypes that sum to one. These k-dimensional representations become the inputs to a linear surrogate model trained to reproduce the original black-box model&#8217;s prediction target. Because the surrogate is linear and non-black-box, its coefficients can be read directly as a global feature ranking, giving analysts a transparent approximation of how the underlying model behaves across the whole data distribution rather than at a single point.</p>
<p>One of the paper&#8217;s most technically interesting contributions is a precise characterization of where ARCHEX&#8217;s sparsity comes from. Sparse explanations—ones that highlight only a few features—are prized in interpretability research, and many methods engineer them through L1 regularization, which penalizes the sum of absolute coefficient values. Weinberg shows that ARCHEX needs no such penalty. Because the archetype membership weights lie on the probability simplex, their L1 norm is constant by construction: the weights always sum to one. Sparsity therefore emerges from the simplex projection itself, a geometric constraint rather than a tuning knob. This observation connects the framework to efficient projection algorithms onto the L1 ball and gives the method a form of built-in parsimony that does not have to be traded off against accuracy.</p>
<p>The evaluation is notable for its statistical caution, a quality often missing from explainability research. ARCHEX was tested on five public benchmark datasets spanning four domains, drawn from standard repositories such as UCI and scikit-learn. Rather than reporting single point estimates, the study uses bootstrap confidence intervals on both predictive performance and on the rank correlation between ARCHEX&#8217;s archetype-derived feature ranking and SHAP&#8217;s local attribution ranking. The results are striking: in four of the five datasets, that rank correlation is small in magnitude and, once sampling uncertainty is quantified, statistically indistinguishable from zero. In other words, there is no reliable evidence that ARCHEX is simply rediscovering what SHAP already tells you.</p>
<p>The fifth dataset complicates the story in an instructive way. On the Digits dataset, the confidence interval excludes zero, and the correlation between the two rankings is moderate and positive rather than low. Weinberg argues that this cuts against a common practice in the field: treating a low correlation point estimate as established proof that a new method offers complementary explanatory content. Without confidence intervals, a researcher might see a weak correlation and claim novelty; with proper uncertainty quantification, the claim may dissolve. The paper thus doubles as a methodological warning about how interpretability methods are compared, echoing earlier sanity-check studies that questioned whether saliency methods measure what they claim.</p>
<p>Beyond the ranking analysis, the paper includes an ablation that tests the value of soft membership. ARCHEX assigns each point a convex blend of archetype weights, while a hard variant based on k-means cluster membership forces each point into a single cluster. Across all five datasets, the soft convex membership outperformed the hard ablation on held-out predictive accuracy. The result makes intuitive sense: real data rarely falls neatly into discrete buckets, and allowing partial membership preserves geometric information that hard assignments discard. For practitioners, it suggests that the smoothness of the archetype representation, not merely the choice of extreme profiles, is doing real work in approximating the decision boundary.</p>
<p>Among the concrete findings, the Breast Cancer Wisconsin dataset offers the most vivid illustration. ARCHEX&#8217;s top-ranked features there are measurements related to concavity and concave points—shape characteristics of cell nuclei that describe how irregular a cell&#8217;s outline is. Several of these features are ranked far lower by SHAP. Weinberg is careful to present this as a preliminary, descriptive observation rather than a validated clinical finding, and that restraint matters: no prospective clinical study supports a diagnostic claim, and the author explicitly frames the result as a hypothesis-generating signal. Still, it shows how a global, archetype-based lens can surface feature relationships that local attribution methods, focused on individual predictions, may systematically underweight.</p>
<p>Where does ARCHEX fit in the crowded landscape of explainable AI? Weinberg positions it as a global data-structure and subgroup-discovery method that complements, rather than replaces, local attribution tools. SHAP and LIME answer the question of why this particular prediction was made; ARCHEX answers which extreme profiles define the data and how the model&#8217;s boundary behaves across them. This distinction echoes a broader debate in the field, from Cynthia Rudin&#8217;s argument that high-stakes decisions should rely on inherently interpretable models to concept-based approaches like TCAV that look beyond per-feature attributions. ARCHEX adds a geometric, archetype-centered voice to that conversation, grounded in a decomposition technique with a three-decade pedigree.</p>
<p>The framework does come with caveats. The core optimization implementation is proprietary and under development for commercial use, so the source code is not publicly available—a limitation for a paper whose central claim is about statistical rigor and reproducibility. To mitigate this, the manuscript provides complete algorithmic pseudocode, optimization hyperparameters, preprocessing procedures, and a machine-readable archive of the numerical results behind the tables and figures, allowing independent verification of the reported outcomes even without the exact code. Whether ARCHEX&#8217;s archetype lens becomes a standard complement to SHAP will depend on replication by other groups, but the paper&#8217;s insistence on confidence intervals before claiming complementarity is a standard the rest of explainable AI would do well to adopt.</p>
<p><strong>Subject of Research:</strong> Archetype-based global interpretability framework for explaining machine learning models</p>
<p><strong>Article Title:</strong> ARCHEX: explaining models through archetypal decomposition for discovering complementary patterns in model interpretability</p>
<p><strong>Article References:</strong> Weinberg, A. I. (2026). ARCHEX: explaining models through archetypal decomposition for discovering complementary patterns in model interpretability. <em>Applied Intelligence, 56</em>(15), Article 475. <a href="https://doi.org/10.1007/s10489-026-07506-5" rel="noopener noreferrer">https://doi.org/10.1007/s10489-026-07506-5</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s10489-026-07506-5" rel="noopener noreferrer">10.1007/s10489-026-07506-5</a></p>
<p><strong>Keywords:</strong> explainable AI, archetypal analysis, model interpretability, SHAP, LIME, machine learning, statistical rigor, feature ranking, surrogate models, Applied Intelligence, data science, black-box models</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">240002</post-id>	</item>
		<item>
		<title>Auditable AI Software Brings Trustworthy Neural Operators to Composite Curing</title>
		<link>https://scienmag.com/auditable-ai-software-brings-trustworthy-neural-operators-to-composite-curing/</link>
		
		<dc:creator><![CDATA[Cassandra Pierce]]></dc:creator>
		<pubDate>Fri, 02 Oct 2026 06:43:49 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[AI auditability in industrial applications]]></category>
		<category><![CDATA[carbon-fiber composite curing process optimization]]></category>
		<category><![CDATA[causality]]></category>
		<category><![CDATA[composite curing]]></category>
		<category><![CDATA[composite material manufacturing process modeling]]></category>
		<category><![CDATA[convergence testing]]></category>
		<category><![CDATA[Fourier Neural Operator]]></category>
		<category><![CDATA[Fourier neural operators for process simulation]]></category>
		<category><![CDATA[high-fidelity thermal simulation using neural networks]]></category>
		<category><![CDATA[manufacturing simulation]]></category>
		<category><![CDATA[Neural operator surrogates for composite curing]]></category>
		<category><![CDATA[neural operators]]></category>
		<category><![CDATA[open-source AI software for thermal process control]]></category>
		<category><![CDATA[open-source software]]></category>
		<category><![CDATA[provenance]]></category>
		<category><![CDATA[Python-based scientific AI tools]]></category>
		<category><![CDATA[reproducibility]]></category>
		<category><![CDATA[reproducibility and provenance in AI software]]></category>
		<category><![CDATA[rigorous verification of neural network models]]></category>
		<category><![CDATA[software verification]]></category>
		<category><![CDATA[surrogate modeling for thermochemical processes]]></category>
		<category><![CDATA[surrogate models]]></category>
		<category><![CDATA[thermochemical modeling]]></category>
		<category><![CDATA[trustworthy AI in manufacturing]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=226230</guid>

					<description><![CDATA[Researchers have released CD-CureNO, an open-source software package that wraps neural-operator surrogates for composite curing in a rigorous framework of provenance tracking, restriction-preserving transfer, causal verification and independently recomputed metrics.]]></description>
										<content:encoded><![CDATA[<p>Manufacturing a carbon-fibre composite part is a delicate thermal balancing act. The resin must be heated along a carefully programmed cycle, and the way heat diffuses through the laminate determines how completely the polymer cures, whether temperatures overshoot their targets, and how steep the conversion gradients become inside the material. High-fidelity simulations of this thermochemical process are indispensable for process design, but they are expensive to run repeatedly. Neural operators—deep learning models that learn mappings between entire functions rather than single values—promise cheap surrogates, and Fourier neural operators in particular have become an influential template. The catch is trust: in a scientific setting, a surrogate that cannot be audited is a surrogate that cannot be adopted.</p>
<p>A team led by Hyojin Park, Nam-Hyun Yoo and Jinhong Yang has now released CD-CureNO, an open-source software package published in the journal SoftwareX that wraps neural-operator surrogates for composite curing in an unusually rigorous verification and provenance framework. The software, versioned at v0.0.3 and archived in Software Heritage under an Apache 2.0 license, is written in Python 3.10 through 3.12 using PyTorch, NumPy, SciPy and pandas, and it runs its audits and test suites on ordinary CPUs, with CUDA reserved for optional model training. What distinguishes the project is not a claim of new mathematics—the authors explicitly disclaim novelty for Fourier operators, joint temperature-and-cure prediction, or physics-based losses—but a domain-specific integration and verification layer built around inspectable scientific controls.</p>
<p>The motivating problem came from a legacy residual Fourier neural operator that stacked up to 51 independently trained positional models for curing analysis. Rather than silently replacing that code, CD-CureNO keeps two reproduction paths separately named: one that reproduces the legacy behavior exactly, and one that corrects known defects. The team regression-tested nine legacy defects, including normalization leakage and incomplete field outputs, so that the historical behavior is documented rather than erased. From that audited foundation, the software extends to joint temperature-and-degree-of-cure fields, causal response, conservative label generation, and transfer of learned knowledge across spatial dimensions.</p>
<p>The architecture is organized around explicit development and validation phases, labeled P0 through P6. P0 audits legacy behavior, P1 reproduces and corrects it, P2 develops joint one-plus-one-dimensional operators over the thickness-time field, P3 supplies a one-dimensional solver stage, P4 generates a true two-dimensional benchmark with two spatial coordinates evolving in time, P5 tests transfer, restriction and causality, and P6 is reserved for a future confirmatory held-out campaign. Each stage emits inspectable artifacts—configuration snapshots, Git and environment records, checkpoints, predictions and machine-readable receipts—that the next stage consumes, with cross-cutting guardrails verified again during postflight aggregation and reporting.</p>
<p>The authors define a result as auditable when its input, code, configuration, split, seed and checkpoint are all identifiable; when its reported metric can be recalculated from saved case outputs; when integrity violations or missing prerequisites fail explicitly; and when every claim can be traced back to those artifacts. Hashes, manifests, tensor mappings and receipts supply the observable evidence. Failed and superseded runs remain on record instead of being replaced by result-selected retries, and complete cure cycles or simulation cases—not individual grid points—define the samples. Crucially, the team distinguishes its independent postflight calculations, which are separate from the training and execution path, from validation by an unaffiliated research group or an independent physical model. A new NumPy-only example verifier imports neither training nor model code and reconstructs metrics from saved fields; a deliberately altered metric is rejected even after its file checksum is updated, and a failed recheck replaces any earlier success receipt.</p>
<p>Two technical contracts give the software its distinctive character. The first is restriction preservation: for laterally uniform coefficients, initial conditions and boundary data with zero lateral flux, an x-independent solution should reduce to the same through-thickness problem at every lateral position. The two-dimensional target therefore implements the lifted one-dimensional field plus a residual correction, with copied source tensors and zero-residual lateral branches that recover the source mapping exactly at initialization up to floating-point ordering. A verifier reconstructs the target state from the checkpoint, since output equality alone would not prove provenance. The second contract is causality: the causal model replaces temporal Fourier transforms with dilated convolutions, and perturbing all future input channels must leave every preceding output field unchanged. A nonnegative latent cure rate guarantees bounded, monotone cure evolution, and finite-volume-compatible fluxes handle material interfaces.</p>
<p>The verification evidence is unusually detailed. A GitHub Actions run on Ubuntu 24.04 passed 224 tests with 30 explicitly documented skips in about 42 seconds, and a fresh Windows run on a separate commit passed the same counts in about 48 seconds; continuous integration even exposed an input-stride-dependent CPU rounding difference, fixed by making shared target channels contiguous, with a new regression test added. The historical two-dimensional benchmark holds 512 temperature-and-cure cases of roughly 918 megabytes, anchored by 27 prespecified solver checks, zero failed cases, a maximum relative global energy residual of 1.345 times ten to the minus twelve, and exact float32 slice-hash replay for selected cases. Fresh regeneration of all 512 cases took 214.5 seconds on the documented Windows machine, and independent hashing confirmed that all eight output arrays match the frozen originals byte for byte. The team is candid that strict Linux plan validation rejected 33 heat-transfer values differing from the frozen Windows plan by at most 2.842 times ten to the minus fourteen, so canonical replay requires the documented compatible environment.</p>
<p>New in the revision is an independent mathematical verification of the production solver against elementary continuous reference solutions. Four synthetic case families—an anisotropic insulated transient, steady manufactured fields exercising all four Robin boundary faces with uniform and then discontinuous conductivity, and a solvable cure specialization coupled to reaction heating—were specified in an archived protocol with acceptance criteria fixed before execution. All 17 scheduled configurations and 72 recorded acceptance items passed. Measured spatial convergence orders clustered near two, temporal orders near one, consistent with backward Euler, and the synthetic cure equation gave orders near four, consistent with the production fourth-order Runge-Kutta update. Finest temperature relative errors ranged from 5.82 times ten to the minus five to 5.10 times ten to the minus four, and the finest cure maximum error was 1.80 times ten to the minus eight. The authors stress the boundaries: these studies verify anisotropic diffusion, boundary treatment and a special reaction-heat coupling, but not general kinetic validity, proprietary solver parity or experimental validation.</p>
<p>A fully public end-to-end example demonstrates the entire workflow without any private data or checkpoints. It generates eight true two-dimensional cases with laterally varying Robin coefficients, trains production causal operators with a fixed seed for 32 source and 48 target epochs, transfers knowledge through tensor copying and zero-residual initialization, and then recomputes held-out metrics in a separate NumPy-only postflight process. The two held-out cases yielded temperature relative errors of 0.006408 and 0.009741, temperature root-mean-square errors of 2.199 and 3.385 kelvin, and cure errors near 0.028, with future-prefix change exactly zero and independently recomputed metrics agreeing to within 4.44 times ten to the minus sixteen. The complete run took under ten seconds on a desktop CPU. The authors deliberately label this a functionality example, not evidence of superiority, and the figures show visible prediction errors rather than flattering comparisons.</p>
<p>Looking forward, the team has published a registered plan for a confirmatory P6 campaign: 40 held-out out-of-distribution cases across five target seeds, paired with a generic causal transfer baseline and analyzed with 10,000 crossed paired bootstrap replicates. A confirmatory claim would require a registered positive 95 percent improvement-interval bound, at least a 10 percent case-median improvement, physical-error guardrails and complete verified execution without test-based selection—and the authors note plainly that 40 cases across five seeds are not 200 independent cases and do not guarantee adequate statistical power. That discipline is the real story of CD-CureNO. In a field where surrogate models often arrive with cherry-picked demos and unverifiable numbers, the package demonstrates that neural operators for manufacturing can ship with the same provenance discipline as the experiments they hope to replace: every number traceable to a run, every failed attempt preserved, and every claim bounded by the evidence that produced it.</p>
<p><strong>Subject of Research:</strong> Auditable restriction-preserving causal neural operators for thermochemical composite curing simulation</p>
<p><strong>Article Title:</strong> CD-CureNO: Auditable software for restriction-preserving causal neural operators in composite curing</p>
<p><strong>Article References:</strong> Park, H., Yoo, N.-H., &amp; Yang, J. (2026). CD-CureNO: Auditable software for restriction-preserving causal neural operators in composite curing. <em>SoftwareX, 36</em>, Article 103086. <a href="https://doi.org/10.1016/j.softx.2026.103086" rel="noopener noreferrer">https://doi.org/10.1016/j.softx.2026.103086</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> Not provided</p>
<p><strong>Keywords:</strong> neural operators, composite curing, Fourier neural operator, software verification, reproducibility, provenance, causality, thermochemical modeling, surrogate models, manufacturing simulation, open-source software, convergence testing</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">226230</post-id>	</item>
		<item>
		<title>How Evolutionary Algorithms Learn to Juggle Conflicting Goals and Hard Constraints</title>
		<link>https://scienmag.com/how-evolutionary-algorithms-learn-to-juggle-conflicting-goals-and-hard-constraints/</link>
		
		<dc:creator><![CDATA[Gavin Prescott]]></dc:creator>
		<pubDate>Fri, 02 Oct 2026 04:36:18 +0000</pubDate>
				<category><![CDATA[Earth Science]]></category>
		<category><![CDATA[algorithmic approaches to constrained problems]]></category>
		<category><![CDATA[auxiliary populations]]></category>
		<category><![CDATA[balancing multiple objectives in engineering]]></category>
		<category><![CDATA[benchmark test problems]]></category>
		<category><![CDATA[conflicting goals in engineering design]]></category>
		<category><![CDATA[constrained multi-objective optimization]]></category>
		<category><![CDATA[constraint handling]]></category>
		<category><![CDATA[differential evolution]]></category>
		<category><![CDATA[epsilon constrained method]]></category>
		<category><![CDATA[evolutionary algorithm techniques]]></category>
		<category><![CDATA[evolutionary algorithms]]></category>
		<category><![CDATA[evolutionary algorithms for multi-objective optimization]]></category>
		<category><![CDATA[handling inequality and equality constraints]]></category>
		<category><![CDATA[hard constraints in optimization problems]]></category>
		<category><![CDATA[multi-task optimization]]></category>
		<category><![CDATA[NSGA-II]]></category>
		<category><![CDATA[optimization of aircraft wing design]]></category>
		<category><![CDATA[Pareto front]]></category>
		<category><![CDATA[power grid resource allocation]]></category>
		<category><![CDATA[push and pull search]]></category>
		<category><![CDATA[surrogate models]]></category>
		<category><![CDATA[trade-offs in engineering solutions]]></category>
		<category><![CDATA[vehicle fleet scheduling algorithms]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=225718</guid>

					<description><![CDATA[A comprehensive new review maps six families of constrained multi-objective evolutionary algorithms, their constraint handling techniques, benchmark tests, and the open challenges shaping the field's future.]]></description>
										<content:encoded><![CDATA[<p>Some of the hardest problems in modern engineering do not have a single right answer. Designing an aircraft wing, scheduling a fleet of vehicles, or allocating power across a grid all involve objectives that fight against each other: improving one metric tends to worsen another. On top of that, real-world designs must obey hard rules, from material strength limits to emission caps, that no amount of clever trade-offs can break. A new comprehensive review published in the open-access journal Vicinagearth by Jing Liang, Hongyu Lin, Caitong Yue, Xuanxuan Ban, and Kunjie Yu of Zhengzhou University maps the entire landscape of algorithms built to solve these so-called constrained multi-objective optimization problems, and it arrives at a moment when the field is both maturing rapidly and facing a fresh wave of challenges.</p>
<p>The mathematical skeleton of such a problem is deceptively simple. There is a vector of decision variables, a set of objective functions to be minimized simultaneously, and two families of constraints: inequality constraints that must not be exceeded and equality constraints that must be met exactly. A solution is called feasible if it violates none of the constraints, and the total constraint violation is measured as the sum of the individual violations across every rule. What makes these problems treacherous is that the objectives usually conflict, so no single solution can be best at everything. Instead, the goal is to find a whole set of solutions that are mutually non-dominated, meaning no member of the set is better than another in every objective. This collection of trade-off points, projected into objective space, forms the Pareto front, and when constraints are added, the algorithm must chase the constrained Pareto front, a potentially very different target.</p>
<p>The review highlights a crucial subtlety that often surprises newcomers: the unconstrained Pareto front and the constrained Pareto front can relate to each other in four distinct ways. In the easiest case, the two fronts completely overlap, so every optimal trade-off point is already feasible. In harder cases, the constrained front is a subset of the unconstrained one, or the two only partially overlap, or, most brutally, they are entirely separate, with the unconstrained front lying wholly in infeasible territory. This taxonomy matters because many algorithms work by first pushing a population toward the unconstrained front and then pulling it back toward feasibility. That strategy shines when the two fronts are close, but it can fail badly when they point in different directions, a lesson the authors demonstrate with new experimental comparisons.</p>
<p>At the heart of every constrained multi-objective evolutionary algorithm sit two components: a multi-objective evolutionary engine, which may be dominance-based, decomposition-based, or indicator-based, and a constraint handling technique that decides how to weigh feasible against infeasible solutions during selection. The review walks through the classic techniques with unusual clarity. Penalty function methods fold constraint violation into the fitness value, and the entire art lies in choosing the penalty coefficient: too small and the population never finds feasible solutions, too large and it gets trapped in the first feasible pocket it stumbles into. Static, dynamic, and adaptive variants each try to thread this needle differently.</p>
<p>The constrained domination principle, introduced by Deb and colleagues in the landmark NSGA-II algorithm, takes a more direct route: any feasible solution beats any infeasible one, two infeasible solutions are compared by their degree of violation, and two feasible solutions are compared by ordinary Pareto dominance. It is simple, easy to embed, and converges fast, but its bias toward feasibility makes it hard for a population to cross large infeasible barriers that may separate it from the true constrained front. The epsilon constrained method, due to Takahama and Sakai, softens this by relaxing the constraint boundary with a tolerance parameter, treating mildly infeasible solutions as if they were feasible. As epsilon shrinks toward zero over the run, the method gradually recovers the strict rules while harvesting useful information from high-quality infeasible solutions along the way. Stochastic ranking adds a probabilistic twist, occasionally ignoring constraints entirely with a tunable probability, while multi-objective methods go further and promote constraint violation itself to the status of an extra objective, though this risks inflating the problem into a many-objective one where selection pressure evaporates.</p>
<p>Beyond these foundational techniques, the review organizes the modern algorithm zoo into six families, and this classification is arguably its most useful contribution. Some methods design new fitness functions that blend objectives and constraint violations with adaptive or dynamic weights, often using the proportion of feasible solutions in the population as a feedback signal. Others enhance the classical constraint handling techniques themselves, for example by folding angle information into the constrained domination principle so that high-quality infeasible solutions survive selection and help the population leap across infeasible regions. Constraint relaxation methods tune the epsilon boundary dynamically, sometimes using the maximum and minimum violation values observed among infeasible individuals, and even deploy detect-and-escape mechanisms that relax constraints when the population shows signs of evolutionary stagnation.</p>
<p>Two of the six families have produced some of the field&#8217;s most striking recent successes. Two-stage optimization methods split the run into phases: the push-and-pull search framework, for instance, ignores all constraints in the first phase so the population can race to the unconstrained front unimpeded, then uses an improved epsilon method to pull it back toward the constrained front. Auxiliary population methods run a second population alongside the main one, often letting the helper ignore constraints entirely so it can scout the unconstrained front and transfer that knowledge back. The coevolutionary framework CCMO, which pairs a feasibility-focused main population with a constraint-free helper, and the multitasking approach MTCMO, which treats the helper as a separate optimization task with an improved epsilon mechanism, both exemplify this cooperative philosophy. A sixth family rewrites the reproduction operators themselves, using differential evolution mutation strategies that explicitly exploit promising infeasible solutions to generate offspring with better convergence and diversity.</p>
<p>The review does not stop at taxonomy; it also runs head-to-head experiments. Six representative algorithms, including DSPCMDE, MOEADDAE, PPS, CCMO, MTCMO, and CMOCSO, were tested on three benchmark suites with contrasting personalities: LIR-CMOP, whose narrow feasible regions are riddled with infeasible blocks; MW, with tiny, discontinuous feasible regions defined in objective space; and SDC, a newer suite where constraints and objectives share a chaotic, non-monotonic relationship. Each algorithm ran twenty times per problem with a population of one hundred and one hundred thousand evaluations, scored by the inverted generational distance metric on the PlatEMO platform. The results are instructive rather than triumphant for any single method. CMOCSO, with its competitive-and-cooperative swarm operators, dominated LIR-CMOP, taking best or second-best on thirteen of fourteen problems. CCMO and MTCMO led on MW, confirming that mining unconstrained-front information pays off there. But on SDC, where chasing the unconstrained front can actually move the population away from the constrained one, DSPCMDE&#8217;s redesigned fitness function proved superior. Friedman&#8217;s test ranked CMOCSO first overall and MTCMO second, yet the differences were modest, underscoring the review&#8217;s central message that no algorithm wins everywhere and that matching strategy to problem structure is the real game.</p>
<p>The benchmark section itself doubles as a history of the field&#8217;s growing rigor, from the pioneering SRN, TNK, and OSY problems of the 1990s through the adjustable-difficulty CTP and DAS-CMOP suites to recent constructions like ZXH_CF, DCMOP with its deceptive constraints that punish solutions precisely where they seem closest to the front, and the real-world RWCMOP collection of fifty engineering problems spanning mechanical design, chemical engineering, power electronics, and power systems. This progression reflects a deliberate effort to expose algorithms to narrow feasible regions, disconnected fronts, multimodality, and deception, features that older, gentler benchmarks simply never tested.</p>
<p>Looking forward, the authors identify five frontiers where current methods strain. Computationally expensive problems, such as those involving computational fluid dynamics, allow only a few thousand real evaluations, making surrogate models that mimic the fitness landscape essential. Multimodal problems hide multiple distinct sets of Pareto-optimal solutions mapping to the same front, demanding strong diversity preservation and mode recognition. Large-scale problems, with hundreds or thousands of decision variables, require dimension reduction and variable grouping to avoid drowning in the search space. Multi-task optimization, which could let experience transfer between related constrained problems, still needs principled knowledge-transfer mechanisms. And dynamic problems, where objectives or constraints shift over time as in fluid catalytic cracking or mineral beneficiation, force algorithms to detect environmental change in real time and decide whether yesterday&#8217;s population is worth keeping. With the field already logging thousands of accesses and dozens of citations for this review, its map of what works, what fails, and what remains unsolved is likely to steer the next generation of algorithms toward exactly these open problems.</p>
<p><strong>Subject of Research:</strong> Evolutionary algorithms for constrained multi-objective optimization problems</p>
<p><strong>Article Title:</strong> Evolutionary constrained multi-objective optimization: a review</p>
<p><strong>Article References:</strong> Liang, J., Lin, H., Yue, C., Ban, X., &amp; Yu, K. (2024). Evolutionary constrained multi-objective optimization: a review. <em>Vicinagearth, 1</em>(1), Article 5. <a href="https://doi.org/10.1007/s44336-024-00006-5" rel="noopener noreferrer">https://doi.org/10.1007/s44336-024-00006-5</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s44336-024-00006-5" rel="noopener noreferrer">10.1007/s44336-024-00006-5</a></p>
<p><strong>Keywords:</strong> constrained multi-objective optimization, evolutionary algorithms, constraint handling, Pareto front, NSGA-II, epsilon constrained method, benchmark test problems, push and pull search, auxiliary populations, differential evolution, surrogate models, multi-task optimization</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">225718</post-id>	</item>
		<item>
		<title>New Attack Forges Neural Network Watermarks Without Touching the Victim Model</title>
		<link>https://scienmag.com/new-attack-forges-neural-network-watermarks-without-touching-the-victim-model/</link>
		
		<dc:creator><![CDATA[Blake Davidson]]></dc:creator>
		<pubDate>Tue, 22 Sep 2026 13:40:53 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[adversarial transferability]]></category>
		<category><![CDATA[adversarial watermark forgery]]></category>
		<category><![CDATA[AI model intellectual property rights]]></category>
		<category><![CDATA[behavior-only interface security flaws]]></category>
		<category><![CDATA[black-box model watermarking vulnerabilities]]></category>
		<category><![CDATA[CIFAR-10]]></category>
		<category><![CDATA[cybersecurity]]></category>
		<category><![CDATA[cybersecurity threats to model watermarking]]></category>
		<category><![CDATA[deep learning model copyright protection]]></category>
		<category><![CDATA[deep learning model theft countermeasures]]></category>
		<category><![CDATA[deep neural network security]]></category>
		<category><![CDATA[fake model ownership verification]]></category>
		<category><![CDATA[FakeMark]]></category>
		<category><![CDATA[false ownership claims]]></category>
		<category><![CDATA[feature fusion]]></category>
		<category><![CDATA[ImageNet]]></category>
		<category><![CDATA[intellectual property protection]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[model ownership proof bypass methods]]></category>
		<category><![CDATA[model watermark fabrication techniques]]></category>
		<category><![CDATA[model watermarking]]></category>
		<category><![CDATA[neural network theft prevention]]></category>
		<category><![CDATA[Neural network watermarking attack]]></category>
		<category><![CDATA[surrogate models]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=205395</guid>

					<description><![CDATA[Researchers have unveiled FakeMark, an attack that fabricates convincing neural network watermark evidence without ever accessing the victim model, exposing a structural weakness in behavior-only ownership verification.]]></description>
										<content:encoded><![CDATA[<p>Deep learning models can cost millions of dollars to train, and as stolen or pirated copies of these models spread across the internet, researchers have increasingly turned to watermarking as a way to prove ownership. The idea is elegant: a model owner embeds secret trigger inputs into a network during training, so that later, in a dispute, the owner can show that the suspect model responds to those triggers in a predictable way. If the model answers the secret questions correctly, the argument goes, it must be the stolen property. But a new study published in the journal Cybersecurity shows just how fragile that logic can be. An attack framework called FakeMark demonstrates that an adversary who never touches the victim model can fabricate evidence that looks exactly like a legitimate watermark, achieving near-perfect agreement on the behavioral tests that many ownership protocols rely on.</p>
<p>The research, led by Yutong Wu and Songfeng Lu of Huazhong University of Science and Technology together with colleagues at Guangxi Normal University, Chongqing University, and other institutions, targets a structural weakness in what the authors call the behavior-only interface of model watermarking. Most black-box watermarking schemes, whether they use backdoored trigger sets, adversarial frontier stitching, dynamic adversarial watermarking, or parameter encoding, ultimately verify ownership by checking whether a suspect model maps certain inputs to certain labels. The problem, the researchers argue, is that a successful trigger response does not reveal how that response came to be. A third-party arbiter sees the model answering the secret key correctly, but cannot distinguish between a model that was genuinely watermarked at training time and one that simply happens to respond to a cleverly constructed set of forged inputs.</p>
<p>FakeMark exploits this ambiguity through a two-stage pipeline that operates entirely on a surrogate model. In the first stage, called clean feature caching, the attacker runs ordinary, unmodified images through a publicly available surrogate network and records the outputs of selected convolutional and fully connected layers. These cached activations serve as a library of clean reference features. In the second stage, adversarial feature fusion, the attacker optimizes a bounded perturbation on a batch of inputs while repeatedly injecting randomly permuted copies of the cached clean references into the surrogate&#8217;s forward pass. The injection is stochastic: at each iteration, a Bernoulli gate decides whether a given layer participates, a fresh permutation shuffles the reference features across the batch, and per-channel interpolation weights are drawn uniformly from a fixed range. The perturbation itself is guided by the gradient of the target-class logit, so the inputs are pushed toward whatever internal representations make the surrogate predict the attacker&#8217;s chosen label.</p>
<p>The crucial design insight is that this stochastic fusion prevents the perturbation from overfitting to architecture-specific quirks of the surrogate. Because the forward map changes at every iteration, successive gradients need not follow a single narrow trajectory, and the resulting adversarial inputs generalize better to unseen victim architectures. The authors support this intuition with a simplified theoretical analysis. Under a regularized linear classifier, they prove that a bounded perturbation maximizing the margin toward a watermarked class necessarily acquires a nonzero component along the watermark-trigger direction, with an explicit lower bound on the fraction of the perturbation&#8217;s norm that lies in that direction. The result is deliberately conditional: it explains why a successful linear watermark creates a trigger-aligned vulnerability, but the authors are careful to note that it does not prove the nonlinear optimization recovers the actual trigger of a deep network.</p>
<p>The empirical results are striking, though the authors frame them with unusual care. Across sixteen distinct architectures, an additional adversarially trained ResNet-50 checkpoint, and eight watermark variants spanning model-independent, model-dependent, active, and parameter-encoding schemes, the retrospective best-case behavioral target-label accuracy reaches 1.00 on CIFAR-10 and 0.99 on ImageNet. In other words, when the best surrogate in the pool is used, forged keys can induce the victim model to answer with the attacker&#8217;s chosen label on essentially every verification sample. On a fixed ADI-watermarked target under a matched protocol, FakeMark achieved a targeted attack success rate of 0.990, compared with 0.860 to 0.940 for canonical transferable attacks such as PGD, MI-FGSM, and DI-FGSM. A controlled component analysis showed that removing the batch permutation or the channel-wise mixing reduced transfer, and that replacing the clean cache with current-batch features actually raised surrogate-side success while lowering victim-side accuracy, indicating that the stochastic fusion genuinely shapes cross-model behavior rather than merely optimizing the surrogate.</p>
<p>Yet the study is equally notable for what it does not claim. ImageNet transfer varied enormously across checkpoint and surrogate settings, ranging from near zero to 0.99. Two ImageNet surrogates, ResNet-18 and VGG16, produced almost no transfer at all, while the adversarially trained ResNet-50 checkpoint and DenseNet-121 performed strongly. The ordering was not explained by model size or computational cost: MobileNetV2, the smallest and cheapest surrogate tested, matched ResNet-50 at 0.990, while Inception-v3, with the largest FLOP count, reached only 0.727. Architecture and checkpoint compatibility, the authors conclude, is a central condition for the attack, and the reported maxima are retrospective upper bounds over the surrogate pool rather than an operational guarantee that an attacker could pick a good surrogate without any victim feedback. Auxiliary experiments on clean, unwatermarked models showed comparable targeted transfer in several settings, so the evidence does not establish a universal watermark-specific amplification effect.</p>
<p>The threat model is deliberately austere. The attacker has white-box gradient access to one or more surrogate models trained on public data, but during attack construction has no query access, no gradients, no parameters, and no architectural knowledge of the victim. The victim model is invoked only at the end, through the normal ownership-verification procedure. The attacker does not modify the victim&#8217;s parameters, does not observe the legitimate watermark key, and does not need to estimate any victim-specific decision threshold. On CIFAR-10, the surrogates and watermarked targets share the standard public training split, so the controlled setting involves substantial sample-level overlap; on ImageNet, only dataset-level overlap can be asserted because external pretrained checkpoints do not publish training manifests. The authors acknowledge these limitations explicitly, along with the fact that the evaluation is restricted to image classification and uses a single surrogate at a time.</p>
<p>The study also probes whether forged keys can be detected. Perceptual measurements showed that the perturbations remain bounded, with PSNR between roughly 24 and 27 decibels and structural similarity scores between 0.72 and 0.91 depending on the setting. Input-level screening told a mixed story: the ADI scheme&#8217;s forged keys stayed below all four diagnostic thresholds tested, while Frontier and DAWN exceeded three of them, and median filtering changed forged predictions more often than legitimate ones. Feature-based detectors, including Mahalanobis distance, local intrinsic dimensionality, and ODIN, flagged forged groups at higher rates than clean calibration data, with the largest increase under LID. The authors interpret these results as evidence of detector-dependent artifacts rather than broad evasion, and they caution that the detector study is not a fully tuned learned-detector benchmark. The picture that emerges is nuanced: some schemes leave fingerprints on forged evidence, others do not, and no single screening method reliably separates genuine from fabricated triggers.</p>
<p>The practical implications reach beyond the laboratory. As models are increasingly deployed through APIs, fine-tuned by third parties, and extracted by adversaries, ownership disputes will hinge on verification evidence that courts, marketplaces, and standards bodies can trust. FakeMark shows that behavioral evidence alone cannot carry that weight, because a successful trigger response is compatible with both legitimate embedding and independent forgery. The authors argue that future ownership protocols should pair behavioral tests with signals that are harder to construct after the fact: cryptographic commitments to watermark keys or training provenance, structural fingerprints that cannot be derived from black-box outputs, and multi-factor validation that survives model post-processing. Parameter-encoding schemes that bind selected weights to a hashed secret, such as NeuralMark, represent one step in this direction, but the broader lesson is procedural rather than technical. Any verification protocol that accepts a single trigger-response score as proof of ownership is, in principle, exposed to fabricated evidence.</p>
<p>The FakeMark code and experimental data have been released publicly, and the authors hope the framework will serve as a stress test for the next generation of watermarking defenses rather than as a tool for intellectual property fraud. The study&#8217;s most enduring contribution may be its reframing of the ownership question. Watermarking research has largely focused on making triggers robust to removal, extraction, and fine-tuning; FakeMark shifts attention to the provenance of the evidence itself. A model that answers the secret questions correctly is not necessarily a stolen model, and a claimant who produces the right answers is not necessarily the rightful owner. Until verification protocols can distinguish between these possibilities, the researchers conclude, the evidentiary value of any behavioral watermark remains fundamentally limited, and the growing market for machine learning intellectual property will need stronger, multi-factor guarantees to stay ahead of those who would forge them.</p>
<p><strong>Subject of Research:</strong> Gradient-guided forgery of behavioral model watermarks in deep neural networks</p>
<p><strong>Article Title:</strong> FakeMark: gradient-guided false watermark claims via robust feature fusion</p>
<p><strong>Article References:</strong> Wu, Y., Li, W., Nie, H., Zhou, Z., Li, J., Gong, Y., &amp; Lu, S. (2026). FakeMark: gradient-guided false watermark claims via robust feature fusion. <em>Cybersecurity, 9</em>(1), Article 217. <a href="https://doi.org/10.1186/s42400-026-00654-8" rel="noopener noreferrer">https://doi.org/10.1186/s42400-026-00654-8</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1186/s42400-026-00654-8" rel="noopener noreferrer">10.1186/s42400-026-00654-8</a></p>
<p><strong>Keywords:</strong> model watermarking, false ownership claims, adversarial transferability, deep neural network security, intellectual property protection, FakeMark, surrogate models, feature fusion, CIFAR-10, ImageNet, Cybersecurity, machine learning</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">205395</post-id>	</item>
		<item>
		<title>AI-Enhanced Adaptive Virtual Screening Accelerates Ligand Discovery in Huge Compound Libraries</title>
		<link>https://scienmag.com/ai-enhanced-adaptive-virtual-screening-accelerates-ligand-discovery-in-huge-compound-libraries/</link>
		
		<dc:creator><![CDATA[Blake Davidson]]></dc:creator>
		<pubDate>Sun, 20 Sep 2026 23:50:11 +0000</pubDate>
				<category><![CDATA[Medicine]]></category>
		<category><![CDATA[accelerated ligand discovery with artificial intelligence]]></category>
		<category><![CDATA[active learning]]></category>
		<category><![CDATA[active learning in virtual screening]]></category>
		<category><![CDATA[adaptive computational drug discovery]]></category>
		<category><![CDATA[AI-driven virtual screening]]></category>
		<category><![CDATA[Artificial Intelligence]]></category>
		<category><![CDATA[drug discovery]]></category>
		<category><![CDATA[efficient drug candidate prioritization]]></category>
		<category><![CDATA[high-throughput virtual screening techniques]]></category>
		<category><![CDATA[hit identification]]></category>
		<category><![CDATA[large chemical libraries]]></category>
		<category><![CDATA[large-scale chemical library screening]]></category>
		<category><![CDATA[ligand design]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[machine learning surrogate models for molecular docking]]></category>
		<category><![CDATA[molecular docking]]></category>
		<category><![CDATA[Nature Biotechnology]]></category>
		<category><![CDATA[nature biotechnology advancements in drug discovery]]></category>
		<category><![CDATA[physics-based vs. AI-based scoring methods]]></category>
		<category><![CDATA[real-time machine learning in ligand identification]]></category>
		<category><![CDATA[scalable computational methods for billions of molecules]]></category>
		<category><![CDATA[structure-based drug design]]></category>
		<category><![CDATA[surrogate models]]></category>
		<category><![CDATA[virtual screening]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=204120</guid>

					<description><![CDATA[An adaptive artificial intelligence framework slashes the computational cost of screening billion-molecule chemical libraries while experimentally confirming high-quality ligand hits across multiple drug targets.]]></description>
										<content:encoded><![CDATA[<p>The search for new drug candidates has always been a numbers game, and the numbers have grown staggering. Commercial and open-source chemical catalogs now contain billions of purchasable molecules, a collection so vast that no laboratory could ever test it directly. Virtual screening, the computational triage of these libraries against disease-relevant protein targets, promised to tame the flood, yet in practice it has been limited by a stubborn trade-off: the most accurate physics-based scoring methods are far too slow to apply to billions of structures, while the fastest methods are too crude to prioritize the right molecules. A new approach published in Nature Biotechnology tackles this dilemma head-on by making the screening process itself adaptive, allowing artificial intelligence to learn, in real time, which regions of a massive chemical library deserve the expensive computational attention and which can be safely set aside.</p>
<p>The core idea behind the method is an active learning loop that governs how computational resources are spent across a screening campaign. Rather than scoring every molecule in a library with the same level of rigor, the algorithm samples molecules at random, sends them through a computationally demanding reference pipeline, and trains a surrogate machine learning model on the results. That model then estimates the likelihood that each untested molecule would rank highly under the full reference protocol. Only the compounds the model is most uncertain about, or most confident will rank near the top, are promoted for exhaustive evaluation. Everything else is filtered out cheaply. The loop repeats, with each round refining the model and progressively concentrating the expensive calculations on the small fraction of the library where the true hits are most likely to reside.</p>
<p>What distinguishes the new work is the sophistication of the reference scoring layer and the way the adaptive policy is engineered around it. The reference pipeline pairs structure-based docking with physics-inspired rescoring and, for a select tier of candidates, more rigorous free-energy estimates, so that the labels the surrogate model learns from are far more reliable than simple docking scores alone. The authors describe an ensemble-based selection strategy in which multiple independently trained models must agree before a molecule is discarded, reducing the risk that a single overconfident network silently eliminates a genuine ligand. Early-stopping criteria and calibration checks are built into the workflow, ensuring that the screening continues until the expected yield of new top-ranked molecules falls below a defined threshold, rather than running for an arbitrary number of cycles.</p>
<p>Benchmarking against billion-scale libraries was central to the study. The team screened catalogs of synthetically accessible compounds against multiple protein targets with well-characterized ligands, including kinase and G-protein-coupled receptor systems, and compared the adaptive workflow against conventional exhaustive docking and against simpler one-shot machine learning filters. The results showed that the adaptive approach recovered the overwhelming majority of the true top-ranking hits identified by exhaustive screening while requiring only a small percentage of the full computational cost. In practical terms, campaigns that would have demanded months of processor time on conventional infrastructure were compressed into days, without sacrificing the quality of the enriched hit lists. The savings scaled with library size, which is precisely the regime where modern screening campaigns need help most.</p>
<p>The authors also stress an important conceptual shift that the method embodies: virtual screening treated not as a static ranking problem but as a sequential decision problem. Each docking calculation performed generates information, and a well-designed screening campaign should choose the next calculation to maximize what is learned about the library as a whole. This framing, borrowed from Bayesian optimization and active learning in other scientific domains, explains why the method improves so dramatically over naive random subsampling. Random subsampling wastes effort on molecules that are obviously unpromising; the adaptive policy instead balances exploration of chemically novel regions against exploitation of regions the models already associate with strong predicted binding, maintaining diversity in the candidate pool and guarding against the collapse onto narrow, easily predicted chemical motifs.</p>
<p>Experimental validation was a critical component of the work, elevating the study beyond a purely computational exercise. Compounds prioritized by the adaptive pipeline were synthesized and tested in biochemical and biophysical assays for several targets, and the experimentally confirmed hit rates substantially exceeded what is typically reported for high-throughput screening of comparably sized collections. Several confirmed ligands occupied binding poses consistent with the computational predictions, and structure-guided analoguing of a subset produced measurable potency improvements. The authors report ligandable chemotypes that had not previously been associated with the targets in question, suggesting that the adaptive workflow does more than rediscover known pharmacophores; it surfaces genuinely novel starting points that conventional screens tend to miss.</p>
<p>For medicinal chemistry teams, the implications are considerable. Ultra-large library screening has already demonstrated, in previous landmark studies, that docked catalogs containing hundreds of millions to billions of molecules can yield potent inhibitors with favorable ligand efficiency. But the computational price of those successes has limited their routine use, concentrating the technique in a handful of well-resourced laboratories. By cutting the compute budget by an order of magnitude or more while preserving hit quality, the adaptive method effectively democratizes the practice. Medium-sized academic groups and biotech companies with modest clusters can now contemplate libraries at scales that were previously out of reach, and industrial groups can run many more target campaigns in parallel, widening the front of early-stage discovery.</p>
<p>The method is not without limitations, and the authors are candid about them. The surrogate models inherit biases from the reference pipeline, so systematic errors in docking or rescoring propagate into the learned filter, albeit in diluted form. Libraries dominated by chemical series very different from the training sample can degrade model confidence, requiring longer exploration phases. There are also open questions about how best to adapt the approach when multiple objectives, such as predicted potency, synthetic accessibility, and off-target liability, must be balanced simultaneously. The team suggests that extending the active learning framework to multi-property optimization is a natural next step, as is coupling the workflow with generative models that can propose and evaluate hypothetical compounds beyond the boundaries of any fixed catalog.</p>
<p>Broader significance aside, the study lands at a moment when the pharmaceutical industry is re-evaluating how artificial intelligence should be woven into discovery pipelines. The lesson of this work is a sober and practical one: the value of machine learning in early drug discovery does not come from replacing physics-based evaluation with end-to-end prediction, but from orchestrating where expensive, trustworthy calculations are spent. The adaptive screening framework treats the AI model as an intelligent allocation layer over rigorous science, and the empirical results suggest that this division of labor is the right one. As chemical catalogs continue to expand and as cryo-electron microscopy and AlphaFold-style structure prediction make more protein targets accessible to structure-based design, methods that make billion-scale screening affordable will likely become standard infrastructure. If the reported hit rates and compute savings hold across a wider range of targets, this approach could mark one of those quiet inflection points that reshapes how the field searches for its next generation of medicines.</p>
<p><strong>Subject of Research:</strong> AI-enhanced adaptive virtual screening of large compound libraries for ligand discovery</p>
<p><strong>Article Title:</strong> AI-enhanced adaptive virtual screening of large libraries for ligand discovery</p>
<p><strong>Article References:</strong> Cecchini, D., Nigam, A., Tang, M., Reis, J., Koop, M., Gottinger, A., Nicoll, C. R., Wang, Y., Jayaraj, A., Çınaroglu˘, S. S., Törner, R., Malets, Y., Gehev, M., Padmanabha Das, K. M., Churion, K., Kim, J., Thomas, N., Li, Y., Seo, H.-S., &#8230; Gorgulla, C. (2026). AI-enhanced adaptive virtual screening of large libraries for ligand discovery. <em>Nature Biotechnology</em>. <a href="https://doi.org/10.1038/s41587-026-03217-x" rel="noopener noreferrer">https://doi.org/10.1038/s41587-026-03217-x</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1038/s41587-026-03217-x" rel="noopener noreferrer">10.1038/s41587-026-03217-x</a></p>
<p><strong>Keywords:</strong> virtual screening, artificial intelligence, drug discovery, active learning, molecular docking, ligand design, machine learning, large chemical libraries, hit identification, structure-based drug design, surrogate models, Nature Biotechnology</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">204120</post-id>	</item>
	</channel>
</rss>
