<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>symbolic regression &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/symbolic-regression/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Sat, 12 Sep 2026 15:32:40 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>symbolic regression &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>AI Learns to Draw Decision Boundaries as Readable Equations</title>
		<link>https://scienmag.com/ai-learns-to-draw-decision-boundaries-as-readable-equations/</link>
		
		<dc:creator><![CDATA[Denise Maddox]]></dc:creator>
		<pubDate>Sat, 12 Sep 2026 15:32:40 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[AI transparency and explainability]]></category>
		<category><![CDATA[automated equation discovery for data classification]]></category>
		<category><![CDATA[beam search]]></category>
		<category><![CDATA[binary classification]]></category>
		<category><![CDATA[classification models with analytical formulas]]></category>
		<category><![CDATA[context-free grammar]]></category>
		<category><![CDATA[decision boundary]]></category>
		<category><![CDATA[decision boundary equations in AI]]></category>
		<category><![CDATA[decision tree simplification]]></category>
		<category><![CDATA[Equation]]></category>
		<category><![CDATA[equation discovery]]></category>
		<category><![CDATA[equation discovery in machine learning]]></category>
		<category><![CDATA[explainable AI]]></category>
		<category><![CDATA[human-readable machine learning models]]></category>
		<category><![CDATA[interpretable machine learning]]></category>
		<category><![CDATA[interpretable neural networks]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[machine learning model interpretability]]></category>
		<category><![CDATA[readable mathematical models in AI]]></category>
		<category><![CDATA[symbolic classification]]></category>
		<category><![CDATA[symbolic regression]]></category>
		<category><![CDATA[symbolic regression for classification]]></category>
		<category><![CDATA[transparent AI decision-making]]></category>
		<category><![CDATA[UCI datasets]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=195947</guid>

					<description><![CDATA[Researchers at Leiden University have developed EDC, a framework that discovers single readable equations defining decision boundaries, rivaling black-box classifiers while remaining fully interpretable.]]></description>
										<content:encoded><![CDATA[<p>Machine learning models are famous for their uncanny accuracy and infamous for their opacity. A random forest or a neural network can sift through thousands of patient records, financial transactions, or sensor readings and deliver a verdict in milliseconds, but when practitioners ask why the model reached its conclusion, the answer is usually buried in millions of weighted connections or hundreds of tangled decision trees. A new study published in the journal Machine Learning challenges this trade-off between power and transparency, introducing a framework that discovers a single, human-readable equation whose sign alone tells you which class a data point belongs to.</p>
<p>The method, called Equation Discovery for Classification, or EDC, was developed by Guus Toussaint and Arno Knobbe of Leiden University. It extends a research tradition that has long flourished in regression: symbolic regression, the automated search for analytical formulas that fit numerical data. Theauthors of the study point out that equation discovery has historically been applied almost exclusively to problems where the target is a continuous number, such as recovering physical laws like the relationship between the pressure, volume, and temperature of a gas. Their contribution is to redirect that machinery toward binary classification, where the goal is to separate two classes cleanly and explainably.</p>
<p>The core idea is elegantly simple in concept. Rather than learning an opaque scoring function, EDC searches for a concise mathematical expression f(x) and a threshold theta, such that a data point is assigned to the positive class whenever f(x) meets or exceeds theta. For a linear equation, this is essentially the geometry underlying logistic regression or a linear support vector machine: a hyperplane slicing the feature space into two halves. What sets EDC apart is that the search is not confined to straight lines. By allowing nonlinear building blocks such as products of features and exponential terms, the algorithm can trace curved, interaction-driven boundaries while still producing an expression short enough for a domain expert to read, critique, and even correct by hand.</p>
<p>Technically, the framework rests on two pillars: a structured search and a dedicated numerical optimizer. The search space of candidate equations is defined by a configurable context-free grammar, a design choice inherited from classic work on declarative bias in equation discovery. The grammar constrains equations to sums of simple summands, including linear terms, products of two features, and exponential expressions, each parameterized by constants. Crucially, the grammar is redundancy-aware: constructions that would produce syntactically different but semantically identical equations are pruned in advance, since, for example, subtraction between summands is unnecessary when a constant coefficient can simply be negative. Traversing this space exhaustively is impossible for all but the smallest problems, so the algorithm employs beam search, iteratively refining the most promising candidate equations level by level while keeping only a fixed number of survivors at each stage.</p>
<p>Fitting the constants inside each candidate equation proved surprisingly subtle. Although every equation in the grammar is differentiable, gradient descent performed poorly in practice, largely because of the exponential terms that pepper the search space. The authors instead adopted a tailored hill-climbing procedure that first samples a large pool of random constant configurations, then concentrates its remaining budget on refining the best ones. In a systematic comparison against off-the-shelf optimizers from the SciPy library, including Powell, Cobyqa, Cobyla, Nelder-Mead, and stochastic gradient descent, this hill climber achieved the highest mean area under the ROC curve across one hundred randomly generated test problems. Perhaps most strikingly, a simple random-sampling baseline already reached a mean AUC of 0.9974 with a budget of one thousand evaluations, revealing that the inner optimization problem is more tractable than one might fear.</p>
<p>The experiments on synthetic data produced one of the study&#8217;s most intriguing findings. When Gaussian noise was injected into datasets whose generating decision boundaries were known, EDC did not merely match the original boundary; beyond a certain noise level it outperformed it. The explanation is a phenomenon the authors describe carefully: noise pushes data points near a curved boundary across it, and concave sections of that boundary collect more stray points than they lose, causing the effective boundary embedded in the data to drift away from its original position and gradually straighten. EDC, fitting the data rather than the hidden formula, tracks this shifted, smoothed boundary, achieving higher AUC than the very equation that generated the data. As noise grew from negligible to substantial across seventeen hundred artificial datasets, even a restricted linear version of EDC eventually beat the original nonlinear boundary.</p>
<p>The framework also proved capable of reconstructing notoriously difficult structures, including XOR-like and interaction-driven boundaries, and of fitting data produced by Gaussian clusters where no explicit target equation exists at all. On these cluster-based problems, which the authors describe as closest to real-world conditions among their artificial experiments, EDC outperformed existing symbolic classification approaches while landing near state-of-the-art black-box methods such as random forests, multi-layer perceptrons, and radial-basis-function support vector machines, all of which posted mean AUC values around 0.97.</p>
<p>Real-world benchmarks reinforced the pattern. Across nine binary classification datasets from the UCI repository, spanning tasks from banknote authentication to income prediction, EDC achieved a higher AUC than every competing equation-discovery-based classifier on every dataset, and beat simple decision trees across the board. A critical distance analysis showed that random forests, neural networks, and SVMs did not statistically significantly outperform EDC, even though those black-box methods won on several datasets, particularly ionosphere and sonar, where feature-class relationships appear to lie outside EDC&#8217;s current set of building blocks. On the Adult income dataset, the discovered equation offered a vivid demonstration of interpretability in action: one term acted as a penalty on years of education only when the individual appeared as a child in the household, while an exponential term over marital status effectively functioned as an if-else statement, adding roughly 28,586 to the score for married individuals living with a spouse and a negligible 8 otherwise.</p>
<p>The method&#8217;s main weakness is computational cost. In its default configuration, with a search depth of six and a beam width of ten, EDC takes dramatically longer than the sub-second runtimes of conventional classifiers, largely because pairwise interaction terms grow quadratically with the number of features and because one-hot encoding of categorical variables inflates the feature count. The authors show, however, that the expense is largely optional. Restricting the search depth, narrowing the beam, or replacing interaction terms with simple quadratic terms produced speed-ups of up to a factor of fifty at only a marginal loss of accuracy, with no statistically significant difference in performance across the benchmark suite. The authors also note that the framework is not limited to two classes: standard one-versus-rest schemes can extend EDC to multi-class problems, with each of the resulting equations explaining when a particular class label prevails.</p>
<p>The work arrives amid a broader movement toward explainable machine learning, driven by domains such as medicine, finance, and criminal justice where decisions carry real consequences and regulators increasingly demand justification. EDC offers a principled bridge between symbolic regression and classification, delivering models that are simultaneously compact, expressive, and auditable. While it may not dethrone random forests or deep networks on raw accuracy, it demonstrates that the gap is small enough, and the interpretability dividend large enough, that equations once again deserve a seat at the machine learning table. The authors have released their code, experimental setup, and datasets openly so that the entire study can be reproduced and the grammar extended by practitioners in new domains.</p>
<p><strong>Subject of Research:</strong> Equation discovery for binary classification using interpretable symbolic decision boundaries</p>
<p><strong>Article Title:</strong> Equation Discovery for Classification: Finding Interpretable Symbolic Specifications of the Decision Boundary</p>
<p><strong>Article References:</strong> Toussaint, G., &amp; Knobbe, A. (2026). Equation Discovery for Classification: Finding Interpretable Symbolic Specifications of the Decision Boundary. <em>Machine Learning, 115</em>(9), Article 215. <a href="https://doi.org/10.1007/s10994-026-07156-1" rel="noopener noreferrer">https://doi.org/10.1007/s10994-026-07156-1</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s10994-026-07156-1" rel="noopener noreferrer">10.1007/s10994-026-07156-1</a></p>
<p><strong>Keywords:</strong> equation discovery, binary classification, symbolic regression, decision boundary, interpretable machine learning, beam search, symbolic classification, explainable AI, context-free grammar, UCI datasets, machine learning, Equation</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">195947</post-id>	</item>
		<item>
		<title>AI Writes Its Own Physics Equations for Materials from Raw Data</title>
		<link>https://scienmag.com/ai-writes-its-own-physics-equations-for-materials-from-raw-data/</link>
		
		<dc:creator><![CDATA[Katie Riggs]]></dc:creator>
		<pubDate>Sat, 12 Sep 2026 13:53:21 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[advanced modeling of filled rubbers]]></category>
		<category><![CDATA[AI in predicting failure and hardening of materials]]></category>
		<category><![CDATA[AI-based material property prediction]]></category>
		<category><![CDATA[AI-driven materials modeling]]></category>
		<category><![CDATA[alloy steels]]></category>
		<category><![CDATA[application of AI in alloy steels and lithium metal]]></category>
		<category><![CDATA[autonomous discovery of constitutive laws]]></category>
		<category><![CDATA[constitutive models]]></category>
		<category><![CDATA[data-driven equations for material deformation]]></category>
		<category><![CDATA[experimental data for material behavior]]></category>
		<category><![CDATA[filled rubbers]]></category>
		<category><![CDATA[graph-based equation discovery]]></category>
		<category><![CDATA[hyperelasticity]]></category>
		<category><![CDATA[interpretability of AI-generated physics equations]]></category>
		<category><![CDATA[interpretable AI]]></category>
		<category><![CDATA[lithium metal]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[machine learning in solid mechanics]]></category>
		<category><![CDATA[overcoming limitations of traditional empirical models]]></category>
		<category><![CDATA[Science Advances]]></category>
		<category><![CDATA[scientific advancements in materials science using artificial intelligence]]></category>
		<category><![CDATA[solid mechanics]]></category>
		<category><![CDATA[strain hardening]]></category>
		<category><![CDATA[symbolic regression]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=194815</guid>

					<description><![CDATA[Researchers at the Eastern Institute of Technology, Ningbo, developed a graph-based AI framework called GraphED that autonomously discovers interpretable constitutive equations for materials directly from experimental data.]]></description>
										<content:encoded><![CDATA[<p>For more than a century, the way engineers describe how materials deform, harden, and fail has followed a familiar ritual. A researcher with deep physical intuition proposes a mathematical formula, guesses its general shape, and then spends weeks or months calibrating its adjustable constants against experiments. The formula works, more or less, until a new material, a new temperature range, or a new loading condition pushes it beyond the assumptions baked into its structure. Now a team at the Eastern Institute of Technology, Ningbo, has built an artificial intelligence framework that breaks this ritual apart. Instead of beginning with a human-chosen equation, the method starts purely from experimental data and autonomously searches for the optimal constitutive law itself, producing compact, explicit, human-readable equations that outperform the empirical models engineers have relied on for decades. The work, published in Science Advances, demonstrates the approach on alloy steels, lithium metal, and filled rubbers, three material classes whose mechanical behavior has long resisted clean mathematical description.</p>
<p>Constitutive models are the connective tissue of solid mechanics. They are the mathematical rules that translate stress, strain, temperature, and strain rate into predictions of how a real object will bend, dent, stretch, or shatter. Every crash simulation of a car body, every structural analysis of a bridge, and every design study for a battery electrode depends on them. Hao Xu, a postdoctoral researcher at EIT and lead author of the study, explains that the traditional paradigm, while enormously successful, carries an inherent ceiling. Researchers derive a mathematical form based on physical intuition and then calibrate its parameters with data, but the predetermined equation structure restricts what the model can ultimately describe and predict. If the true material behavior involves a functional dependence the chosen formula cannot express, no amount of parameter fitting will recover it. The new framework, called GraphED, inverts the workflow: data comes first, and the equation emerges from it.</p>
<p>The central technical obstacle to such data-driven discovery is representation. Machine learning systems can only optimize what they can encode, and mathematical equations are notoriously awkward objects for algorithms to manipulate. Yuntian Chen, EIT associate professor and co-corresponding author of the study, describes the challenge as efficiently encoding both equation architectures and material-specific parameters into a searchable format for computational algorithms. Conventional symbolic regression approaches typically represent expressions as trees, with operators branching from a root node down to leaves. Trees, however, make it difficult to share structure across different materials and to embed tunable parameters cleanly. The EIT team&#8217;s solution is to abandon the tree entirely and encode each candidate equation as a directed graph. In this graph architecture, nodes correspond to mathematical operators and physical variables, while directed edges define the logical and computational connections between them.</p>
<p>The graph representation is more than a cosmetic change, because the edges themselves become carriers of physical meaning. An edge can hold a fixed physical constant, such as a universal exponent, or a tunable material-dependent parameter that varies from one alloy or polymer to another. This design allows GraphED to do two things simultaneously that earlier methods treated separately. It can identify a universal mathematical structure shared across diverse materials and experimental conditions, and it can adaptively calibrate personalized parameters for each individual material scenario. The framework then iterates: it generates candidate graph-structured equations, evaluates how well each one reproduces the experimental data, and optimizes both the topology of the graph and the values of its parameters. The output is a constitutive law that is physically consistent, mathematically compact, and, crucially, interpretable by a human engineer who can read the equation and understand what each term contributes.</p>
<p>That last quality, interpretability, is what separates this work from the black-box models that dominate much of modern machine learning. A neural network trained on material data may predict well, but it offers no equation, no insight, and no transferable knowledge. GraphED delivers the opposite: an explicit formula that a researcher can inspect, question, embed in existing simulation software, and carry into new contexts. The team validated this promise across three very different solid material systems, each chosen for its distinct mechanical character and its historical resistance to empirical modeling. The breadth of the test cases matters, because a discovery method that only works on one material family is a curiosity; one that generalizes across steels, metals, and elastomers is a tool.</p>
<p>The first test case was alloy steel, the workhorse of structural engineering, whose behavior under high strain rates and large plastic deformation is governed by strain-rate dependence and strain hardening. These are precisely the phenomena captured by the Johnson–Cook model, one of the most widely adopted constitutive formulations in impact and crashworthiness analysis. GraphED discovered explicit equations governing both strain-rate dependence and strain hardening directly from experimental data, and when the team integrated these data-driven equations into a complete constitutive model, the result delivered more precise mechanical predictions than the Johnson–Cook benchmark. For a field in which the incumbent model has been refined over decades, the demonstration that an autonomously discovered equation can beat it on predictive accuracy is a significant result.</p>
<p>The second material pushed the method into far harder territory. Lithium metal is a critical component of next-generation energy devices, including high-energy-density batteries, but its mechanical behavior is notoriously difficult to model. Its plastic flow is highly sensitive to both temperature and strain rate, and the interplay of these dependencies has made traditional empirical formulations unreliable. GraphED yielded concise plastic-flow constitutive equations that achieve superior agreement with experimental measurements compared with conventional empirical models. For battery designers, whose simulations of electrode deformation and dendrite-related failure depend on accurate mechanical descriptions, the ability to extract a trustworthy equation directly from data could shorten development cycles and improve the fidelity of electrochemical-mechanical coupled models.</p>
<p>The third validation targeted filled rubbers, a class of elastomers whose hyperelastic response changes with filler composition and temperature. Rubber components in tires, seals, and vibration isolators experience large, reversible deformations, and their stiffness depends on both what they are made of and how hot they are. GraphED captured a compact hyperelastic constitutive equation that maintains robust accuracy across varying material compositions and temperature conditions, meaning a single discovered formula could serve a family of rubber compounds rather than requiring a bespoke model for each formulation. Together, the three case studies span rate-dependent plasticity, temperature-sensitive metal plasticity, and finite-strain hyperelasticity, covering a remarkably wide slice of solid mechanics with one framework.</p>
<p>The researchers see the implications reaching well beyond computational mechanics. Dongxiao Zhang, Chair Professor at EIT and corresponding author of the study, notes that many complex material mechanical behaviors cannot be well described by existing empirical constitutive models, and that equation discovery via graph-based AI provides a powerful new paradigm to assist researchers in deriving rigorous mathematical descriptions when traditional model forms fall short. The phrase that matters is rigorous mathematical description: the goal is not a surrogate that mimics data, but a law, expressed in symbols, that captures the underlying behavior. Because the framework requires only experimental or observational data as input, its potential applications extend to any discipline seeking interpretable physical laws, from soft matter physics and biomechanics to geophysics and beyond, wherever data exists but the governing equation does not.</p>
<p>The study also signals a broader shift in how science may handle the equation-discovery step itself. Symbolic regression and equation learning have been growing fields, but the graph-based encoding developed here resolves a key bottleneck by enabling simultaneous optimization of equation topology and material parameterization, a combination that makes the search tractable for realistic, multi-material datasets. If the paradigm spreads, the role of the mechanician may evolve from proposing formulas to curating data, specifying physical constraints, and interpreting the equations that the algorithm returns. The authors declare no competing interest, and the research was published as a peer-reviewed contribution in Science Advances. For a field built on equations handed down through generations of researchers, the prospect of machines that read raw experimental data and hand back compact, readable laws of material behavior is a genuinely new way of doing physics, and the steels, lithium, and rubber that proved it are only the beginning.</p>
<p><strong>Subject of Research:</strong> Graph-based AI equation discovery of constitutive laws in solid materials from experimental data</p>
<p><strong>Article Title:</strong> AI discovers interpretable constitutive laws in solids directly from data</p>
<p><strong>Article References:</strong> AI discovers interpretable constitutive laws in solids directly from data. (n.d.). <a href="https://www.eurekalert.org/news-releases/1143179" rel="noopener noreferrer">Original publication</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> Not provided</p>
<p><strong>Keywords:</strong> constitutive models, solid mechanics, graph-based equation discovery, symbolic regression, alloy steels, lithium metal, filled rubbers, machine learning, Science Advances, interpretable AI, strain hardening, hyperelasticity</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">194815</post-id>	</item>
	</channel>
</rss>
