<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>regularization &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/regularization/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Sat, 12 Sep 2026 21:48:50 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>regularization &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>AI Learns From Its Own Past to Sharpen Graph Neural Networks</title>
		<link>https://scienmag.com/ai-learns-from-its-own-past-to-sharpen-graph-neural-networks/</link>
		
		<dc:creator><![CDATA[Cassandra Pierce]]></dc:creator>
		<pubDate>Sat, 12 Sep 2026 21:48:50 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[data mining and knowledge discovery in graph models]]></category>
		<category><![CDATA[deep learning]]></category>
		<category><![CDATA[generalization challenges in GNNs]]></category>
		<category><![CDATA[graph data applications in finance and science]]></category>
		<category><![CDATA[graph neural network training techniques]]></category>
		<category><![CDATA[Graph Neural Networks]]></category>
		<category><![CDATA[improving GNN robustness]]></category>
		<category><![CDATA[innovative methods in graph machine learning]]></category>
		<category><![CDATA[knowledge distillation]]></category>
		<category><![CDATA[lessons]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[memory augmentation]]></category>
		<category><![CDATA[memory-augmented self-distillation]]></category>
		<category><![CDATA[neural network history learning]]></category>
		<category><![CDATA[node classification]]></category>
		<category><![CDATA[overfitting]]></category>
		<category><![CDATA[overfitting in graph models]]></category>
		<category><![CDATA[oversmoothing]]></category>
		<category><![CDATA[regularization]]></category>
		<category><![CDATA[relationship-based data analysis]]></category>
		<category><![CDATA[self-distillation]]></category>
		<category><![CDATA[self-learning in machine learning]]></category>
		<category><![CDATA[Taking]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=198856</guid>

					<description><![CDATA[Researchers have developed a memory-augmented self-distillation framework that boosts graph neural network accuracy by 2.5 to 6 percent across benchmark datasets.]]></description>
										<content:encoded><![CDATA[<p>Graph neural networks have become one of the most powerful tools in modern machine learning for making sense of data that lives in relationships rather than rows of a spreadsheet. From detecting money laundering in Bitcoin transaction networks to classifying scientific papers by their citation patterns, these models excel at learning from graphs, the mathematical structures that capture how entities connect. Yet despite their success, graph neural networks carry a persistent weakness that has frustrated researchers for years: they overfit. A model that performs brilliantly on the data it was trained on can falter badly when confronted with nodes it has never seen, undermining the very generalization that makes graph learning valuable in the real world.</p>
<p>A new study published in Data Mining and Knowledge Discovery by Saurabh Sharma and Joydeep Chandra of the Indian Institute of Technology Patna, together with Souvik Chowdhury of Jadavpur University, proposes an elegant way out of this trap. Their approach, called Memory Augmented Self Distillation, teaches a graph neural network to learn from its own history. Rather than relying on a separate, fully trained teacher model to guide a smaller student, the framework builds a memory of the network&#8217;s past states and draws on that stored knowledge to regularize and refine its current training. The result is a measurable improvement of 2.5 to 6 percent in accuracy across a range of benchmark datasets compared with existing graph neural network training and self-distillation methods.</p>
<p>To understand why this matters, it helps to look at the technique the new method builds upon: knowledge distillation. First popularized by Geoffrey Hinton and colleagues in 2015, knowledge distillation is a compression and regularization strategy in which a large, well-trained teacher network transfers its knowledge to a smaller student network. The student learns not only from the ground-truth labels but also from the teacher&#8217;s softer, richer output distributions, which encode subtle information about how confident the teacher is and which classes resemble one another. In domains like computer vision and natural language processing, distillation has become a standard tool for building compact, robust models.</p>
<p>Applying distillation to graph neural networks, however, has proven surprisingly difficult. The core obstacle is a phenomenon known as oversmoothing. Graph neural networks learn by passing messages along edges, allowing each node to aggregate information from its neighbors. When the network is deep, this repeated aggregation causes the representations of all nodes to converge toward indistinguishable similarity, effectively washing out the distinctive features that make classification possible. A teacher graph neural network that has been trained to convergence may therefore produce representations that are smooth but information-poor, offering the student little of value to learn from. Conventional teacher-student distillation, so effective elsewhere, struggles to deliver meaningful guidance in the graph setting.</p>
<p>Self-distillation, in which a network serves as its own teacher, sidesteps the need for a separate teacher but runs into a different problem: the information bottleneck. Because the student and the teacher share the same architecture and training data, the knowledge available to transfer is limited by what the model already contains. Without an external source of diverse information, self-distillation can become an echo chamber, reinforcing the model&#8217;s existing biases rather than correcting them. Previous efforts to make graph distillation work have explored multi-teacher setups, adversarial distillation, and structure-aware multilayer perceptrons, but each carries its own computational or methodological trade-offs.</p>
<p>The framework introduced by Sharma, Chowdhury, and Chandra takes a fundamentally different route. Instead of a single teacher or a fixed set of teachers, the method constructs a memory bank that captures diverse snapshots of the learning process as it unfolds. These memory entries act as multiple, heterogeneous knowledge sources drawn from the student model&#8217;s own trajectory through training. Because each snapshot reflects a different stage of learning, the memory collectively encodes a richer and more varied body of knowledge than any single model state could offer, directly addressing the information bottleneck that plagues conventional self-distillation.</p>
<p>The crucial question then becomes which of these stored sources the network should listen to at any given moment. Listening to a poorly trained early snapshot could mislead the model, while relying exclusively on the most recent state would recreate the echo chamber problem. The researchers solve this with a competency-based knowledge source selection mechanism. This mechanism dynamically evaluates how competent each memory source is relative to the current learning objective and selects the most pertinent one for distillation at each step. In effect, the network continuously asks which of its past selves has the most useful lesson to teach, and adapts its supervision accordingly. This adaptive selection transforms the memory from a static archive into an active, evolving curriculum.</p>
<p>The technical payoff of this design is twofold. First, the distillation signal from competent memory sources acts as a powerful regularizer, discouraging the network from drifting into the overconfident, overfit solutions that plague graph learning on limited labeled data. Second, because the memory sources are diverse, the student is exposed to a broader distribution of knowledge than it could generate on its own, improving the quality of its learned representations. The authors demonstrate these gains across multiple benchmark datasets, including widely used citation networks such as Cora, Citeseer, and Pubmed, as well as graph kernel benchmarks and an elliptic Bitcoin transaction dataset used for anti-money laundering research, where only a subset of classes was analyzed for the experiments.</p>
<p>The evaluation methodology reflects careful statistical practice, with paired t-tests used to establish the significance of the improvements over baseline methods. The comparisons span the landscape of graph neural network architectures, including graph convolutional networks, graph attention networks, and jumping knowledge networks, alongside recent distillation frameworks designed specifically for graphs. The consistency of the accuracy gains across datasets and architectures suggests that the benefit stems from the underlying principle of memory-augmented self-supervision rather than from tuning to any particular benchmark. The final student models also exhibited better generalization, retaining their performance advantages when evaluated beyond the training distribution.</p>
<p>Beyond the immediate results, the study points toward a broader shift in how researchers think about the training of graph-based models. The idea that a model&#8217;s own training history is a resource worth preserving and mining is a departure from the standard paradigm in which intermediate states are discarded the moment a new set of weights is computed. It resonates with a simple intuition: lessons from the past, properly curated, can guide better decisions in the present. For graph neural networks, whose vulnerability to overfitting and oversmoothing has limited the depth and reliability of the models practitioners can deploy, that intuition now has concrete, quantified support. As graph learning continues to expand into finance, chemistry, recommendation systems, and network security, techniques like Memory Augmented Self Distillation could become a standard component of the training pipeline, helping models not only to learn from data but to learn from themselves.</p>
<p><strong>Subject of Research:</strong> A memory-augmented self-distillation framework for improving the generalization of graph neural networks</p>
<p><strong>Article Title:</strong> Taking lessons from history: Memory Augmented Self Distillation for graph neural networks</p>
<p><strong>Article References:</strong> Taking lessons from history: Memory Augmented Self Distillation for graph neural networks. (n.d.). <a href="https://doi.org/10.1007/s10618-026-01258-z" rel="noopener noreferrer">https://doi.org/10.1007/s10618-026-01258-z</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s10618-026-01258-z" rel="noopener noreferrer">10.1007/s10618-026-01258-z</a></p>
<p><strong>Keywords:</strong> graph neural networks, knowledge distillation, self-distillation, overfitting, oversmoothing, memory augmentation, node classification, regularization, deep learning, machine learning, Taking, lessons</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">198856</post-id>	</item>
		<item>
		<title>Projective Coordinates Turn Kepler&#8217;s Classic Orbits Into Simple Harmonic Motion</title>
		<link>https://scienmag.com/projective-coordinates-turn-keplers-classic-orbits-into-simple-harmonic-motion/</link>
		
		<dc:creator><![CDATA[Grant Pearson]]></dc:creator>
		<pubDate>Sat, 12 Sep 2026 19:27:05 +0000</pubDate>
				<category><![CDATA[Space]]></category>
		<category><![CDATA[canonical transformations]]></category>
		<category><![CDATA[canonical transformations in dynamical systems]]></category>
		<category><![CDATA[celestial mechanics]]></category>
		<category><![CDATA[celestial mechanics numerical integration]]></category>
		<category><![CDATA[Hamiltonian dynamics]]></category>
		<category><![CDATA[Hamiltonian mechanics in celestial mechanics]]></category>
		<category><![CDATA[harmonic oscillator representation of orbits]]></category>
		<category><![CDATA[innovative methods in classical mechanics]]></category>
		<category><![CDATA[J2 perturbation]]></category>
		<category><![CDATA[Kepler problem]]></category>
		<category><![CDATA[linearization]]></category>
		<category><![CDATA[linearization of two-body problem]]></category>
		<category><![CDATA[Manev potential]]></category>
		<category><![CDATA[orbital dynamics simplification]]></category>
		<category><![CDATA[orbital mechanics]]></category>
		<category><![CDATA[perturbation analysis in orbital mechanics]]></category>
		<category><![CDATA[projective coordinate transformations]]></category>
		<category><![CDATA[projective coordinates]]></category>
		<category><![CDATA[regularization]]></category>
		<category><![CDATA[relativistic corrections in orbital models]]></category>
		<category><![CDATA[state transition matrices]]></category>
		<category><![CDATA[symplectic structure]]></category>
		<category><![CDATA[symplectic structure preservation]]></category>
		<category><![CDATA[two-body problem singularities]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=197880</guid>

					<description><![CDATA[Researchers have developed a family of projective canonical transformations that linearize Kepler and Manev orbital dynamics within a Hamiltonian framework, yielding closed-form solutions, state transition matrices, and polynomial treatment of perturbations such as Earth's J2 oblateness effect.]]></description>
										<content:encoded><![CDATA[<p>For more than three centuries, the two-body problem has stood as both the crowning achievement and the persistent headache of classical mechanics. Newton&#8217;s inverse-square law of gravitation yields elegant conic-section orbits, yet the equations of motion themselves are stubbornly nonlinear, singular at collision, and awkward to integrate numerically when perturbations creep in. A new study by Joseph T. A. Peterson, Manoranjan Majji, and John L. Junkins of Texas A&amp;M University&#8217;s Department of Aerospace Engineering, published in Celestial Mechanics and Dynamical Astronomy, offers a fresh and remarkably general way out: a family of projective coordinate transformations that, when extended canonically to a Hamiltonian framework, convert central-force orbital dynamics into linear harmonic-oscillator motion. The work applies not only to Kepler&#8217;s familiar gravity but also to Manev-type potentials that mimic relativistic corrections, and it accommodates arbitrary perturbing forces throughout.</p>
<p>The mathematical heart of the paper lies in a systematic method for extending dimension-raising point transformations into canonical transformations—the special class of coordinate changes that preserve Hamilton&#8217;s equations and the symplectic structure of phase space. The authors begin with Hamilton&#8217;s principle, requiring the action integral to be stationary along physical trajectories in both the original and the transformed coordinate systems. Because the new coordinates are redundant—four coordinates describe a three-dimensional position—they must satisfy a constraint, and the researchers handle this by introducing a Lagrange multiplier directly into the canonical structure. This yields explicit formulas for the transformed momenta and the new Hamiltonian, and it generalizes earlier work by Ferrándiz and Sansaturio while allowing for time-dependent transformations and constraints.</p>
<p>The particular family of projective transformations considered takes a position vector r and decomposes it as r = u raised to the power n, multiplied by q raised to the power m, times a vector q constrained to unit length. The parameters n and m act as adjustable knobs that generate an entire family of canonical coordinate systems, with the well-known Burdet–Ferrándiz (BF) transformation recovered for specific values. After analyzing the properties of each member of the family, the authors settle on a preferred transformation with n = m = -1, meaning r = (1/u) times the unit vector along q. This choice, they argue, subtly improves on BF at the configuration level and significantly at the momentum level, and it places attitude dynamics and angular momentum at the center of the formulation rather than at its margins.</p>
<p>Several technical virtues of the preferred choice stand out. First, radial distance becomes r = 1/u, a function only of the scalar coordinate u, so central forces are isolated entirely within the one-dimensional radial subsystem. Second, the six coordinates describing the rotational part of the motion directly characterize the attitude kinematics of the local vertical local horizontal (LVLH) frame—the rotating orthonormal basis used constantly in spacecraft operations. Third, rotational and radial motion decouple completely, and the rotational subsystem is linear for any central force whatsoever, not merely for gravity. Finally, and perhaps most strikingly, the transformation fully linearizes both Kepler and Manev dynamics without any need to carefully impose constraint conditions; the two extra integrals of motion built into the redundant formulation may be used when convenient but are never required for linearity itself.</p>
<p>Coordinate changes alone do not finish the job. The key additional step is a reparameterization of time: replacing physical time t with a new evolution parameter s defined by dt = r squared times ds, or alternatively a parameter tau defined by dt = (r squared divided by angular momentum) times d tau. For Kepler orbits, tau corresponds to the true anomaly up to an additive constant—a physically meaningful quantity long used by orbital analysts. Under either parameterization, the equations for the rotational coordinates (q, p) become those of a special orthogonal rotation, solvable in closed form via the Rodrigues rotation formula or equivalently via matrix exponentials. The radial equations, meanwhile, become those of a perturbed linear harmonic oscillator precisely when the central potential is of Manev type, V = -k1/r &#8211; k2/(2 r squared), with Kepler gravity appearing as the special case k2 = 0.</p>
<p>The Manev potential deserves particular attention. Introduced originally as a classical approximation to certain relativistic corrections, it preserves the conic orbits and integrability of the Kepler problem while adding perihelion precession—the slow rotation of an orbit&#8217;s closest approach point that also afflicts Mercury in Einstein&#8217;s general relativity. That the same projective machinery linearizes Manev dynamics as readily as Kepler dynamics suggests the framework extends naturally beyond the idealized inverse-square law. In the transformed system, the radial oscillator has a frequency determined by the difference between the squared angular momentum and the Manev parameter k2, and closed-form solutions follow immediately from standard oscillator theory.</p>
<p>Practical consequences for astrodynamics flow directly from these formal results. Because the linearized system admits closed-form solutions, closed-form state transition matrices—the sensitivity matrices that map small changes in initial conditions to changes in the final state—become readily available. State transition matrices are workhorses of modern mission design: they underpin orbit determination, uncertainty propagation, station-keeping, and targeting algorithms. Obtaining them in closed form, rather than integrating the variational equations numerically, promises both computational savings and improved long-term accuracy for orbit propagation, a goal that earlier linearization schemes such as the Kustaanheimo–Stiefel (KS) quaternionic transformation and the BF transformation have long pursued.</p>
<p>The authors demonstrate the reach of the method with a worked example of considerable practical importance: the J2-perturbed Kepler problem, often called the main satellite problem. Here the perturbation arises from the oblateness of the central body, such as Earth, whose equatorial bulge modifies the gravitational potential at leading order. Under the projective transformation, the perturbing potential becomes a polynomial expression in the new coordinates, and the resulting perturbation forces enter the linearized equations as explicit, tractable terms. This is exactly the sort of structure that perturbation theorists and designers of symplectic integrators prize, since polynomial perturbations of a linear oscillator can be handled by well-developed analytical and numerical machinery. Notably, the transformation preserves the form of angular momentum across the entire family of projective coordinates, and the constraint function and Lagrange multiplier turn out to be integrals of motion even in the presence of arbitrary, possibly nonconservative, perturbing forces—a robustness the authors verify with Poisson-bracket calculations.</p>
<p>The historical lineage of this work is rich. Regularization techniques—coordinate and time transformations that tame the singularity at collision—date back over a century, with contributions from Burdet, Vitins, Silver, Schumacher, Bond, Cid and colleagues, and a modern renaissance by Roa, Majji, Baù and collaborators. The KS transformation, rooted in spinor algebra and now commonly interpreted through quaternions, remains the most celebrated Hamiltonian linearization of the Kepler problem. Moser&#8217;s stereographic-projection approach offers yet another route. What the new paper adds is a projective alternative developed from first principles via Hamilton&#8217;s principle, unifying and extending the Burdet–Ferrándiz line of work while clarifying which choices of transformation parameters yield linear dynamics and why. The derivation that n must equal -1 for Kepler–Manev linearization is given explicitly, removing guesswork from the construction.</p>
<p>The authors point toward several promising future directions: quantitative comparisons with the KS transformation in terms of computational efficiency and long-term propagation stability, the derivation of action–angle coordinates for use in symplectic perturbation methods, and exploitation of the combined linearity and Hamiltonian structure for symplectic integration or even the quantization of classical central-force dynamics. For a field in which spacecraft navigation, debris tracking, and interplanetary mission design all hinge on accurately propagating orbits over long spans of time, a transformation that renders the underlying dynamics linear while preserving the full Hamiltonian structure is more than a mathematical curiosity—it is a practical instrument. If the promised efficiency gains materialize, these projective coordinates may soon find their way from the pages of a dynamics journal into the flight software of real missions.</p>
<p><strong>Subject of Research:</strong> Projective canonical transformations that linearize central-force orbital dynamics, including Kepler and Manev problems, in a Hamiltonian framework.</p>
<p><strong>Article Title:</strong> Projective transformations for linearized and regularized central-force dynamics: Hamiltonian formulation</p>
<p><strong>Article References:</strong> Peterson, J. T. A., Majji, M., &amp; Junkins, J. L. (2026). Projective transformations for linearized and regularized central-force dynamics: Hamiltonian formulation. <em>Celestial Mechanics and Dynamical Astronomy, 138</em>(5), Article 54. <a href="https://doi.org/10.1007/s10569-026-10304-3" rel="noopener noreferrer">https://doi.org/10.1007/s10569-026-10304-3</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s10569-026-10304-3" rel="noopener noreferrer">10.1007/s10569-026-10304-3</a></p>
<p><strong>Keywords:</strong> celestial mechanics, Hamiltonian dynamics, canonical transformations, projective coordinates, Kepler problem, Manev potential, regularization, linearization, state transition matrices, J2 perturbation, orbital mechanics, symplectic structure</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">197880</post-id>	</item>
	</channel>
</rss>
