<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>machine learning for complex decision-making &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/machine-learning-for-complex-decision-making/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Fri, 02 Oct 2026 14:17:50 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>machine learning for complex decision-making &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>AI Learns to Control Machines Safely by Watching Their Hidden Inner States</title>
		<link>https://scienmag.com/ai-learns-to-control-machines-safely-by-watching-their-hidden-inner-states/</link>
		
		<dc:creator><![CDATA[Denise Maddox]]></dc:creator>
		<pubDate>Fri, 02 Oct 2026 14:17:50 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[actor-critic]]></category>
		<category><![CDATA[adaptive control]]></category>
		<category><![CDATA[AI control of industrial machinery]]></category>
		<category><![CDATA[AI monitoring hidden system variables]]></category>
		<category><![CDATA[barrier function]]></category>
		<category><![CDATA[control algorithms for unknown systems]]></category>
		<category><![CDATA[DC motor control]]></category>
		<category><![CDATA[discrete-time system control]]></category>
		<category><![CDATA[discrete-time systems]]></category>
		<category><![CDATA[fuzzy rules]]></category>
		<category><![CDATA[hidden internal states in machine control]]></category>
		<category><![CDATA[industrial process automation safety]]></category>
		<category><![CDATA[machine learning for complex decision-making]]></category>
		<category><![CDATA[model-free control]]></category>
		<category><![CDATA[neural network-based control schemes]]></category>
		<category><![CDATA[nonlinear system control with AI]]></category>
		<category><![CDATA[output tracking]]></category>
		<category><![CDATA[reinforcement learning]]></category>
		<category><![CDATA[Reinforcement learning safety in control systems]]></category>
		<category><![CDATA[safe autonomous robotics]]></category>
		<category><![CDATA[safe control]]></category>
		<category><![CDATA[spectral entropy]]></category>
		<category><![CDATA[state constraints]]></category>
		<category><![CDATA[trial-and-error learning safety challenges]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=228211</guid>

					<description><![CDATA[A new model-free hierarchical reinforcement learning scheme keeps a DC motor's hidden internal states safely constrained while achieving precise output tracking, with theoretical guarantees and smoother, lower-entropy currents.]]></description>
										<content:encoded><![CDATA[<p>Reinforcement learning has earned a reputation for mastering games, robotics, and complex decision-making tasks, but one stubborn problem has kept it out of many real-world control rooms: safety. An algorithm that learns by trial and error can easily push a motor, a chemical reactor, or a power converter past its physical limits while it is still figuring out how the system behaves. A new study published in Neural Computing and Applications by Chidentree Treesatayapun of the Center for Research and Advanced Studies (CINVESTAV) in Ramos Arizpe, Mexico, tackles this challenge head-on with a control scheme that teaches an artificial agent to track desired outputs precisely while never letting the machine&#8217;s hidden internal variables drift into dangerous territory.</p>
<p>The research addresses a class of systems that engineers find particularly awkward: unknown, non-affine, discrete-time systems. In plain terms, these are processes sampled at fixed time steps, whose mathematical equations are either unknown or too complicated to write down, and in which the control input does not appear in a convenient, separable form. Most industrial equipment behaves this way in practice. Friction, saturation, and nonlinear electrical characteristics mean that the tidy models used in textbooks rarely match the hardware on the factory floor. Rather than trying to identify the system&#8217;s equations, the new approach is model-free: it learns directly from the data the system produces, adjusting its behavior online as measurements arrive.</p>
<p>The central insight of the work is that safety cannot be judged by the output alone. When a DC motor is commanded to follow a position trajectory, the variable that actually determines whether the machine survives is the motor current, an internal state that depends on the output and on the control action. If the controller drives the position aggressively, the current can spike, overheating the windings or tripping protective hardware. The study therefore treats constraints on these dependent internal states as first-class citizens in the control design, explicitly incorporating them into the learning objective rather than hoping that good output tracking will keep them in check as a side effect.</p>
<p>To achieve this, the author built the controller around a fuzzy-rule emulated actor-critic architecture. Actor-critic methods are a staple of reinforcement learning: one network, the actor, proposes control actions, while a second network, the critic, estimates how good those actions are and guides the actor toward better ones. The novelty here lies in the structure of these networks. Instead of conventional neural layers, they are built from fuzzy rules, which partition the input space into overlapping regions and respond with locally tuned outputs. This gives the architecture a piecewise, interpretable character that suits systems sampled in discrete time, and it allows the learning laws to be derived with clear analytical guarantees rather than relying purely on gradient heuristics.</p>
<p>Safety enters through a barrier function embedded in a hierarchical reward function. Barrier functions are a classical control-theory tool: they grow without bound as a state variable approaches a forbidden boundary, so any strategy that minimizes them is naturally repelled from unsafe regions. By folding such a barrier into a layered reward, the scheme makes the agent feel escalating penalties as the motor current nears its limits, while the upper levels of the hierarchy reward accurate tracking of the commanded position. The result is a controller that balances two competing goals in a principled way: it wants to follow the reference signal, but not at the cost of violating the constraints that keep the hardware healthy.</p>
<p>Learning in the scheme is carried out by two distinct online laws, each employing time-varying learning rates. Rather than adjusting the actor and critic with fixed step sizes, the algorithm modulates how aggressively it adapts depending on the operating conditions, which helps the networks converge quickly at the start of a task and then settle into stable, fine-grained refinement. The paper goes beyond simulation and empirical tuning: the effectiveness of these learning laws is theoretically demonstrated, with analysis showing that they ensure robust closed-loop performance. In the control community, where reinforcement learning proposals are often criticized for lacking formal guarantees, this combination of online adaptability and theoretical backing is a meaningful contribution.</p>
<p>The experimental validation is refreshingly concrete. Rather than testing only on numerical benchmarks, the author implemented the method on a real DC motor position control system, exactly the kind of electromechanical plant found in robotics, manufacturing lines, and automation equipment across industry. The results showed that the controller achieved sufficient tracking of the position reference while producing a significant reduction in the accumulating reward function compared with alternative approaches. In reinforcement learning terms, a lower accumulated reward cost means the controller found a way to do the job while incurring far less penalty, which here translates directly into gentler, safer behavior of the physical machine.</p>
<p>One of the most striking findings comes from an information-theoretic lens. The study compared the spectral entropy of the motor current, the constrained internal state, under the proposed scheme and under competing controllers. Spectral entropy measures how spread out and unpredictable a signal&#8217;s frequency content is. Under the new method, the spectral entropy of the current was significantly lower, and the probability distribution of the current was noticeably narrower. In practical terms, the motor current stayed in a tight, predictable band instead of wandering across a wide range of values. Narrow, low-entropy currents mean less thermal stress, less electrical noise, and smoother operation, which is precisely what a safety-oriented controller should deliver.</p>
<p>The implications extend well beyond a single motor on a laboratory bench. Model-free safe control of unknown discrete-time systems is a recurring need in fields as varied as wind turbine generators, electro-hydraulic servo systems, permanent magnet motor drives, and chemical process control, all of which appear in the paper&#8217;s extensive engagement with recent literature on adaptive dynamic programming, control barrier functions, and safe reinforcement learning. Many of those approaches either require a system model, handle constraints only on the input or the output, or offer no formal learning guarantees. By combining fuzzy-rule-based function approximation, hierarchical reward design, barrier-function safety, and provable online learning laws in a single model-free framework, the study offers a template that could be adapted to any plant where an observable output hides safety-critical internal states.</p>
<p>There are, of course, the usual caveats that accompany any single-author experimental study. The validation platform is a DC motor position control system, and extending the guarantees to higher-dimensional, multi-input systems with more complex constraint structures will require further work. The published article does not report external funding, and the author declares no conflict of interest. Still, the core message is likely to resonate widely: reinforcement learning controllers can be made genuinely safe, not by restricting them so heavily that they stop learning, but by building the physics of the constraints into the reward they optimize. As machines increasingly learn their own control policies online, approaches like this one, which watch the hidden variables that keep hardware alive, may become the standard grammar of trustworthy autonomous control.</p>
<p><strong>Subject of Research:</strong> Model-free hierarchical reinforcement learning for safe output tracking with internal state constraints in unknown discrete-time systems</p>
<p><strong>Article Title:</strong> Hierarchical reinforcement learning for safe output tracking: managing constraints of dependent internal states in unknown discrete-time systems</p>
<p><strong>Article References:</strong> Treesatayapun, C. (2026). Hierarchical reinforcement learning for safe output tracking: managing constraints of dependent internal states in unknown discrete-time systems. <em>Neural Computing and Applications, 38</em>(18), Article 739. <a href="https://doi.org/10.1007/s00521-026-12456-7" rel="noopener noreferrer">https://doi.org/10.1007/s00521-026-12456-7</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s00521-026-12456-7" rel="noopener noreferrer">10.1007/s00521-026-12456-7</a></p>
<p><strong>Keywords:</strong> reinforcement learning, safe control, actor-critic, fuzzy rules, barrier function, discrete-time systems, output tracking, DC motor control, state constraints, spectral entropy, model-free control, adaptive control</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">228211</post-id>	</item>
	</channel>
</rss>
