<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>temporal difference learning &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/temporal-difference-learning/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Sun, 04 Oct 2026 04:16:18 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>temporal difference learning &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>Adaptive Reinforcement Learning Algorithm Steers Drones to Moving Chargers Faster</title>
		<link>https://scienmag.com/adaptive-reinforcement-learning-algorithm-steers-drones-to-moving-chargers-faster/</link>
		
		<dc:creator><![CDATA[Grant Pearson]]></dc:creator>
		<pubDate>Sun, 04 Oct 2026 04:16:18 +0000</pubDate>
				<category><![CDATA[Space]]></category>
		<category><![CDATA[adaptive algorithms for dynamic environments]]></category>
		<category><![CDATA[advanced reinforcement learning algorithms in aerospace]]></category>
		<category><![CDATA[AI-driven drone endurance enhancement]]></category>
		<category><![CDATA[ASDE algorithm]]></category>
		<category><![CDATA[autonomous flight]]></category>
		<category><![CDATA[convergence improvement in reinforcement learning]]></category>
		<category><![CDATA[drone battery charging]]></category>
		<category><![CDATA[dynamic environments]]></category>
		<category><![CDATA[dynamic obstacle and target tracking in drone missions]]></category>
		<category><![CDATA[Electric drone battery management]]></category>
		<category><![CDATA[epsilon-greedy exploration]]></category>
		<category><![CDATA[intelligent routing for unmanned aerial vehicles]]></category>
		<category><![CDATA[mobile charging vehicle]]></category>
		<category><![CDATA[mobile drone charging solutions]]></category>
		<category><![CDATA[path planning]]></category>
		<category><![CDATA[persistent flight path optimization]]></category>
		<category><![CDATA[real-time drone charging strategies]]></category>
		<category><![CDATA[reinforcement learning]]></category>
		<category><![CDATA[reinforcement learning for autonomous vehicle navigation]]></category>
		<category><![CDATA[SARSA]]></category>
		<category><![CDATA[Shenyang Jianzhu University]]></category>
		<category><![CDATA[temporal difference learning]]></category>
		<category><![CDATA[UAV]]></category>
		<category><![CDATA[UAV path planning with moving targets]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=233454</guid>

					<description><![CDATA[Researchers in China have developed an adaptive reinforcement learning algorithm called ASDE that plans drone flight paths to moving charging vehicles, converging about 45 percent faster than conventional SARSA while improving success rates and producing shorter, smoother routes.]]></description>
										<content:encoded><![CDATA[<p>Electric drones have a stubborn problem: their batteries run out long before their missions do. Extending flight time by swapping in bigger packs adds weight and erodes the very endurance gains engineers are chasing, so researchers have increasingly turned to an alternative vision of persistent flight in which unmanned aerial vehicles periodically rendezvous with mobile charging vehicles on the ground, topping up their cells mid-mission before returning to work. The catch is that planning a flight path toward a charger that refuses to sit still is a genuinely hard computational problem, and a new study published in the International Journal of Aeronautical and Space Sciences argues that the reinforcement learning methods commonly used for it are not up to the task. A team at Shenyang Jianzhu University led by Dan Shan, Meng Zhang, Jianwei He, Tianyu Zhang and Yanfeng Li has now proposed a redesigned learning algorithm, called ASDE, that they report converges roughly 45 percent faster than the conventional approach it builds upon.</p>
<p>The core difficulty lies in the mismatch between how standard reinforcement learning agents learn and how a moving charging vehicle actually behaves. In the classic SARSA algorithm, an agent learns by trial and error, updating the value of state-action pairs as it experiences rewards and penalties. That works reasonably well when the world is static, because the consequences of a given action in a given cell of a grid map stay consistent across episodes. But a mobile charging vehicle follows irregular motion patterns that are difficult to predict, which means the reward landscape itself shifts beneath the learner. The researchers identify three intertwined failure modes in this setting: exploration is inefficient because the agent wastes effort sampling regions of the environment that no longer matter, the vehicle&#8217;s motion is irregular enough to defeat simple models, and prediction errors about where the charger will be grow large enough to poison the learned policy.</p>
<p>ASDE, which stands for adaptive SARSA for dynamic environments, attacks all three problems at once. The first ingredient is an explicit motion model of the mobile charging vehicle woven directly into the learning framework, so that the drone&#8217;s planner has a structured expectation of how the charger moves rather than treating its position as an arbitrary, unknowable quantity. The second is a restructured reward function that reshapes the feedback the agent receives, steering it toward trajectories that are not merely successful but also efficient and smooth. Together these changes give the learning process a much better-shaped objective, reducing the amount of random wandering the agent must do before it discovers useful behavior.</p>
<p>The third ingredient is perhaps the most conceptually interesting: a time-varying epsilon-greedy strategy. In textbook reinforcement learning, epsilon-greedy exploration means the agent takes a random action with a fixed probability epsilon and otherwise exploits its current knowledge. A fixed epsilon is a blunt instrument. Too high, and the agent never settles into the good policy it has found; too low, and it stops exploring before it has found one. ASDE instead adapts the exploration rate continuously according to environmental feedback, exploring more aggressively when the situation is uncertain or the charger&#8217;s motion has invalidated prior assumptions, and exploiting more heavily when the environment appears stable and the learned value estimates are trustworthy. This dynamic balance is what allows the algorithm to remain responsive in a setting where the target of the entire mission is itself in motion.</p>
<p>The fourth component extends the algorithm&#8217;s memory of its own trajectory. Standard SARSA is a one-step temporal difference method: each update looks only one step into the past. ASDE employs a hybrid temporal difference lambda mechanism with multi-step backtracking, which propagates credit backward across a stretch of recent states and actions rather than a single transition. Crucially, the effective backtracking horizon is not fixed. It adjusts dynamically based on the mobile charging vehicle&#8217;s instantaneous motion, stretching out when the charger&#8217;s behavior is predictable and contracting when it changes abruptly. The researchers report that this adaptive backtracking enhances both predictive accuracy and responsiveness, allowing the drone&#8217;s value estimates to track a moving target without the lag that plagues fixed-horizon methods.</p>
<p>To find out whether these design choices actually matter, the team ran simulation experiments across grid environments ranging from a compact 20 by 20 layout to a more demanding 50 by 50 space, comparing ASDE against benchmark algorithms including conventional SARSA. The results, as summarized in the paper, are consistent across scales. ASDE improved task success rates by between 4.4 and 10.4 percent relative to the benchmarks, a meaningful margin in a domain where a failed rendezvous can mean a drone stranded far from its base. The learned paths were also substantially shorter, with reductions of 21.2 to 30.7 percent in path length, which translates directly into energy saved and mission time recovered.</p>
<p>One of the most practically significant findings concerns path quality rather than raw performance. ASDE reduced the number of inflection points along planned trajectories by 33.0 to 46.0 percent compared with the benchmark algorithms. Inflection points are the sharp corners where a path changes direction abruptly, and for a flying vehicle each one costs energy and imposes maneuvering loads. A smoother path is easier for a flight controller to track accurately, gentler on the airframe, and more predictable for any surrounding traffic. That a learning algorithm can produce paths that are simultaneously shorter, more reliable, and smoother suggests the restructured reward function and adaptive exploration are doing real work in shaping the geometry of the solutions, not just the success statistics.</p>
<p>The convergence result deserves particular attention because learning speed is often the hidden bottleneck in deploying reinforcement learning on real hardware. An agent that needs millions of training episodes to converge is impractical to train in simulation for every new environment a drone might face. By reporting that ASDE reaches convergence approximately 45 percent faster than conventional SARSA, the authors are making a claim about deployability as much as about benchmark performance. Faster convergence means the planner can adapt more quickly when the mobile charging vehicle&#8217;s behavior shifts, and it lowers the computational cost of retraining, both of which matter for operations where conditions change from mission to mission.</p>
<p>The broader context makes the work timely. Drones are being pressed into service for search and rescue, infrastructure inspection, agricultural monitoring, and delivery, and in many of these roles the endurance limit of batteries is the binding constraint on how the technology can be used. Charging infrastructure that moves with the mission, rather than waiting at a fixed pad, is one of the more elegant proposed answers, and the concept has close cousins in research on electric vehicles and mobile robotic refueling. What has been missing is a planning layer robust enough to handle the uncertainty of a charger that is itself navigating traffic, terrain, and its own constraints. The Shenyang team&#8217;s contribution is a concrete, quantitatively evaluated step toward that layer, grounded in one of the workhorse algorithms of reinforcement learning rather than requiring an entirely new theoretical apparatus.</p>
<p>There are, of course, the usual caveats that separate simulation from the sky. The reported gains come from grid-world experiments of up to 50 by 50 cells, and real flight adds wind, sensing noise, communication delays, and safety constraints that no grid abstraction fully captures. The authors state that the datasets generated and analyzed in the study are available from the corresponding author on reasonable request, and the work was published in the International Journal of Aeronautical and Space Sciences on 14 July 2026 under the auspices of the Korean Society for Aeronautical and Space Sciences. Even with those caveats, the pattern of results is striking enough to suggest that the adaptive machinery at the heart of ASDE, the feedback-driven exploration schedule and the motion-aware temporal difference backtracking, could generalize beyond charging rendezvous to any drone task that involves intercepting a moving target. For a field where the difference between a 60 percent and a 70 percent success rate can determine whether an autonomous mission is viable at all, an algorithm that delivers higher success, shorter paths, smoother trajectories, and faster learning in a single package is the kind of incremental engineering advance that quietly enables the next generation of flying robots.</p>
<p><strong>Subject of Research:</strong> Reinforcement learning-based path planning for UAV dynamic charging with mobile charging vehicles</p>
<p><strong>Article Title:</strong> ASDE Algorithm-Based UAV Dynamic Charging Path Planning Method</p>
<p><strong>Article References:</strong> Shan, D., Zhang, M., He, J., Zhang, T., &amp; Li, Y. (2026). ASDE Algorithm-Based UAV Dynamic Charging Path Planning Method. <em>International Journal of Aeronautical and Space Sciences</em>. <a href="https://doi.org/10.1007/s42405-026-01253-0" rel="noopener noreferrer">https://doi.org/10.1007/s42405-026-01253-0</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s42405-026-01253-0" rel="noopener noreferrer">10.1007/s42405-026-01253-0</a></p>
<p><strong>Keywords:</strong> UAV, reinforcement learning, SARSA, ASDE algorithm, mobile charging vehicle, path planning, dynamic environments, temporal difference learning, epsilon-greedy exploration, drone battery charging, autonomous flight, Shenyang Jianzhu University</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">233454</post-id>	</item>
	</channel>
</rss>
