<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>drone navigation &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/drone-navigation/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Thu, 24 Sep 2026 00:05:50 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>drone navigation &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>Certifying drone-brain AI: aviation&#8217;s W-shaped safety process put to the reinforcement learning test</title>
		<link>https://scienmag.com/certifying-drone-brain-ai-aviations-w-shaped-safety-process-put-to-the-reinforcement-learning-test/</link>
		
		<dc:creator><![CDATA[Cassandra Pierce]]></dc:creator>
		<pubDate>Thu, 24 Sep 2026 00:05:50 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[AI safety assurance frameworks]]></category>
		<category><![CDATA[autonomous drone security testing]]></category>
		<category><![CDATA[aviation safety]]></category>
		<category><![CDATA[aviation safety with reinforcement learning]]></category>
		<category><![CDATA[collision avoidance]]></category>
		<category><![CDATA[drone AI certification]]></category>
		<category><![CDATA[drone attack simulation in AI safety]]></category>
		<category><![CDATA[drone navigation]]></category>
		<category><![CDATA[EASA]]></category>
		<category><![CDATA[embedded systems]]></category>
		<category><![CDATA[European Union aviation safety regulations]]></category>
		<category><![CDATA[machine learning certification]]></category>
		<category><![CDATA[machine learning regulation in aviation]]></category>
		<category><![CDATA[neural network certification for aerospace]]></category>
		<category><![CDATA[neural network safety standards]]></category>
		<category><![CDATA[neural networks]]></category>
		<category><![CDATA[operational design domain]]></category>
		<category><![CDATA[reinforcement learning]]></category>
		<category><![CDATA[reinforcement learning in safety-critical systems]]></category>
		<category><![CDATA[safety-critical systems]]></category>
		<category><![CDATA[Soft Actor–Critic]]></category>
		<category><![CDATA[verification of learned models in aviation]]></category>
		<category><![CDATA[W-shaped process]]></category>
		<category><![CDATA[W-shaped safety process in aerospace]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=211506</guid>

					<description><![CDATA[Researchers applied EASA's W-shaped safety certification process, built for supervised learning, to a reinforcement learning drone navigating an urban environment while evading an attacker, identifying key incompatibilities and proving the adapted pipeline works on real embedded hardware.]]></description>
										<content:encoded><![CDATA[<p>Safety-critical software has long been built on a simple assumption: the behavior of a program follows directly from its code. Regulators could inspect control variables, branches, and loops, and certify that a system would do what it was designed to do. Machine learning breaks that assumption, because behavior is not written but learned. Now a team of Spanish researchers has taken one of the most consequential questions in AI safety—can the aviation industry&#8217;s emerging certification framework cope with reinforcement learning, the technique behind some of the most capable autonomous systems—and answered it by building a drone that must survive a determined attacker in a simulated city.</p>
<p>The starting point is the so-called W-shaped development process, proposed by the European Union Aviation Safety Agency together with the company Daedalean in their Concepts of Design Assurance for Neural Networks reports. The classic V-model, which underpins standards such as DO-178C in avionics, ISO 26262 in automotive engineering, and the ECSS framework in aerospace, arranges development as two arms: requirements are refined into code on one side, and hierarchical testing verifies them on the other. The W-shaped process extends this with a second V dedicated to what EASA calls learning assurance—verifying that a learned model generalizes to unseen operational data and behaves robustly within its Operational Design Domain, the envelope of conditions in which it is meant to operate. For supervised learning, where datasets with labels exist before training begins, the process works. But reinforcement learning is different, and that difference matters.</p>
<p>In reinforcement learning, an agent learns by interacting with its environment: it observes a state, takes an action, and receives a reward that reflects how good the action was for achieving its goal. There are no ground-truth labels to check outputs against, and—crucially—there is no dataset before training. Training data is generated by the very interactions that produce learning. The researchers, led by Ángel-Grover Pérez-Muñoz of Universidad Politécnica de Madrid and colleagues, found that this single characteristic fractures the W-shaped process at its foundation. The process assumes data management precedes learning; in reinforcement learning the two are inseparable. EASA itself has deferred guidance on reinforcement learning to 2028, which makes an empirical test of the framework both timely and rare.</p>
<p>To probe the limits, the team designed a use case that is as much a safety exercise as an AI challenge. A defender drone, controlled entirely by a neural-network reinforcement learning agent, must follow a predefined sequence of waypoints through a simulated urban landscape built in Microsoft&#8217;s AirSim simulator, while an adversary drone tries to ram it. Three attacker strategies escalate in difficulty: a deterministic pursuer that simply chases the defender&#8217;s last known position; a predictive pursuer that estimates the defender&#8217;s position one second ahead with random scaling to mimic sensor error; and a still harder variant that randomizes its prediction horizon between one and two seconds. The defender must reach each waypoint within three meters, maintain at least 1.25 meters of separation from the attacker, and repeat this under waypoint altitudes it never saw during training.</p>
<p>The technical machinery underneath is careful and deliberate. The agent&#8217;s state is a twelve-dimensional vector built entirely from relative quantities—distance to the attacker, distance to the next waypoint, relative velocity, and the defender&#8217;s own velocity—so that the learned policy is invariant to where the drone happens to be in the world. Actions are continuous velocity commands on three axes, bounded at seven meters per second. The reward function stacks a cubic distance penalty, large bonuses for reaching waypoints and completing the course, and steep penalties for collisions, proximity violations, and wasted time. Training used the Soft Actor-Critic algorithm with four-layer actor and critic networks of 512 neurons each; notably, the team found SAC outperformed Proximal Policy Optimization because it prioritized minimizing penalties over chasing rewards—an intuitive fit for a system whose primary duty is not crashing.</p>
<p>The results carry a cautionary lesson about overtraining. An agent trained for 400,000 steps against the simple deterministic attacker achieved success rates between roughly 75 and 81 percent against all three attackers, including strategies it had never encountered, evaluated over 1,100 episodes across eleven random seeds. Its sibling trained for a full million steps on the same attacker fared dramatically worse, dropping as low as 27 percent against the predictive pursuers—a statistically significant collapse, confirmed with bootstrap tests of ten million resamples. Extended training, it appears, caused the agent to overspecialize on the deterministic chase pattern. Agents trained directly against the harder attackers did no better, with one achieving near-zero success. The lesson for safety engineers is stark: more compute and more training do not guarantee more general, more trustworthy behavior, and the certification process must be able to catch this.</p>
<p>What elevates the study beyond a simulation exercise is what happened next. The actor network—the part of the agent that turns observations into actions—was automatically translated from Python into C code compliant with MISRA-C, the conservative subset of C used in safety-critical software, which forbids dynamic memory allocation and thereby eliminates memory leaks, heap fragmentation, and unpredictable timing. The code was deployed on a Zynq UltraScale+ MPSoC board under the XtratuM hypervisor, which partitions the hardware in both space and time so that a fault in the AI component cannot propagate to other subsystems. Because the deployed agent must never learn in operation, only the frozen policy flies; the critic networks stay behind.</p>
<p>The embedded system passed its examinations. Running at a required control period of 200 milliseconds with a 100-millisecond deadline, static worst-case execution time analysis using the OTAWA tool bounded a single inference at 38 milliseconds, while 1,500 measured inferences averaged 23.62 milliseconds with a maximum of 71.18 milliseconds—comfortably inside the deadline even with interference from other tasks. Numerical comparison between the Python original and the C implementation, necessarily reduced from 64-bit to 32-bit floating point, showed errors on the order of one hundred-thousandth of a meter per second, negligible against an action range of fourteen meters per second. Success rates on the hardware—79.1, 81.4, and 79.6 percent against the three attackers—fell within the confidence intervals of the original model, demonstrating that the certified translation did not silently degrade the learned behavior.</p>
<p>The broader significance lies in the mapping the researchers drew between reinforcement learning&#8217;s needs and the W-shaped process&#8217;s stages. They found that the entire development-assurance half of the process—implementation, integration, and verification on real hardware—is agnostic to the learning technique and applied cleanly. The learning-assurance half required reinvention: instead of predefined datasets, the team defined scenarios derived from the Operational Design Domain to govern data generation; data quality requirements such as representativeness and completeness, which normally apply to whole datasets, were reinterpreted as properties of those scenarios; and a new step was inserted to define the agent-environment interface—the states, actions, and reward function—before training could begin.</p>
<p>This is, by the authors&#8217; account, the first application of EASA&#8217;s W-shaped process to a reinforcement learning system, and it arrives as EUROCAE and SAE finalize the ARP6983/ED-324 standard for aeronautical AI. The work suggests a realistic path forward: certification frameworks built for supervised learning can stretch to accommodate learning agents that discover their own strategies, provided the emphasis shifts from curating data to curating the scenarios that generate it. The researchers caution that their results live in simulation, and closing the sim-to-real gap—sensor noise, communication delays, environmental perturbations—remains future work, as does confronting learning-based attackers. But for a field waiting on regulators until 2028, the demonstration that a reinforcement learning drone can be trained, certified, and embedded within real-time safety constraints is a milestone worth more than its 71-millisecond inference time might suggest.</p>
<p><strong>Subject of Research:</strong> Applying the EASA W-shaped development process to reinforcement learning for safe drone navigation and collision avoidance</p>
<p><strong>Article Title:</strong> Application of the W-shaped process for a Reinforcement Learning use case on drone navigation</p>
<p><strong>Article References:</strong> Pérez-Muñoz, Á.-G., García-Quijano, H., López-García, G., Alonso, A., &amp; Pérez, M. S. (2026). Application of the W-shaped process for a Reinforcement Learning use case on drone navigation. <em>Machine Learning with Applications, 26</em>, Article 101009. <a href="https://doi.org/10.1016/j.mlwa.2026.101009" rel="noopener noreferrer">https://doi.org/10.1016/j.mlwa.2026.101009</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1016/j.mlwa.2026.101009" rel="noopener noreferrer">10.1016/j.mlwa.2026.101009</a></p>
<p><strong>Keywords:</strong> reinforcement learning, drone navigation, aviation safety, EASA, W-shaped process, machine learning certification, neural networks, embedded systems, collision avoidance, Soft Actor-Critic, operational design domain, safety-critical systems</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">211506</post-id>	</item>
		<item>
		<title>Physics-Guided AI Teaches Drones to Fly Smarter in Three Dimensions</title>
		<link>https://scienmag.com/physics-guided-ai-teaches-drones-to-fly-smarter-in-three-dimensions/</link>
		
		<dc:creator><![CDATA[Katie Riggs]]></dc:creator>
		<pubDate>Sun, 13 Sep 2026 01:52:53 +0000</pubDate>
				<category><![CDATA[Space]]></category>
		<category><![CDATA[3D drone path planning]]></category>
		<category><![CDATA[3D path planning]]></category>
		<category><![CDATA[artificial potential field]]></category>
		<category><![CDATA[autonomous flight]]></category>
		<category><![CDATA[complex environment drone trajectory optimization]]></category>
		<category><![CDATA[convergence speed in drone AI training]]></category>
		<category><![CDATA[deep reinforcement learning]]></category>
		<category><![CDATA[deep reinforcement learning in UAVs]]></category>
		<category><![CDATA[disaster relief drone automation]]></category>
		<category><![CDATA[drone autonomous navigation]]></category>
		<category><![CDATA[drone navigation]]></category>
		<category><![CDATA[energy efficiency]]></category>
		<category><![CDATA[energy-efficient drone flight algorithms]]></category>
		<category><![CDATA[environmental monitoring using autonomous drones]]></category>
		<category><![CDATA[hybrid AI approaches for autonomous flight]]></category>
		<category><![CDATA[long short-term memory]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[obstacle avoidance in drone navigation]]></category>
		<category><![CDATA[physics-guided artificial intelligence for drones]]></category>
		<category><![CDATA[robotics]]></category>
		<category><![CDATA[smooth and feasible drone trajectories]]></category>
		<category><![CDATA[Soft Actor–Critic]]></category>
		<category><![CDATA[trajectory optimization]]></category>
		<category><![CDATA[UAV path planning]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=200604</guid>

					<description><![CDATA[Researchers have developed a hybrid AI method that combines artificial potential field guidance with an LSTM-enhanced Soft Actor–Critic algorithm to plan faster, smoother, and more energy-efficient 3D drone trajectories.]]></description>
										<content:encoded><![CDATA[<p>Unmanned aerial vehicles have become indispensable tools for environmental monitoring, disaster relief, and logistics, yet the task of teaching a drone to chart its own course through a complex three-dimensional world remains one of the hardest problems in autonomous flight. A new study published in the International Journal of Aeronautical and Space Sciences presents a hybrid artificial intelligence approach that promises to make drone path planning faster to learn, smoother to execute, and cheaper to fly. The method, developed by Jianhua Liu, Haitao Zhou, Xia Lei, Xiaoguang Tu, Houqiang Hua, and Xiaofan Wang, combines classical physics-based guidance with modern deep reinforcement learning, and it delivers measurable gains in convergence speed and energy efficiency over conventional learning-based planners.</p>
<p>The core challenge the researchers set out to solve is deceptively simple to state: given a start point and a goal in a cluttered 3D environment, generate a trajectory that is short, smooth, dynamically feasible, and energy-efficient, all at the same time. These objectives frequently conflict with one another. The shortest path may hug obstacles so tightly that a real aircraft could not safely follow it. The smoothest path may waste energy in wide, sweeping arcs. Traditional optimization methods can balance these goals but often struggle in unknown or changing environments, where they must be re-run from scratch whenever conditions shift.</p>
<p>Deep reinforcement learning has emerged in recent years as an attractive end-to-end alternative. In this paradigm, an artificial intelligence agent learns to fly by trial and error, receiving rewards for progress toward the goal and penalties for collisions, erratic motion, or excessive energy use. Over many training episodes, the agent internalizes a policy that maps what it observes directly to the actions it should take. The appeal is obvious: once trained, such a policy can react to new situations in milliseconds without recomputing a full trajectory. The drawback, as the new paper emphasizes, is that learning from scratch is painfully slow. Random exploration in a vast three-dimensional action space means the agent spends most of its early training bumping into obstacles or wandering aimlessly, a problem the authors describe as blind exploration combined with low sample efficiency.</p>
<p>To inject common sense into this process, the team turned to the artificial potential field, a concept that has guided robot navigation since the 1980s. In an artificial potential field, the goal exerts an attractive force that pulls the vehicle toward it, while obstacles exert repulsive forces that push it away. The drone is imagined as a ball rolling downhill on a landscape sculpted by these forces, and at every instant the field suggests a sensible direction of travel. The elegance of the approach is that the guidance is state-dependent: as the drone&#8217;s position and surroundings change, the forces change with them, always pointing toward safer, more productive regions of space.</p>
<p>Rather than replacing the learning algorithm with this classical method, the researchers fused the two. The state-dependent forces generated by the artificial potential field are integrated directly with the policy of a Soft Actor–Critic agent, a state-of-the-art deep reinforcement learning algorithm prized for its stability and its ability to balance exploration against exploitation. Soft Actor–Critic maximizes both the expected reward and the entropy, or randomness, of the agent&#8217;s behavior, which prevents it from collapsing prematurely into a mediocre strategy. By blending the potential field&#8217;s heuristic push into the action-selection process, the hybrid system no longer explores blindly. From the very first training episode, the agent is nudged in directions that the physics suggests are promising, while retaining the freedom to deviate when the heuristic is wrong, as it can be in local minima where attractive and repulsive forces cancel out.</p>
<p>The second innovation addresses a different weakness: memory. Standard actor–critic networks treat each moment in isolation, deciding what to do based only on the current observation. A flying vehicle, however, is a dynamical system whose future depends on its recent past. A sudden change in heading that was perfectly safe at low speed may be catastrophic at high speed, and the network cannot know the difference if it has no access to the sequence of states that led to the present moment. To capture these temporal dependencies, the authors embedded a long short-term memory, or LSTM, structure into both the actor and the critic networks. LSTMs are recurrent neural networks equipped with gating mechanisms that allow them to retain information over many time steps and to forget what is no longer relevant, giving the agent an effective working memory of its own flight history.</p>
<p>The practical consequence of this architectural choice is smoother, more dynamically feasible trajectories. Because the policy can perceive trends, accelerations, and oscillations in the state sequence rather than single snapshots, it learns to produce control commands that flow naturally from one to the next, avoiding the jerky, high-frequency corrections that plague memoryless policies and that translate directly into wasted energy and mechanical stress on real airframes. The combination of potential field guidance and recurrent memory gives the method its name: APF–LSTM–SAC.</p>
<p>The team validated the approach in comprehensive tests within complex simulated three-dimensional environments, comparing it against pure deep reinforcement learning baselines. The results were striking. The hybrid method achieved improvements in convergence speed of at least 15.14 percent, meaning the agent reached competent flight policies substantially faster than its unguided counterparts, and it reduced energy consumption by at least 10.36 percent, a figure that matters enormously for battery-powered aircraft whose mission endurance is measured in minutes. Faster training also carries a practical dividend: fewer simulated flight hours are needed before a policy is deployable, which lowers the computational cost of developing autonomous capabilities for new vehicle types or new environments.</p>
<p>The significance of the work extends beyond the specific percentages. It exemplifies a growing trend in robotics toward physics-informed machine learning, in which decades of classical control theory and heuristic reasoning are used to scaffold, rather than be replaced by, modern data-driven methods. Pure learning systems must rediscover from scratch lessons that engineers already know, such as the fact that obstacles should be avoided and goals approached. By encoding those lessons as inductive biases inside the learning pipeline, researchers can preserve the adaptability of reinforcement learning while dramatically shrinking the search space the algorithm must explore. The potential field component supplies a sensible prior; the LSTM supplies temporal awareness; and the Soft Actor–Critic framework supplies robust, entropy-regularized optimization that can gracefully reconcile the two.</p>
<p>The authors note that the datasets generated and analyzed during the study are available from the corresponding author on reasonable request, and the work was supported by the National Natural Science Foundation of China, the CAAC Key Laboratory of General Aviation Operation, and the Fundamental Research Funds for the Central Universities. As drones take on ever more ambitious roles, from delivering medical supplies to surveying disaster zones, the ability to plan safe, efficient three-dimensional trajectories autonomously will only grow in importance. Hybrid approaches like APF–LSTM–SAC suggest that the fastest route to capable autonomous flight may not be to make learning algorithms bigger, but to make them wiser, by letting the accumulated physics of navigation light the way through the darkness of blind exploration.</p>
<p><strong>Subject of Research:</strong> UAV three-dimensional path planning using artificial potential field guidance and an LSTM-enhanced Soft Actor–Critic deep reinforcement learning algorithm</p>
<p><strong>Article Title:</strong> UAV 3D Path Planning Based on Artificial Potential Field Guidance and LSTM-Enhanced Soft Actor–Critic</p>
<p><strong>Article References:</strong> Liu, J., Zhou, H., Lei, X., Tu, X., Hua, H., &amp; Wang, X. (2026). UAV 3D Path Planning Based on Artificial Potential Field Guidance and LSTM-Enhanced Soft Actor–Critic. <em>International Journal of Aeronautical and Space Sciences</em>. <a href="https://doi.org/10.1007/s42405-026-01292-7" rel="noopener noreferrer">https://doi.org/10.1007/s42405-026-01292-7</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s42405-026-01292-7" rel="noopener noreferrer">10.1007/s42405-026-01292-7</a></p>
<p><strong>Keywords:</strong> UAV path planning, deep reinforcement learning, artificial potential field, Soft Actor–Critic, long short-term memory, drone navigation, trajectory optimization, energy efficiency, 3D path planning, autonomous flight, machine learning, robotics</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">200604</post-id>	</item>
	</channel>
</rss>
