<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>partial observability &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/partial-observability/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Sun, 13 Sep 2026 02:30:37 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.1</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>partial observability &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>New AI Framework Tames Chaotic Teamwork in Multi-Agent Reinforcement Learning</title>
		<link>https://scienmag.com/new-ai-framework-tames-chaotic-teamwork-in-multi-agent-reinforcement-learning/</link>
		
		<dc:creator><![CDATA[Denise Maddox]]></dc:creator>
		<pubDate>Sun, 13 Sep 2026 02:30:37 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[adaptive coalition formation]]></category>
		<category><![CDATA[Bayesian belief fusion]]></category>
		<category><![CDATA[Bayesian-Elite adaptive coalition network]]></category>
		<category><![CDATA[coalition formation]]></category>
		<category><![CDATA[Complex & Intelligent Systems]]></category>
		<category><![CDATA[cooperative AI]]></category>
		<category><![CDATA[cooperative artificial intelligence]]></category>
		<category><![CDATA[graph attention networks]]></category>
		<category><![CDATA[Hanabi]]></category>
		<category><![CDATA[hierarchical hybrid control]]></category>
		<category><![CDATA[MAPPO]]></category>
		<category><![CDATA[multi-agent coordination strategies]]></category>
		<category><![CDATA[multi-agent reinforcement learning]]></category>
		<category><![CDATA[multi-agent reinforcement learning framework]]></category>
		<category><![CDATA[multi-agent teamwork challenges]]></category>
		<category><![CDATA[noisy communication in AI]]></category>
		<category><![CDATA[partial observability]]></category>
		<category><![CDATA[partially observable environments]]></category>
		<category><![CDATA[policy stabilisation]]></category>
		<category><![CDATA[real-world autonomous agent applications]]></category>
		<category><![CDATA[reproducibility]]></category>
		<category><![CDATA[University of Yaoundé I]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=200864</guid>

					<description><![CDATA[Researchers have developed H3C-BEACON, a unified multi-agent reinforcement learning framework that jointly integrates communication, Bayesian belief inference, adaptive coalition formation, and policy stabilisation to achieve major gains and unprecedented reproducibility on cooperative AI benchmarks.]]></description>
										<content:encoded><![CDATA[<p>Teaching a team of artificial intelligence agents to cooperate has long been one of the most stubborn problems in machine learning. Each agent sees only a fragment of the world, the environment shifts beneath them as they learn, and the messages they exchange are often incomplete or noisy. Now, researchers at the University of Yaoundé I in Cameroon have unveiled a unified framework that tackles all of these challenges at once, and the results suggest a meaningful step forward for cooperative artificial intelligence. The framework, called H3C-BEACON — short for Hierarchical Hybrid Heterogeneous Control with Bayesian-Elite Adaptive Coalition Network — is described in a peer-reviewed position paper published open access in the journal Complex &amp; Intelligent Systems.</p>
<p>The problem the researchers set out to solve is deceptively simple to state. In multi-agent reinforcement learning, or MARL, several autonomous agents learn by trial and reward to accomplish tasks together, much like players learning a team sport. When every agent can see the full state of the world, coordination is tractable. But real-world settings — fleets of delivery drones, robotic warehouses, autonomous vehicles negotiating traffic — are only partially observable and constantly changing. Each agent must simultaneously infer what it cannot see, decide what to communicate to its teammates, figure out which teammates it should coordinate with, and keep its learning process stable enough that early mistakes do not cascade into collapsed policies. Most existing methods address these demands with separate, independent mechanisms, and the authors argue that the interactions between those mechanisms have been chronically underexploited.</p>
<p>H3C-BEACON&#8217;s central contribution is to fold six complementary components into a single, coherent optimisation loop. The first is a Dynamic Graph Attention Network, or DGAT, that governs communication. Rather than flooding every agent with information from every other agent, the network learns distance-aware attention weights, so each agent focuses its message exchange on the neighbours that matter most for the task at hand. This keeps the communication overhead manageable while preserving the information that actually drives good coordination.</p>
<p>The second component addresses the epistemic fog of partial observability. Each agent maintains probabilistic beliefs about the hidden state of the environment and fuses those beliefs with the estimates of its teammates using Bayesian inference. When two agents hold slightly different beliefs about the same uncertain variable, the fusion process weighs the evidence and produces a sharper joint estimate than either agent could achieve alone. Third, the framework introduces spectral coalition formation: a mechanism that dynamically groups agents into specialised coalitions based on the structure of their interactions. Instead of fixing roles in advance, the system lets functional specialisation emerge from the spectral properties of the agents&#8217; interaction graph, allowing the team to reorganise itself as the task demands.</p>
<p>The remaining three components concern learning stability, which is where many multi-agent systems quietly fall apart. A dual-critic architecture separates the evaluation of global coordination from local decision making, so that an agent&#8217;s individual contribution can be assessed without conflating it with the noise of its teammates&#8217; behaviour. The fourth and arguably most distinctive mechanism, called RTD++ elite-trajectory anchoring, constrains the evolving policy to stay within a bounded distance — measured as a Kullback-Leibler divergence — of a set of elite trajectories collected during training. The authors provide theoretical support for this idea, proving a covering-number bound showing that policies constrained in this way occupy a small, well-behaved region of parameter space, which in turn supports more reliable optimisation. Finally, bounded entropy control keeps the exploration-exploitation balance from swinging wildly: agents are encouraged to explore, but never so much that the policy dissolves into randomness.</p>
<p>The empirical results are striking in the environments where the framework&#8217;s design assumptions hold. On the Multi-Agent Particle Environments, a standard family of cooperative benchmarks, H3C-BEACON consistently outperformed MAPPO, a widely used and strong baseline algorithm. In the communication-intensive simple_world_comm scenario, the framework achieved a perfect win rate across all five independent random seeds, and lifted the best episode reward from −6.06 ± 0.70 under MAPPO to −2.35 ± 0.62. In simple_spread, a coordination task in which agents must cover landmarks while avoiding collisions, the most telling result was not the raw score but the variance: H3C-BEACON produced a 95 percent confidence interval roughly 28 times narrower than MAPPO&#8217;s, at ±0.57 versus ±15.90. For practitioners, that near-elimination of performance variability across random initialisations may matter as much as the improvement in average performance, because reproducibility has been a chronic weakness of deep multi-agent learning.</p>
<p>The clearest demonstration of the framework&#8217;s stabilisation machinery came from Hanabi-full, a cooperative card game in which players see everyone else&#8217;s cards but never their own. Under this severe partial observability, H3C-BEACON raised the mean score from 2.29 ± 0.23 to 3.96 ± 0.82, a 73 percent improvement, and — crucially — avoided policy collapse in every run. The authors attribute this robustness directly to RTD++, which anchors the policy to elite trajectories and prevents the catastrophic forgetting and sudden performance crashes that frequently end multi-agent training runs prematurely.</p>
<p>The picture is not uniformly rosy, and the authors are candid about it. On StarCraft combat scenarios, MAPPO remained superior. The team argues this is consistent with the structural properties of that environment rather than a flaw in their approach: StarCraft micromanagement involves homogeneous units, a dense and fully observable global state, and no explicit communication channel that would benefit from graph attention or coalition formation. In other words, the very components that give H3C-BEACON its edge in communication-heavy, imperfect-information settings offer little purchase in an environment that strips those challenges away. The authors also report computational costs honestly: the full framework processes roughly 50 environment steps per second in its dense configuration, compared with about 200 for MAPPO, reflecting the price of running six interacting components per episode.</p>
<p>Ablation experiments reinforce the claim that the architecture&#8217;s strength lies in the integration of its parts rather than any single trick. Removing DGAT cost 28 percent of the win rate, while removing either RTD++ or the coalition formation mechanism caused the largest degradation, cutting the win rate by roughly 70 percentage points on simple_spread. Learning-curve analyses showed that variants lacking RTD++ often failed to reach 90 percent of the best reward within 500,000 training steps at all. A sensitivity analysis further confirmed that the qualitative ranking of algorithms was robust to perturbations of the win-rate thresholds, with no rank reversals across seeds, suggesting the reported advantages are not artefacts of how success was measured. All primary results were computed over five independent random seeds with 95 percent confidence intervals.</p>
<p>What emerges from the paper is an argument about philosophy as much as engineering. The authors contend that communication, belief estimation, coalition formation, and stable optimisation should not be bolted together post hoc but jointly modelled from the start, because their benefits compound: better beliefs make communication more informative, coalitions make coordination more targeted, and anchored optimisation preserves the gains long enough for them to materialise. If the framework&#8217;s limitations on fully observable, homogeneous environments are acknowledged, its performance in the messy, partially observable, decentralised settings that resemble real-world deployment is precisely where cooperative AI most needs help. For a field haunted by irreproducible results and collapsed training runs, a method that delivers a perfect win rate on one benchmark, a twenty-eight-fold reduction in variance on another, and zero policy collapses on a third is a result the community will be watching closely.</p>
<p><strong>Subject of Research:</strong> A unified hierarchical framework for cooperative multi-agent reinforcement learning in partially observable environments</p>
<p><strong>Article Title:</strong> H3C-BEACON: hierarchical hybrid heterogeneous control with Bayesian-elite adaptive coalition network for multi-agent reinforcement learning</p>
<p><strong>Article References:</strong> H3C-BEACON: hierarchical hybrid heterogeneous control with Bayesian-elite adaptive coalition network for multi-agent reinforcement learning. (n.d.). <a href="https://doi.org/10.1007/s40747-026-02494-y" rel="noopener noreferrer">https://doi.org/10.1007/s40747-026-02494-y</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s40747-026-02494-y" rel="noopener noreferrer">10.1007/s40747-026-02494-y</a></p>
<p><strong>Keywords:</strong> multi-agent reinforcement learning, cooperative AI, partial observability, Bayesian belief fusion, graph attention networks, coalition formation, policy stabilisation, MAPPO, Hanabi, Complex &amp; Intelligent Systems, University of Yaoundé I, reproducibility</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">200864</post-id>	</item>
		<item>
		<title>AI Guidance Helps Spacecraft and Defenders Outsmart Unknown Attackers in Orbit</title>
		<link>https://scienmag.com/ai-guidance-helps-spacecraft-and-defenders-outsmart-unknown-attackers-in-orbit/</link>
		
		<dc:creator><![CDATA[Grant Pearson]]></dc:creator>
		<pubDate>Sat, 12 Sep 2026 20:13:29 +0000</pubDate>
				<category><![CDATA[Space]]></category>
		<category><![CDATA[active defense]]></category>
		<category><![CDATA[active spacecraft interception methods]]></category>
		<category><![CDATA[AI-assisted orbital defense tactics]]></category>
		<category><![CDATA[convolutional neural network]]></category>
		<category><![CDATA[cooperative satellite maneuver coordination]]></category>
		<category><![CDATA[deep Q-network]]></category>
		<category><![CDATA[gated recurrent unit]]></category>
		<category><![CDATA[multi-agent space engagement dynamics]]></category>
		<category><![CDATA[Multi-POMDP]]></category>
		<category><![CDATA[orbital defense guidance systems]]></category>
		<category><![CDATA[partial observability]]></category>
		<category><![CDATA[pursuit and evasion in congested orbit]]></category>
		<category><![CDATA[pursuit-evasion]]></category>
		<category><![CDATA[reinforcement learning]]></category>
		<category><![CDATA[satellite collision avoidance techniques]]></category>
		<category><![CDATA[space conflict and countermeasures]]></category>
		<category><![CDATA[space security]]></category>
		<category><![CDATA[space situational awareness and defense]]></category>
		<category><![CDATA[spacecraft guidance]]></category>
		<category><![CDATA[spacecraft pursuit–evasion strategies]]></category>
		<category><![CDATA[spacecraft survivability]]></category>
		<category><![CDATA[three-body space engagement modeling]]></category>
		<category><![CDATA[unknown attacker interception in space]]></category>
		<category><![CDATA[zero-effort miss distance]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=198232</guid>

					<description><![CDATA[Researchers at Sun Yat-sen University developed an adaptive deep reinforcement learning guidance method that lets a target spacecraft and its defender cooperatively evade pursuers using unknown strategies under noisy, incomplete information.]]></description>
										<content:encoded><![CDATA[<p>Space is becoming crowded, and the high-value orbital regimes where communications, navigation, and reconnaissance satellites operate are increasingly contested. Among the many challenges that follow from this congestion, one of the most technically demanding is the pursuit–evasion confrontation: a scenario in which a hostile maneuvering spacecraft attempts to intercept a target vehicle that must survive the encounter. A research team at the School of Aeronautics and Astronautics of Sun Yat-sen University has now introduced an active defense guidance method designed for exactly this situation, in which a target spacecraft does not rely on evasion alone but releases a defensive vehicle that counter-intercepts the incoming pursuer. The work, published in Space: Science &amp; Technology, addresses a problem that has long frustrated mission designers: how can two cooperating spacecraft coordinate their maneuvers when they cannot see the full picture and do not know which interception strategy the attacker is using?</p>
<p>The scenario the researchers studied is a three-body engagement involving the target, the pursuer, and the defender. Once the defender is deployed, two coupled pursuit–evasion relationships emerge simultaneously: the pursuer chases the target, while the defender chases the pursuer in an attempt to spoil the interception. Because the defender bends the pursuer&#8217;s trajectory away from the target, the two friendly vehicles must act in a coordinated fashion rather than as independent agents. The difficulty is compounded by the fact that the pursuer may employ any of several established interception laws, including optimal control guidance, proportional navigation, or differential game guidance. A defense scheme tuned to a single assumed attack strategy can be defeated the moment the adversary switches tactics. At the same time, neither the target nor the defender has access to complete state information; both must rely on noisy local measurements of range and line-of-sight angles, which makes the problem one of partial observability as well as multi-strategy adversarial behavior.</p>
<p>Existing guidance approaches fall short in this setting for distinct reasons. Unilateral optimal control methods assume a known, fixed adversary model and cannot adapt when the pursuer changes strategy mid-engagement. Differential game formulations deliver elegant theoretical solutions but typically presume perfect information on both sides, an assumption that collapses in the presence of sensor noise and hidden intentions. Conventional reinforcement learning, meanwhile, has shown promise in adversarial settings, but standard algorithms struggle to converge when observations are incomplete and the opponent&#8217;s policy varies across encounters. The Sun Yat-sen team identified this triple challenge—unknown pursuit strategies, information deficiency, and high-maneuverability confrontation—as the key technological bottleneck standing in the way of practical, cooperative active defense for spacecraft.</p>
<p>To overcome it, the researchers reframed the three-body engagement as a multi-agent partially observable Markov decision process, or Multi-POMDP. In this formalism, each agent receives only a noisy local observation rather than the true global state, and the opponent&#8217;s policy is treated as uncertain and potentially drawn from a set of diverse strategies. The solution architecture is a reinforcement learning guidance framework built on an adaptive dueling double deep Q-network, abbreviated AD3QN. The central idea is that a single neural network learns, from experience, how to issue coordinated maneuver commands to both the target and the defender, so that the pair adapts on the fly to whatever interception strategy the pursuer happens to be executing. The framework separates perception from decision-making: first the raw, incomplete observation history is transformed into a compact situational representation, and then that representation drives the selection of acceleration commands.</p>
<p>The perception pipeline is a fusion of two well-established neural components. Current and historical incomplete observations are stacked along the time dimension into a two-dimensional tensor, which a convolutional neural network processes to extract spatial features—the critical geometric signatures of an unfolding engagement, such as the configuration of lines of sight among the three vehicles. A gated recurrent unit then models the temporal structure of those features, using its internal update and reset gates to produce a history tensor that encapsulates the trajectory characteristics of the encounter. This history encoding is what allows the policy to infer which pursuit strategy the adversary appears to be following, something no single-frame observation can reveal. To make training on partial information effective, the team restructured the experience replay buffer so that it stores the stacked observation tensor, the action taken, the history tensor, and the resulting reward, enabling the network to exploit temporal correlations during learning rather than treating each decision as an isolated snapshot.</p>
<p>The decision-making core of AD3QN builds on the dueling double deep Q-network formulation, in which the state value function and the action advantage function are estimated separately. This decomposition reduces the variance of value estimates, a property that proves crucial in multi-strategy environments where the consequences of a maneuver differ sharply depending on the adversary&#8217;s current policy. Equally important is the reward design. Rather than rewarding success only at the final moment of the engagement, the researchers constructed a continuous reward function based on potential-field differences computed from the zero-effort miss distance—the miss distance that would result if both vehicles stopped maneuvering. Before the defender–pursuer encounter, the reward landscape encourages the defender to shrink its zero-effort miss distance relative to the pursuer, driving it toward interception. After that encounter, the shaping switches, guiding the target to enlarge its zero-effort miss distance from the pursuer and thereby accomplish evasion. The authors provide a theoretical proof that this reward formulation does not alter the optimal policy, which means the shaping improves training stability without sacrificing optimality.</p>
<p>The numerical validation compared AD3QN against a demanding field of baselines, including Deep Deterministic Policy Gradient, Twin Delayed Deep Deterministic Policy Gradient, Proximal Policy Optimization, and Deep Recurrent Q-Learning. Under the combined stresses of multi-strategy adversaries and incomplete information, these mainstream algorithms all struggled to achieve stable policy optimization, whereas AD3QN converged reliably thanks to its fusion architecture and its explicit handling of observation history. Computational efficiency is a decisive factor for flight implementation, and here the method posted a striking result: a decision frequency of 85 Hz in a simulated single-chip microprocessor environment, roughly 30 percent faster than the DRQN comparison and comfortably within the real-time requirements of spacecraft actuation systems. In representative sample engagements, the defender successfully deflected the pursuer&#8217;s trajectory by approximately 20 meters relative to the target—a deviation large enough to convert a lethal interception into a clean miss.</p>
<p>Statistical robustness was assessed through Monte Carlo analysis. Across 1,000 randomized simulations, the proposed method achieved an evasion success rate of 99.7 percent, dramatically outperforming the optimal switching cooperative guidance law at 51.8 percent and various reinforcement learning baselines, which fell below 0.2 percent. The robustness tests are perhaps the most compelling evidence of practical value: even when observation noise was inflated to 40 times the typical level—with range noise of 400 meters and line-of-sight angle noise of 40 milliradians—the method still maintained an evasion success rate of 73.5 percent. Parameter sensitivity studies added further engineering insight. When the pursuer&#8217;s maneuvering capability increased, the target&#8217;s achievable miss distance fell by roughly 30 percent, but boosting the defender&#8217;s agility effectively compensated, improving evasion performance. This trade-off quantifies a design principle for future defensive architectures: investment in the defender&#8217;s maneuverability can offset a faster adversary.</p>
<p>The broader significance of the study lies in demonstrating that cooperative, adaptive active defense is achievable under realistic sensing conditions rather than idealized perfect information. By combining a Multi-POMDP problem formulation, a CNN–GRU fusion network for perception under partial observability, dueling double Q-learning for stable value estimation in adversarial settings, and a theoretically sound potential-field reward, the Sun Yat-sen team has assembled a guidance framework that is simultaneously adaptive, robust, and computationally light enough for onboard implementation. For operators of high-value satellites in congested orbital regions, the work suggests a path toward survivability that does not depend on predicting the attacker&#8217;s playbook in advance. As orbital confrontation scenarios grow more complex, methods of this kind—able to learn coordinated counter-interception from noisy observations and to withstand sensor degradation far beyond nominal levels—are likely to become a cornerstone of spacecraft autonomy and space security engineering.</p>
<p><strong>Subject of Research:</strong> Active defense guidance for spacecraft in three-body pursuit–evasion engagements with incomplete information and multi-strategy adversaries</p>
<p><strong>Article Title:</strong> Active defense guidance for spacecraft in multi-strategy engagement with incomplete information</p>
<p><strong>Article References:</strong> Active defense guidance for spacecraft in multi-strategy engagement with incomplete information. (n.d.). <a href="https://www.eurekalert.org/news-releases/1143394" rel="noopener noreferrer">Original publication</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> Not provided</p>
<p><strong>Keywords:</strong> spacecraft guidance, active defense, pursuit-evasion, reinforcement learning, deep Q-network, partial observability, Multi-POMDP, space security, zero-effort miss distance, convolutional neural network, gated recurrent unit, spacecraft survivability</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">198232</post-id>	</item>
	</channel>
</rss>
