<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>HDBSCAN &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/hdbscan/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Wed, 07 Oct 2026 14:14:24 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.3</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>HDBSCAN &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>New Open-Source Tool Turns Tangled Optimization Trade-Offs Into Clear Infrastructure Decisions</title>
		<link>https://scienmag.com/new-open-source-tool-turns-tangled-optimization-trade-offs-into-clear-infrastructure-decisions/</link>
		
		<dc:creator><![CDATA[Faith Mcneil]]></dc:creator>
		<pubDate>Wed, 07 Oct 2026 14:14:24 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[computational tools for environmental impact assessment]]></category>
		<category><![CDATA[data center infrastructure decision-making]]></category>
		<category><![CDATA[Decision Support Systems]]></category>
		<category><![CDATA[decision-making in renewable energy projects]]></category>
		<category><![CDATA[energy systems]]></category>
		<category><![CDATA[HDBSCAN]]></category>
		<category><![CDATA[infrastructure planning]]></category>
		<category><![CDATA[large-scale multi-objective optimization interpretation]]></category>
		<category><![CDATA[machine learning in decision support]]></category>
		<category><![CDATA[multi-criteria decision analysis in infrastructure development]]></category>
		<category><![CDATA[multi-objective optimization]]></category>
		<category><![CDATA[multi-objective optimization analysis]]></category>
		<category><![CDATA[offshore wind]]></category>
		<category><![CDATA[open-source infrastructure planning tools]]></category>
		<category><![CDATA[open-source software]]></category>
		<category><![CDATA[Pacific Northwest National Laboratory]]></category>
		<category><![CDATA[Pareto front]]></category>
		<category><![CDATA[Pareto front visualization]]></category>
		<category><![CDATA[pyMOODS]]></category>
		<category><![CDATA[stakeholder-driven optimization solutions]]></category>
		<category><![CDATA[UMAP]]></category>
		<category><![CDATA[visual analytics]]></category>
		<category><![CDATA[visual analytics for complex trade-offs]]></category>
		<category><![CDATA[wind farm and power grid optimization]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=244713</guid>

					<description><![CDATA[Researchers at Pacific Northwest National Laboratory have released pyMOODS, an open-source visual analytics framework that uses machine learning and rank-based analysis to help planners navigate the complex trade-offs produced by multi-objective optimization in large-scale infrastructure design.]]></description>
										<content:encoded><![CDATA[<p>When planners design a wind farm, a power grid, or a data center, they rarely face a single right answer. Instead, they confront a bewildering landscape of competing goals: cut costs, boost reliability, shrink emissions, and survive extreme weather, all at once. Optimization algorithms can churn through these conflicts and produce a Pareto front, a vast set of mathematically non-dominated solutions, but the output is often a spreadsheet of thousands of rows that no human can meaningfully interpret. A team at Pacific Northwest National Laboratory (PNNL) believes it has found a way through the fog. In a paper published in the journal SoftwareX, researchers led by Palak Mattoo and Jennifer Pham introduce pyMOODS, an open-source, machine-learning-assisted visual analytics framework designed to help decision-makers navigate the trade-offs that emerge from large-scale multi-objective optimization in infrastructure planning.</p>
<p>The core problem pyMOODS addresses is one that optimization researchers have long acknowledged but rarely solved. Multi-objective optimization does not yield one optimal design; it yields an ensemble of Pareto-optimal candidates, each representing a different compromise among objectives such as cost, generation capacity, and environmental impact. Choosing among them is a human judgment call, informed by policy goals and stakeholder priorities. Yet the tools available for this post-optimization stage are fragmented. Algorithmic toolkits like pymoo, PlatEMO, and Borg generate Pareto fronts but offer little for the decision-maker. Multi-criteria decision analysis libraries such as pymcdm and pyDecision rank alternatives once objectives are fixed, but expose no exploratory interface. Visualization tools like PAVED and Parasol support interactive exploration but stop short of preference-weighted ranking. What has been missing, the PNNL team argues, is a single interactive platform that connects the visualization layer to a complete decision-support workflow on fronts that have already been computed.</p>
<p>pyMOODS fills that gap with a decoupled software architecture that separates heavy computation from interactive exploration. The frontend is a React and MUI application rendering linked interactive plots built with React Plotly, while a Python backend handles the analytics. User actions issue asynchronous requests to a Flask REST API, which mediates all communication between the two. Because interaction state lives in the frontend while analysis runs in the backend, computation never blocks the interface. A data access module loads, validates, and caches datasets held as CSV and JSON files, and a data processing module performs transformations and filtering before invoking the analytical routines. Notably, the dashboard evolved from early Plotly Dash prototypes into the current React-based implementation, gaining dynamic rendering of parameters, plots, and objective functions, along with better responsiveness under high-dimensional workloads.</p>
<p>One of the framework&#8217;s most elegant design decisions is its declarative input schema. A problem formulation is supplied as a single JSON configuration file that links to the data files and assigns each column to one of five categories: input parameters, hyperparameters, decision variables, objective functions, and control inputs. This means a single deployment can serve formulations whose objectives, decision variables, and hyperparameters differ in number and meaning, without any code changes. That flexibility matters because the framework targets real infrastructure applications, such as offshore wind co-design, MTDC-ESS system design, and load-shedding-constrained network expansion, where problem structures vary dramatically across use cases. It is a practical requirement that benchmark-oriented tools, typically evaluated on synthetic test problems like ZDT and DTLZ, simply do not address.</p>
<p>The analytical heart of pyMOODS combines two very different classes of methods. For structuring the solution space, the framework turns to unsupervised machine learning: a Uniform Manifold Approximation and Projection (UMAP) embedding projects the high-dimensional objective and decision space into two dimensions, and density-based HDBSCAN clustering groups similar solutions within that latent space. This gives users a navigable map of the entire front, where clusters of comparable candidates are visible at a glance and outliers stand out immediately. Users can also plot any pair of objectives or decision variables directly, or restrict the working set to solutions sharing a chosen hyperparameter configuration, such as a particular technology option.</p>
<p>For trade-off interpretation, pyMOODS takes a deliberately different path: a deterministic, parameter-free rank analysis. Every solution is ranked on each objective, and a solution&#8217;s generalizability is defined as its worst rank across all objectives. A solution is said to specialize in an objective when its rank there is strictly better than that of every more generalizable solution. Because this characterization derives entirely from the rank matrix, it introduces no trained model, no tuned threshold, and no random seed; identical inputs always produce identical generalizer and specializer assignments. This answers a question that ranking alone cannot: not merely which solutions score well under a given weighting, but which specific compromise each one represents. Solutions can additionally be ordered by an aggregate score reflecting user-specified objective weights, giving stakeholders a familiar mechanism for expressing their priorities, and the ranking is computed from the stored front rather than by re-solving, so stakeholders with different priorities can interrogate the same results interactively.</p>
<p>The framework&#8217;s workflow comes alive in its demonstration on the MoCoDo dataset for offshore wind farm planning with battery energy storage. The problem involves six objectives, including battery cost, cable material cost, day-ahead revenue, real-time revenue, and two reserve revenue streams, along with two decision variables: cable capacity and battery rated power. A conventional approach would aggregate all revenue streams into one total and all costs into another, reducing the problem to a simple two-way trade-off, but that aggregation obscures how individual solutions perform across specific revenue mechanisms. pyMOODS instead guides the user through four linked views: a scatter plot offering a first overview of the solution space, a ranking table for selecting a candidate, a parallel coordinates plot distinguishing generalizers from specializers across all six objectives, and beeswarm plots showing the distribution of all candidates per objective and decision variable. Selection propagates across views, so a candidate can be examined from several perspectives without losing context.</p>
<p>Performance testing suggests the architecture delivers on its responsiveness promise. On a local deployment running on an Apple M2 Pro with 16 GB of RAM, the team profiled three representative user tasks across two built-in formulations, Cameo_datacenter and MoCoDo_v2, repeating each measurement three times. Formulation initialization took roughly five milliseconds in both cases, while full dashboard load sequences averaged 3.838 seconds for the larger formulation and 0.475 seconds for the smaller one, with filter and selection interactions showing nearly identical latencies. The release also includes automated frontend tests using Playwright, which verify that dashboard components load correctly, the solution table sorts properly, and the generalizer-specializer controls update results, using simulated backend responses to validate the interface without requiring the full backend to run.</p>
<p>A capability comparison against representative open-source tools underscores what makes pyMOODS distinctive. Front-generation frameworks and metaheuristic platforms are not designed to be driven by a decision-maker; multi-criteria decision analysis libraries rank alternatives but expose no exploratory interface; and interactive visualization tools support exploration but stop short of preference-weighted ranking. In the published comparison, pyMOODS is the only tool that simultaneously offers an interactive user interface, operates on precomputed fronts from any solver, supports user-weighted ranking, provides automatic trade-off characterization, and accepts schema-agnostic problem formulations. The authors are candid about the release&#8217;s limitations, however: performance has been characterized on only two formulations on a single machine, rank-and-weight aggregation cannot recover non-convex regions of a Pareto front, and no usability evaluation has yet been conducted. Planned features include additional MCDM methods for comparison and an AI chatbot assistant.</p>
<p>The implications reach well beyond software engineering. As power systems worldwide absorb rising penetrations of onsite energy sources and face intensifying extreme weather, the co-optimization of generation, transmission, and storage investments is becoming one of the defining computational challenges of the energy transition. Tools like pyMOODS aim to ensure that the enormous computational effort poured into optimizing these systems does not end in an unreadable list of trade-offs. Released under the BSD 3-Clause License and archived on GitHub and Zenodo, version v0.0.3 puts a complete post-optimization decision-support pipeline into the hands of the planners, engineers, and researchers who must ultimately decide which compromises our future infrastructure will embody. Whether it changes real-world decisions, the authors acknowledge, is a question that future work, and future users, will have to answer.</p>
<p><strong>Subject of Research:</strong> An open-source visual analytics framework for multi-objective decision support in large-scale infrastructure planning</p>
<p><strong>Article Title:</strong> pyMOODS: An open-source framework for multi-objective decision support in large-scale infrastructure planning</p>
<p><strong>Article References:</strong> Mattoo, P., Pham, J. N., Jain, M., Arendt, D., Avila, P., Ramachandran, T., Adetola, V., Yun, J. Y., &amp; Wenskovitch, J. (2026). pyMOODS: An open-source framework for multi-objective decision support in large-scale infrastructure planning. <em>SoftwareX, 36</em>, Article 103095. <a href="https://doi.org/10.1016/j.softx.2026.103095" rel="noopener noreferrer">https://doi.org/10.1016/j.softx.2026.103095</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1016/j.softx.2026.103095" rel="noopener noreferrer">10.1016/j.softx.2026.103095</a></p>
<p><strong>Keywords:</strong> pyMOODS, multi-objective optimization, Pareto front, visual analytics, infrastructure planning, offshore wind, UMAP, HDBSCAN, decision support systems, open-source software, energy systems, Pacific Northwest National Laboratory</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">244713</post-id>	</item>
		<item>
		<title>Shooting Data Joins a Centralized Platform for Basketball Analytics in Spain</title>
		<link>https://scienmag.com/shooting-data-joins-a-centralized-platform-for-basketball-analytics-in-spain/</link>
		
		<dc:creator><![CDATA[Denise Maddox]]></dc:creator>
		<pubDate>Sat, 26 Sep 2026 01:07:05 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[ACB]]></category>
		<category><![CDATA[advanced basketball statistics]]></category>
		<category><![CDATA[basketball analytics]]></category>
		<category><![CDATA[basketball data visualization]]></category>
		<category><![CDATA[basketball game data collection]]></category>
		<category><![CDATA[centralized sports data platform]]></category>
		<category><![CDATA[data visualization]]></category>
		<category><![CDATA[DBSCAN]]></category>
		<category><![CDATA[HDBSCAN]]></category>
		<category><![CDATA[K-means clustering]]></category>
		<category><![CDATA[player shooting performance metrics]]></category>
		<category><![CDATA[R package]]></category>
		<category><![CDATA[Shiny web application]]></category>
		<category><![CDATA[shooting data]]></category>
		<category><![CDATA[shooting data in professional basketball]]></category>
		<category><![CDATA[shot charts]]></category>
		<category><![CDATA[Spanish ACB league data management]]></category>
		<category><![CDATA[spatial basketball data]]></category>
		<category><![CDATA[sport analytics]]></category>
		<category><![CDATA[sports data integration]]></category>
		<category><![CDATA[sports performance analysis]]></category>
		<category><![CDATA[sports technology innovation]]></category>
		<category><![CDATA[web scraping]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=215831</guid>

					<description><![CDATA[A new study completes a three-part project by adding spatial shooting data from Spain's ACB league to a centralized basketball analytics platform, complete with an R package and interactive web application.]]></description>
										<content:encoded><![CDATA[<p>Basketball has become one of the most data-hungry sports on the planet, and a new study published in Multimedia Tools and Applications shows how the Spanish professional league is catching up with the analytics revolution. Guillermo Vinué, an independent researcher based in Valencia, has completed the final piece of a three-part research project devoted to improving basketball data management in Spain. The work focuses on the ACB, the country&#8217;s top male professional competition, and tackles a source of information that had remained untouched in his earlier papers: shooting data. By capturing exactly where every shot is taken, what type of shot it is, and whether it goes in, the research opens the door to a far richer picture of player performance than traditional box score statistics can provide.</p>
<p>The new paper is the third and last part of what the author describes as a research triptych. The first two installments dealt with box score data and play-by-play data respectively, each feeding into a growing analytical infrastructure for Spanish basketball. Box scores capture cumulative statistics such as points, rebounds, and assists, while play-by-play records describe the sequence of events during a game. Shooting data, however, adds a crucial spatial dimension. For each attempt, the dataset records the location of the shot on the court, the shot type, and the outcome. This is precisely the kind of information that has transformed how analysts, coaches, and fans understand the modern game, from the rise of the three-point shot to the strategic value of shots near the rim.</p>
<p>To build the dataset, Vinué developed code in the R programming language that automatically collects shooting information from the official ACB website using a web scraping procedure. Web scraping involves programmatically extracting data from web pages, and the author designed the procedure to be friendly and reproducible. The approach draws on established R tools for harvesting web content, including packages such as rvest and RSelenium, and respects the target website&#8217;s robots.txt file, which specifies rules for automated access. Once collected, the raw data required ordinary cleaning to be structured into a standard format suitable for analysis. The author then validated the accuracy of the scraped data by plotting the shots and comparing their coordinates and outcomes against the shot charts displayed on the source website, ensuring that the automated pipeline faithfully reproduces the original information.</p>
<p>The scale of the resulting dataset is substantial. It covers the 2024-2025 ACB season and includes 39,550 shots drawn from 306 games and 293 players. Each record contains, among other details, the location of the attempted shot, the shot type, and whether the attempt was successful. A dataset of this size allows for detailed shooting profiles of individual players, showing where on the floor they prefer to shoot and how effective they are from each area. Such profiles go far beyond a simple field goal percentage, which collapses all attempts into a single number regardless of where they were taken. Spatially aware statistics have become central to basketball analytics worldwide, and this work brings that capability to Spanish league data in a systematic, reproducible way.</p>
<p>A key contribution of the study is the integration of shooting data into a previously developed data platform, so that all analyses are centralized in a single place. The platform, described in an earlier paper, already handled box score and play-by-play information; adding shooting data completes the triptych and turns the platform into a comprehensive resource for Spanish professional basketball. Alongside the platform, the author updated BAwiR, an accompanying R package that is freely available from CRAN, the comprehensive archive network for R software. The package includes documentation for all its files as well as three vignettes that guide users through its functionality. Making both the data and the code openly available means that other researchers, analysts, and enthusiasts can reproduce the results and extend the analysis to their own questions.</p>
<p>To make the data accessible to non-programmers, the study also provides a web application for easy interaction. Built with Shiny, a popular framework for turning R code into interactive web tools, the application lets users explore shooting data without writing any code themselves. Interactive visualization has become an essential bridge between raw data and decision-makers in sport, allowing coaches and analysts to query specific players, games, or shot types on demand. The author also gathered feedback from users through a questionnaire, reflecting a growing recognition in the visualization research community that sports tools should be designed in consultation with the people who will actually use them.</p>
<p>One of the more technically interesting parts of the work concerns how to divide the basketball court into meaningful zones. Classical zones such as the paint, the mid-range areas, and the corners are usually defined manually, but the paper explores whether automatic clustering algorithms can do the job instead. Three methods were tested on the shooting dataset: DBSCAN, HDBSCAN, and k-means. DBSCAN, which stands for Density-Based Spatial Clustering of Applications with Noise, groups points that lie close together in space and marks isolated points as noise, but it requires the user to specify a neighborhood radius and a minimum number of points in advance. HDBSCAN extends this idea by building a hierarchy of clusters of varying density, reducing the need to tune the radius parameter. K-means, the simplest of the three, iteratively assigns each shot to its nearest cluster center and then recalculates those centers as the average of the assigned points, with the number of clusters chosen beforehand, often through silhouette analysis.</p>
<p>The results of the clustering comparison were instructive. K-means produced the closest representation of the classical basketball regions, identifying six zones, although it did not distinguish between two-point and three-point shots, and some regions contained both. HDBSCAN roughly identified some areas, such as the paint and the left and right three-point zones, but its other regions were messy. DBSCAN fared worst, returning just two large regions that carry little basketball meaning. When the analysis was refined by filtering first by shot type, k-means divided the two-point area into three regions approximating the paint and the left and right mid-ranges, while the three-point area split into six clearly identifiable regions: left and right corners, left and right elbows, center three-point shots, and long-distance three-point attempts. These regions largely coincide with the manual partition proposed by the author, and the clustering even suggested numerical limits for the court zones, such as paint boundaries at y-coordinates of around plus or minus 2000. The conclusion is that k-means can serve as a useful starting point for determining the coordinates needed to build court regions manually.</p>
<p>The study is candid about the limitations of automated data collection. The scraping code is reproducible with similar shot data, but if the target website changes its structure, manual updates and further cleaning may be required. This is a familiar trade-off in sports analytics, where researchers often depend on websites they do not control. The author acknowledges the official ACB website for making the data accessible, and the paper&#8217;s appendices provide a step-by-step guide to the data collection strategy, along with a glossary of basketball terms to help readers less familiar with the sport&#8217;s vocabulary. Such transparency supports reproducibility, a value increasingly emphasized in the R sports analytics community, where packages for basketball, football, and other sports are systematically reviewed and shared.</p>
<p>The broader significance of the work lies in democratizing access to professional basketball data outside the NBA. While the NBA offers well-documented application programming interfaces that have spawned a rich ecosystem of analytical tools, European leagues have historically lagged in data availability. Research has shown that analytics investment can influence team performance in professional basketball, and systematic reviews have highlighted both the potential and the uneven adoption of analytical techniques in the sport. By centralizing box score, play-by-play, and shooting data for the ACB in one platform, backed by a free R package and an interactive web application, this research gives Spanish basketball a data infrastructure comparable in spirit to what NBA analysts enjoy. From shot charts and spatial clustering to player shooting profiles, the tools are now in place for coaches, journalists, and fans to interrogate the game at a level of detail that was previously out of reach for Spain&#8217;s top league.</p>
<p><strong>Subject of Research:</strong> Web scraping and visualization of shooting data from the Spanish ACB basketball league</p>
<p><strong>Article Title:</strong> Adding shooting data to a centralized platform for basketball data visualization</p>
<p><strong>Article References:</strong> Adding shooting data to a centralized platform for basketball data visualization. (n.d.). <a href="https://doi.org/10.1007/s11042-026-21930-2" rel="noopener noreferrer">https://doi.org/10.1007/s11042-026-21930-2</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s11042-026-21930-2" rel="noopener noreferrer">10.1007/s11042-026-21930-2</a></p>
<p><strong>Keywords:</strong> basketball analytics, ACB, shooting data, web scraping, R package, data visualization, Shiny web application, k-means clustering, DBSCAN, HDBSCAN, shot charts, sport analytics</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">215831</post-id>	</item>
		<item>
		<title>Machine Learning Maps Hidden Orbit Structures in Four-Dimensional Space</title>
		<link>https://scienmag.com/machine-learning-maps-hidden-orbit-structures-in-four-dimensional-space/</link>
		
		<dc:creator><![CDATA[Grant Pearson]]></dc:creator>
		<pubDate>Sun, 20 Sep 2026 22:18:04 +0000</pubDate>
				<category><![CDATA[Space]]></category>
		<category><![CDATA[astrodynamics]]></category>
		<category><![CDATA[celestial mechanics]]></category>
		<category><![CDATA[chaos]]></category>
		<category><![CDATA[chaos and stability in celestial orbits]]></category>
		<category><![CDATA[circular restricted three-body problem]]></category>
		<category><![CDATA[clustering]]></category>
		<category><![CDATA[data-driven space trajectory prediction]]></category>
		<category><![CDATA[dynamical systems analysis]]></category>
		<category><![CDATA[Earth–Moon system]]></category>
		<category><![CDATA[Fast Lyapunov Indicator]]></category>
		<category><![CDATA[four-dimensional Poincaré maps]]></category>
		<category><![CDATA[HDBSCAN]]></category>
		<category><![CDATA[hidden orbit structures]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[machine learning in space navigation]]></category>
		<category><![CDATA[Poincaré map]]></category>
		<category><![CDATA[Principal Component Analysis]]></category>
		<category><![CDATA[quasi-periodic orbits]]></category>
		<category><![CDATA[space mission route optimization]]></category>
		<category><![CDATA[three-body problem]]></category>
		<category><![CDATA[trajectory classification in astrodynamics]]></category>
		<category><![CDATA[trajectory design]]></category>
		<category><![CDATA[unsupervised clustering in dynamical systems]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=203384</guid>

					<description><![CDATA[Researchers at the Air Force Institute of Technology have developed an unsupervised machine learning pipeline that automatically discovers and classifies dynamical structures in four-dimensional Poincaré maps of the circular restricted three-body problem.]]></description>
										<content:encoded><![CDATA[<p>For more than a century, the three-body problem has stood as one of the most famously intractable puzzles in celestial mechanics. When a spacecraft drifts through the gravitational landscape of two large bodies, such as the Earth and the Moon, its path can bend, loop, and scatter in ways that defy simple prediction. Now, a team of researchers at the Air Force Institute of Technology has unveiled a machine learning pipeline that can automatically classify the hidden architecture of these trajectories, potentially transforming how mission designers chart routes through cislunar space and beyond. The study, published in Astrophysics and Space Science, applies unsupervised clustering to four-dimensional Poincaré maps in the circular restricted three-body problem, or CR3BP, and demonstrates that algorithms can recover meaningful dynamical structures from data that would overwhelm even the most patient human analyst.</p>
<p>The CR3BP reduces the gravitational dance of three bodies to its essential form: an infinitesimally small mass, such as a spacecraft, moving under the influence of two primaries that orbit their shared center of mass in perfect circles. Within this simplified but still chaotic system, trajectories fall into three fundamental categories. Periodic orbits repeat themselves exactly, tracing closed loops through phase space. Quasi-periodic orbits wander across the surface of an invisible torus, never quite repeating but never straying far from home. Chaotic trajectories, by contrast, diverge unpredictably from their initial conditions, sensitive to the smallest perturbation. Distinguishing among these behaviors is the central task of trajectory design, because a mission planner must know whether a candidate path will remain stable or spiral into chaos.</p>
<p>The traditional tool for this task is the Poincaré map, a technique dating back to Henri Poincaré&#8217;s foundational work in 1890. Rather than tracking a trajectory continuously, the map records only the points where the path pierces a chosen surface in phase space, converting a continuous curve into a discrete scatter of crossings. In the planar version of the problem, this produces a two-dimensional plot in which periodic orbits appear as fixed points, quasi-periodic orbits form distinctive chains of islands, and chaotic motion scatters into structureless dust. Mission designers have long relied on visual inspection of these maps to identify promising orbits, a process that works well enough in two dimensions but collapses entirely when the problem extends into three spatial dimensions.</p>
<p>Extending the Poincaré map to four dimensions is necessary because real spacecraft do not confine themselves to planes. To capture out-of-plane motion, researchers use a space-plus-color representation: two spatial coordinates occupy the horizontal axes, the vertical coordinate is plotted on a third axis, and the out-of-plane velocity is encoded as color. The result is a rich but visually cluttered dataset in which the familiar island chains of two-dimensional maps vanish, replaced by three-dimensional structures with no established taxonomy. Previous efforts to sift through these maps relied on filtering by the Fast Lyapunov Indicator, a numerical measure of chaos, but this remained a manual, time-intensive operation that yielded only preliminary catalogs of structures. With each map containing thousands to millions of data points, the need for automation became unmistakable.</p>
<p>The new pipeline, developed by Kevin M. Trigg, Daniel J. Broyles, Robert A. Bettinger, and Tyler J. Kapolka, addresses this challenge through a carefully engineered sequence of computational steps. The researchers simulated 10,000 seed trajectories for each of 11 different Jacobi constants, a parameter that acts like an energy level and determines which regions of phase space a spacecraft can reach. Each trajectory was propagated for 1,000 time units using high-precision numerical integration, with early termination for paths that collided with a primary body or escaped the system. The team then faced a fundamental data problem: trajectories produce variable numbers of Poincaré crossings, and clustering algorithms require inputs of fixed length. Their solution was a feature engineering scheme that compresses each trajectory into a 23-dimensional numerical fingerprint.</p>
<p>That fingerprint draws on three complementary strategies. Statistical and dynamical descriptors capture the bounding box of each trajectory&#8217;s crossings across position and velocity coordinates, along with the Fast Lyapunov Indicator, which quantifies sensitivity to initial conditions and separates regular from chaotic behavior. Geometric descriptors treat each trajectory&#8217;s crossings as a miniature dataset, applying clustering within the trajectory itself to measure how cohesive or fragmented its structure is; a well-formed invariant torus yields dense, orderly crossings, while chaos produces scattered noise. Finally, frequency-domain analysis applies a Fast Fourier Transform to the ordered sequence of crossings, extracting the three strongest oscillatory modes and their amplitudes as signatures of recurrence. Periodic and quasi-periodic trajectories concentrate their spectral energy in sharp peaks, whereas chaotic trajectories spread it broadly.</p>
<p>With feature vectors in hand, the pipeline applies Principal Component Analysis to compress the representation, retaining components that explain at least 95 percent of the variance and typically reducing 23 features to 11. The reduced vectors then feed into HDBSCAN, a hierarchical density-based clustering algorithm chosen for its ability to handle clusters of varying density and shape without requiring the number of clusters to be specified in advance. This flexibility matters because Poincaré maps contain structures with irregular boundaries and non-uniform density, conditions that defeat simpler methods like DBSCAN, which relies on a single global density threshold. HDBSCAN also explicitly labels low-density trajectories as noise, providing a natural mechanism for isolating chaotic orbits that belong to no coherent structure. The researchers compared HDBSCAN against agglomerative clustering, spectral clustering, affinity propagation, and Gaussian mixture models, finding that while competitors achieved slightly higher Silhouette scores, HDBSCAN consistently produced lower Davies-Bouldin scores, better structural similarity, and the crucial advantage of automatic noise detection.</p>
<p>The analysis revealed eight distinct dynamical structures in the four-dimensional maps, which the authors named descriptively in the absence of any formal taxonomy: Figure-Eight, Tube, Scorpion Tail with Dots, Scorpion Tail without Dots, Pillar and Shield, Pillar and Shapes, Two Planes, and Three Pillars. These geometries bear no resemblance to the island chains of planar maps, underscoring how radically the topology changes when out-of-plane motion is included. The Figure-Eight structures, associated with quasi-periodic motion around Lagrange points, dominated the regular regions of the maps. An ablation study confirmed that every feature group contributed essential information: removing the Fast Lyapunov Indicator degraded the separation between regular and chaotic regions, while removing statistical features collapsed intra-cluster cohesion. When the tuned clustering configuration was applied unchanged to maps at other Jacobi constants, it recovered consistent structures with noise proportions holding steady between roughly 20 and 30 percent, demonstrating genuine robustness across the energy landscape.</p>
<p>Perhaps the most ambitious component of the work is its approach to orbit continuation. In classical astrodynamics, tracing how an orbit family evolves requires differential correction and continuation methods that follow a trajectory as parameters change. The researchers instead built a data-driven approximation: they treated each Jacobi constant level as a layer in a directed graph, connected each trajectory to its three nearest neighbors in feature space within the adjacent layer, and used depth-first search to find chains spanning the entire range from C equals 3.18 down to 2.68. Chains were ranked by average feature distance, with the smoothest paths representing candidate continuations of dynamical families. The results confirmed that the learned feature representations preserve meaningful dynamical similarity across energy levels, with Figure-Eight structures persisting coherently across maps. Yet the method also exposed its own limitation: only 24 unique structures appeared among the top 1,000 chains, revealing a strong bias toward the dominant, tightly clustered geometries and a need for additional constraints to encourage exploration of rarer structures.</p>
<p>The implications extend beyond the Earth-Moon system. The authors emphasize that the methodology applies to any multi-body gravitational environment, and that four-dimensional Poincaré maps represent a growing frontier in astrodynamics research. Future work includes cataloging how these structures evolve with mass parameter and energy, comparing the discovered orbits against known periodic orbit families and invariant tori generated by traditional continuation methods, and developing algorithms to locate fixed points in maps where, unlike the planar case, they are not intuitively positioned at the centers of island chains. The pipeline also contributes to a broader scientific goal: automated detection of invariant structures in high-dimensional Hamiltonian systems, a challenge that reaches into plasma physics, celestial mechanics, and accelerator design. For mission planners navigating the increasingly crowded cislunar arena, where spacecraft such as those supporting lunar exploration must exploit subtle gravitational structures to conserve fuel, a tool that converts millions of trajectory crossings into organized, labeled dynamical families could shorten the path from concept to flight-ready trajectory design.</p>
<p><strong>Subject of Research:</strong> Unsupervised machine learning for discovering dynamical structures in 4D Poincaré maps of the circular restricted three-body problem</p>
<p><strong>Article Title:</strong> Application of machine learning to discover dynamical structures in 4D Poincaré maps in the circular restricted three-body problem</p>
<p><strong>Article References:</strong> Trigg, K. M., Broyles, D. J., Bettinger, R. A., &amp; Kapolka, T. J. (2026). Application of machine learning to discover dynamical structures in 4D Poincaré maps in the circular restricted three-body problem. <em>Astrophysics and Space Science, 371</em>(9), Article 107. <a href="https://doi.org/10.1007/s10509-026-04638-5" rel="noopener noreferrer">https://doi.org/10.1007/s10509-026-04638-5</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s10509-026-04638-5" rel="noopener noreferrer">10.1007/s10509-026-04638-5</a></p>
<p><strong>Keywords:</strong> Poincaré map, circular restricted three-body problem, machine learning, HDBSCAN, quasi-periodic orbits, chaos, Fast Lyapunov Indicator, principal component analysis, astrodynamics, Earth-Moon system, trajectory design, clustering</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">203384</post-id>	</item>
	</channel>
</rss>
