Extreme weather is becoming harder to plan for precisely because the events that cause the greatest damage are often the least familiar. A seawall may be designed around the strongest storm on record, yet a future storm could last longer, cover a wider area, or deliver far more rain. A power grid may survive historical heat waves but fail under a combination of record temperatures and prolonged demand. Wildfire crews may be prepared for the largest blaze documented in a region, only to face a fire that spreads through an unusual pattern of wind, dryness, and vegetation. A new machine-learning method developed by engineers at MIT is designed to help communities explore these possibilities before they happen. Called Extreme Event Aware, or “η-learning,” the approach generates statistically plausible extreme events even when the historical record contains few or no examples of comparable disasters.
Conventional risk assessments generally depend on the past. Researchers analyze observed storms, heat waves, floods, fires, or other hazards, then use statistical models and computer simulations to estimate what an event of a particular return period might look like. A “once-in-100-years” storm, for example, is typically inferred from the distribution of storms that have already been observed. That approach becomes increasingly uncertain at the far end of the distribution, where data are scarce. The most damaging event in the historical record may not represent the upper limit of what is physically possible. Climate change can further complicate the calculation by shifting the underlying conditions, making past observations a less reliable guide to future extremes.
The MIT method takes a different approach. Rather than requiring examples of the most extreme events during training, it combines information about how often certain values occur with information about how those values are arranged across space. The first type of information consists of point statistics: numerical descriptions of the probability that a variable, such as the maximum daily rainfall over a region, will reach a particular level. The second consists of spatial maps, which show how weather or other environmental conditions are distributed over an area. By learning the relationship between these two forms of information, the algorithm can generate new maps that remain consistent with the statistical behavior of the region while extending beyond the extremes directly represented in the training data.
The researchers demonstrated the system using precipitation across the continental United States. They began with 25 years of hourly rainfall maps and aggregated the observations into daily maps. From the complete record, they calculated statistics describing the frequency of different maximum-rainfall levels. However, the spatial model was trained using paired low- and high-resolution maps from only the first six months of the record. That abbreviated training period contained few, if any, examples of the most extreme rainfall events. The design created a demanding test: the algorithm had to learn the structure and geography of precipitation without simply memorizing the rarest storms.
In technical terms, the spatial component learned how broad, lower-resolution patterns correspond to detailed, high-resolution precipitation fields. A low-resolution map might indicate a large atmospheric system moving across a region, while the corresponding high-resolution map captures localized bands of intense rainfall, gaps between storm cells, and sharp variations in accumulation. The point-statistical component then constrained the generated fields so that their maximum values followed a specified extreme-value distribution. Together, the two components allow η-learning to create many possible spatial realizations of an event with a chosen rarity, such as a storm expected to occur once every 100 years.
This distinction is important because an extreme event is not defined only by a single maximum measurement. For emergency managers and infrastructure designers, the location, footprint, duration, and internal structure of a storm can be as consequential as its peak intensity. Two storms might produce the same maximum rainfall at one location but create very different risks if one remains concentrated over a city while the other spreads across an entire watershed. The new method can generate scenarios that vary these characteristics while preserving statistical plausibility. A planner could therefore examine thousands of possible storms rather than relying on one synthetic event that may accidentally overlook the most vulnerable combination of intensity and geographic coverage.
The researchers say the system could address questions that standard forecasting and simulation tools struggle to answer: What might a storm more intense than anything previously recorded look like? Where could its heaviest rainfall occur? How large an area might be affected, and how long might the event persist? Such scenarios could help cities evaluate drainage systems, reservoirs, bridges, transportation networks, and coastal defenses. The same logic could be used to explore unprecedented floods and wildfires, provided that suitable spatial observations and statistical information are available. In each case, the goal is not to predict one specific disaster, but to characterize a distribution of plausible disasters that may occupy the farthest reaches of risk.
The approach also has potential beyond environmental hazards. In robotic navigation, an autonomous system may need to reason about rare combinations of obstacles, sensor failures, or unusual movements that were absent from its training data. In financial markets, a crash can emerge from interactions among multiple sectors rather than from a single isolated variable. η-learning could be used to investigate how unusual but statistically credible combinations of conditions might produce system-wide disruptions. The method is particularly suited to problems in which extreme outcomes arise from complex interactions and where direct examples of the worst cases are too limited to support conventional data-hungry models.
MIT researchers Kai Chang and Themis Sapsis describe the work as an effort to model events that have not yet been observed but are still consistent with the known behavior of a system. The method does not claim to reveal exactly when or where the next catastrophe will occur. Instead, it provides a framework for generating a large ensemble of possible futures and assigning those scenarios a frequency or probability. That distinction could be valuable for decision-makers who must prepare for events more severe than historical experience without treating every imaginable scenario as equally likely. By filtering out implausible combinations while retaining rare, high-impact possibilities, the algorithm aims to make worst-case planning more quantitative.
As extreme events place growing pressure on energy systems, supply chains, food production, insurance markets, and public infrastructure, the ability to estimate unprecedented risk is becoming a strategic concern. A single storm, wildfire, or heat wave can trigger cascading failures across systems designed for efficiency rather than spare capacity. The MIT researchers’ open-access study, published in Nature Communications, suggests that machine learning can help close the gap between what has happened and what could plausibly happen next. If the technique proves effective across a wider range of hazards, it could give communities a new way to visualize disasters that history has not yet recorded—and to strengthen defenses before those events arrive.
Subject of Research: Machine-learning generation of statistically plausible unprecedented extreme events and worst-case environmental scenarios.
Article Title: Extreme Event Aware (η-) Learning
News Publication Date: 20 August 2026
Web References: https://www.nature.com/articles/s41467-026-76811-x
References: Nature Communications; DOI: 10.1038/s41467-026-76811-x
Keywords: Extreme weather events, artificial intelligence, machine learning, extreme-value statistics, precipitation modeling, natural disasters, climate risk, computer modeling, infrastructure resilience, η-learning

