<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>open data &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/open-data/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Sat, 10 Oct 2026 11:47:26 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.3</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>open data &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>New Global Database Puts Every Definition of Marine Heatwave on One Map</title>
		<link>https://scienmag.com/new-global-database-puts-every-definition-of-marine-heatwave-on-one-map/</link>
		
		<dc:creator><![CDATA[Violet Maxwell]]></dc:creator>
		<pubDate>Sat, 10 Oct 2026 11:47:26 +0000</pubDate>
				<category><![CDATA[Earth Science]]></category>
		<category><![CDATA[climate change]]></category>
		<category><![CDATA[climate change impact on oceans]]></category>
		<category><![CDATA[climatology]]></category>
		<category><![CDATA[cross-comparison of heatwave criteria]]></category>
		<category><![CDATA[detrending]]></category>
		<category><![CDATA[earth system science data]]></category>
		<category><![CDATA[Earth System Science Data publication]]></category>
		<category><![CDATA[effects of marine heatwaves on marine ecosystems]]></category>
		<category><![CDATA[ESA Climate Change Initiative]]></category>
		<category><![CDATA[global marine heatwave definitions]]></category>
		<category><![CDATA[marine cold spells]]></category>
		<category><![CDATA[marine ecology]]></category>
		<category><![CDATA[marine heatwave analysis]]></category>
		<category><![CDATA[marine heatwave database]]></category>
		<category><![CDATA[marine heatwave duration and intensity]]></category>
		<category><![CDATA[marine heatwave research methods]]></category>
		<category><![CDATA[Marine Heatwaves]]></category>
		<category><![CDATA[ocean extremes]]></category>
		<category><![CDATA[ocean temperature anomalies]]></category>
		<category><![CDATA[open data]]></category>
		<category><![CDATA[satellite data]]></category>
		<category><![CDATA[satellite sea surface temperature observations]]></category>
		<category><![CDATA[sea surface temperature]]></category>
		<category><![CDATA[standardized marine heatwave metrics]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=258466</guid>

					<description><![CDATA[A new 43-year global database applies ten different marine heatwave definitions in parallel, revealing how methodological choices reshape our picture of ocean extremes.]]></description>
										<content:encoded><![CDATA[<p>The ocean is getting hotter, and the way scientists count those hot spells has become a scientific problem in its own right. A team led by the Danish Meteorological Institute, working with partners across Europe and India, has now released a global database that tackles the issue head-on. Called MHW-MAD, short for Marine Heat Waves and Cold Spells – Multiple Analysis/Definitions, the dataset covers more than four decades of satellite sea surface temperature observations from 1982 to 2024 and applies ten different definitions of what counts as a marine heatwave, all in parallel. Published in the journal Earth System Science Data, the resource is designed to settle a question that has quietly plagued the field for years: how much of the disagreement between marine heatwave studies is real, and how much is simply an artefact of the definitions researchers choose.</p>
<p>Marine heatwaves are discrete, prolonged periods of abnormally warm water relative to local conditions, and their consequences can be dramatic. The most commonly cited definition, established by Alistair Hobday and colleagues in 2016, identifies a heatwave wherever sea surface temperature exceeds the local 90th percentile for five or more consecutive days. By that measure, events can persist for months and stretch across regions larger than a thousand square kilometres. The ecological stakes are high. An intense heatwave off Western Australia in 2011 stripped roughly 100 kilometres of kelp forest, about 90 percent of the region&#8217;s kelp, and triggered a lasting regime shift in the local ecosystem. Coral bleaching, harmful algal blooms, mass mortalities, shifting species distributions and collapsing fisheries have all been linked to these events, and a 2013 to 2016 warm anomaly in the Northeast Pacific, known simply as the Blob, caused widespread seabird and marine mammal deaths.</p>
<p>The evidence that heatwaves are worsening is unambiguous. Between 1925 and 2016, the average annual number of marine heatwave days worldwide increased by more than 50 percent, and many recent events have been attributed to anthropogenic warming. Models project further increases in frequency by the end of the century. Yet the field lacks a single agreed way to detect these events. Some studies anchor their statistics to a fixed historical baseline, such as the World Meteorological Organization&#8217;s standard 30-year climate normal of 1991 to 2020, while others use a shifting baseline that moves forward with each year. In a warming ocean, a fixed baseline will flag more recent warm spells as extreme, whereas a moving baseline filters out the long-term trend and highlights year-to-year variability instead. Threshold percentiles, minimum durations and detrending choices add further layers of divergence.</p>
<p>MHW-MAD addresses this fragmentation by computing everything at once. The underlying sea surface temperature record comes from the European Space Agency&#8217;s Sea Surface Temperature Climate Change Initiative, version 3.0, which merges infrared and microwave retrievals from 22 different satellite missions into a daily, gap-free global grid at 0.05 degrees resolution, roughly five kilometres. The product is generated with the climate configuration of the OSTIA analysis system, standardises temperatures to a depth of about 20 centimetres, and deliberately avoids assimilating in situ measurements so that satellite-derived trends remain intact. An interim climate data record extends the series from 2022 through 2024 using the same software, ensuring a consistent time series across the full 43-year span.</p>
<p>The methodological machinery behind the database is detailed and reproducible. For every grid cell and every day of the year, the team computed climatological distributions of temperature, including the mean and the 1st, 5th, 10th, 50th, 90th, 95th and 99th percentiles. Two baseline approaches were implemented: a fixed 30-year WMO baseline of 1991 to 2020, and a rolling 30-year window that advances one year at a time from 1982 to 2011 up to 1995 to 2024. Following Hobday&#8217;s original two-stage smoothing, percentiles were estimated from an 11-day window centred on each day of the year and then smoothed with a 31-day moving average, which suppresses sampling noise and prevents abrupt day-to-day jumps in the thresholds. A 366-day reference calendar, with 29 February interpolated in non-leap years, keeps anomalies correctly aligned across the record.</p>
<p>Detrending is one of the database&#8217;s most consequential features. Because long-term warming raises the baseline temperature, a static climatology will produce ever more frequent heatwave detections in later years, conflating climate change with natural variability. To separate the two, the team performed an ordinary least squares linear regression of temperature against year for each calendar date at each grid cell, pooling all values for that date across 1982 to 2024, and subtracted the fitted trend. The result is a parallel, trend-free version of the record in which detected heatwaves can be interpreted as events driven by variability rather than by the slowly shifting mean. Detrending was applied at every grid cell regardless of statistical significance to keep the database gap-free, though only for the standard Hobday definition.</p>
<p>Beyond the standard 90th-percentile, five-day criterion, the database offers stricter thresholds at the 95th and 99th percentiles and longer persistence requirements of 10 and 30 consecutive days, which isolate only the most durable basin-scale episodes. Each detected event is also assigned a categorical severity index following the Hobday 2018 scheme: category 1 for moderate anomalies just above the threshold, category 2 for strong, category 3 for severe, and category 4 for extreme events where the anomaly exceeds four times the local threshold difference. The same framework is applied symmetrically to the cold tail of the distribution, producing marine cold spell indices based on the 1st, 5th and 10th percentiles, so that warm and cold extremes can be analysed within a single consistent system.</p>
<p>The team tested how these choices matter using a case study of the Blob on 1 January 2014. The event proved robust across all definitions, but its apparent size and intensity shifted considerably. A fixed WMO baseline flagged more area as heatwave, and at higher severity, than warmer later climatologies. Raising the threshold from the 90th to the 95th or 99th percentile progressively screened out moderate anomalies and shrank the affected extent. Globally, requiring 10 consecutive days instead of five reduced heatwave coverage by only about 0.5 percent, while a 30-day minimum cut it by roughly 4 percent. Detrending removed about 1.2 percent of heatwave area worldwide, yet had little effect on the Blob itself, a finding consistent with earlier work attributing that event to natural variability rather than long-term warming. Switching between the WMO and reanalysis baselines changed global coverage by just 0.4 percent, though the Blob appeared more severe under the WMO baseline.</p>
<p>All outputs are delivered as CF-compliant NetCDF files through an ERDDAP server hosted by the Danish Meteorological Institute, with no registration required. Users can download entire files or extract subsets by time, latitude, longitude and variable through a point-and-click web interface, direct URLs, or the rerddap and erddapy packages for R and Python. Three file types are provided: daily climatology files containing the seasonal mean and percentile thresholds, daily anomaly maps for both raw and detrended inputs, and daily category files encoding severity from zero to four. The dataset is also archived on Zenodo with a persistent identifier for citation, and everything is released under a Creative Commons Attribution 4.0 licence.</p>
<p>The practical implications reach well beyond academic bookkeeping. Ecologists can now test whether an impact such as coral bleaching correlates only with the longest, most intense heatwaves, which would argue for management focused on multi-week thermal stress rather than short-lived spikes. Attribution studies can use the detrended anomalies to separate the fingerprint of warming from natural fluctuations. The authors caution that the database covers surface waters only, since subsurface heatwaves may not align with surface events, and that polar regions require targeted products because sea-ice-covered areas carry artificial freezing-point values that would otherwise register as heatwave hot spots. They plan annual updates, additional sea surface temperature products, non-linear detrending and a Python package for users to build their own indicators. By embracing the full spectrum of definitions rather than picking one, MHW-MAD turns a source of confusion into a lens, letting researchers see exactly how much of what we think we know about ocean extremes depends on where we draw the line.</p>
<p><strong>Subject of Research:</strong> A multi-definition global database of marine heatwaves and cold spells derived from satellite sea surface temperature data</p>
<p><strong>Article Title:</strong> Marine Heat Waves and Cold Spells – Multiple Analysis/Definitions (MHW-MAD): A Multi-Definition Global Marine Heatwave Database from Satellite Sea Surface Temperature Data</p>
<p><strong>Article References:</strong> Hayward, A., Dasgupta, N., McAdam, R., Payne, M. R., Raj, R. P., Bonino, G., Chatterjee, S., Combes, V., Denaxa, D., De Rovere, F., Englyst, P., Haapaniemi, V., Hargous, P., Høyer, J., Joseph, K. A., Lopes, B., Oliveira, A., Paixão, J., Silva, F., &#8230; Olsen, S. M. (2026). Marine Heat Waves and Cold Spells – Multiple Analysis/Definitions (MHW-MAD): A Multi-Definition Global Marine Heatwave Database from Satellite Sea Surface Temperature Data. <em>Earth System Science Data, 18</em>(10), 7181-7197. <a href="https://doi.org/10.5194/essd-18-7181-2026" rel="noopener noreferrer">https://doi.org/10.5194/essd-18-7181-2026</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.5194/essd-18-7181-2026" rel="noopener noreferrer">10.5194/essd-18-7181-2026</a></p>
<p><strong>Keywords:</strong> marine heatwaves, sea surface temperature, satellite data, climate change, ocean extremes, marine cold spells, ESA Climate Change Initiative, detrending, climatology, marine ecology, Earth System Science Data, open data</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">258466</post-id>	</item>
		<item>
		<title>World&#8217;s Highest Weather Stations Reveal Everest&#8217;s Extreme Climate From Valley to Summit</title>
		<link>https://scienmag.com/worlds-highest-weather-stations-reveal-everests-extreme-climate-from-valley-to-summit/</link>
		
		<dc:creator><![CDATA[Violet Maxwell]]></dc:creator>
		<pubDate>Sat, 10 Oct 2026 05:44:23 +0000</pubDate>
				<category><![CDATA[Earth Science]]></category>
		<category><![CDATA[atmospheric measurements at Everest summit]]></category>
		<category><![CDATA[automatic weather stations]]></category>
		<category><![CDATA[automatic weather stations in high mountains]]></category>
		<category><![CDATA[climate change insights from Himalayan weather stations]]></category>
		<category><![CDATA[climate data]]></category>
		<category><![CDATA[climate variation from valley to summit]]></category>
		<category><![CDATA[ERA5 reanalysis]]></category>
		<category><![CDATA[Everest weather stations]]></category>
		<category><![CDATA[extreme mountain climate]]></category>
		<category><![CDATA[glaciology]]></category>
		<category><![CDATA[high elevation weather monitoring challenges]]></category>
		<category><![CDATA[high-altitude meteorological data]]></category>
		<category><![CDATA[high-altitude meteorology]]></category>
		<category><![CDATA[Himalaya]]></category>
		<category><![CDATA[Himalayan glacier climate research]]></category>
		<category><![CDATA[impact of extreme weather on scientific instruments]]></category>
		<category><![CDATA[Khumbu]]></category>
		<category><![CDATA[monsoon]]></category>
		<category><![CDATA[Mount Everest]]></category>
		<category><![CDATA[open data]]></category>
		<category><![CDATA[open-access climate data Himalayan region]]></category>
		<category><![CDATA[quality control]]></category>
		<category><![CDATA[temperature lapse rate]]></category>
		<category><![CDATA[vertical climate profiling Mount Everest]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=257602</guid>

					<description><![CDATA[A new open-access dataset from six automatic weather stations spanning 3,810 to 8,810 metres on Mount Everest reveals how temperature, humidity, wind and radiation vary with altitude and season, and shows where the widely used ERA5 reanalysis falls short.]]></description>
										<content:encoded><![CDATA[<p>High on the flanks of Mount Everest, where the air holds barely a third of the oxygen found at sea level and winter storms can bury equipment under metres of snow, a network of automatic weather stations has been quietly recording some of the most extraordinary meteorological measurements ever collected. Between 2019 and 2025, a team of National Geographic Explorers, scientists and elite climbing Sherpas installed and maintained six automatic weather stations stretching from the village of Phortse at 3,810 metres to Bishop Rock at 8,810 metres, just below the summit of the world&#8217;s highest mountain. The resulting dataset, now published in the journal Earth System Science Data as a quality-controlled, open-access archive, spans nearly five vertical kilometres of elevation and offers scientists an unprecedented window into how temperature, humidity, wind and radiation behave across the full elevational range of Himalayan glaciers.</p>
<p>The effort behind the data was anything but routine. High Mountain Asia contains the largest glacierized area outside the polar regions, yet field research there is notoriously difficult. Low barometric pressure, extreme weather and steep terrain challenge both personnel and instruments, and the region&#8217;s stations are frequently disabled by riming, heavy snowfall, battery depletion and communication failures. The Everest stations were installed as part of the National Geographic and Rolex Perpetual Planet Everest Expeditions, with Sherpa climbers bolting custom aluminium tripods designed by Campbell Scientific directly to rock and guying them to additional anchor points. The Balcony station at 8,430 metres was toppled by extreme winds and anchor failure in January 2020, prompting a return expedition in May 2022 that installed the new Bishop Rock station at 8,810 metres and upgraded the South Col site.</p>
<p>Each station carries a carefully engineered suite of sensors. Air temperature and relative humidity are measured with Vaisala HMP155A-L5-PT probes housed in naturally ventilated 14-plate solar radiation shields, while barometric pressure is recorded with Vaisala PTB110 and PTB210 sensors inside the datalogger enclosures. Wind speed and direction initially came from R. M. Young 05108-45 anemometers, but all four units installed at the two highest sites failed in the extreme conditions, and the team replaced them with a polycarbonate version of the sensor plus a custom Richards C5C anemometer and pitot tube made by the Mount Washington Observatory. Radiation components are captured with Hukseflux pyranometers and pyrgeometers, and precipitation is measured at the two lowest stations with OTT Pluvio2 weighing gauges fitted with double-Alter wind shields to reduce the wind-induced undercatch that plagues snowfall measurement in exposed mountain terrain.</p>
<p>The engineering details reveal how altitude dictated design choices. At the lower stations, sensors sit 2 metres above the ground and batteries live inside the datalogger box; at the highest sites, sensors were mounted lower, at about 1.5 metres, to reduce both wind loading and the leverage, or torque, that a taller pole would exert on the tripod base during severe gusts. Batteries at the upper stations were placed in separate insulated boxes to protect them from extreme cold, and both boxes were bolted to rock to cut wind drag. The lower stations weigh roughly 60 kilograms each, the upper ones about 52 kilograms including the experimental pitot tube. Data are logged on Campbell Scientific CR1000X dataloggers at intervals from 10 minutes to daily, and the published archive provides quality-controlled hourly values in Coordinated Universal Time.</p>
<p>Because raw readings from such a hostile environment are riddled with artefacts, the team applied a multistage quality control procedure. Relative humidity values exceeding 100 percent were capped at that physical limit, and artificially low readings caused by the logger computing humidity with respect to water rather than ice at sub-zero temperatures were corrected using Buck&#8217;s equations for saturation vapour pressure. Periods of zero wind speed combined with zero directional variability were flagged as sensor freezing and removed rather than mistaken for calm conditions. Night-time incoming shortwave radiation below 7 watts per square metre was set to zero, and an albedo-based correction recalculated compromised radiation values whenever fresh snow or rime on the upward-facing sensor pushed the apparent surface albedo above the realistic threshold of 0.95.</p>
<p>The resulting climatology, though based on a short record, paints a vivid seasonal picture. Mean annual temperatures were 4.1 degrees Celsius at Phortse, minus 3.1 degrees at Base Camp and minus 10.2 degrees at Camp II. At South Col, July was the warmest month with a mean of minus 12.2 degrees and a striking diurnal range of 10.2 degrees, while February averaged minus 29.7 degrees. Precipitation clearly delineates the seasons: winters are predominantly dry, amounts build through the pre-monsoon, and the June-to-September monsoon delivers 72 percent of annual precipitation at Phortse and 77 percent at Base Camp. The mean annual precipitation gradient between the two sites, separated by roughly 1,500 metres, was minus 107 millimetres per kilometre, but it weakens by almost half during the monsoon, indicating that relative precipitation drops with elevation are far steeper outside the monsoon season.</p>
<p>Temperature gradients proved equally revealing. Between Phortse and Base Camp the mean temperature gradient was minus 4.8 degrees Celsius per kilometre, least negative in winter and most negative in the pre-monsoon. When the four stations with minimal data gaps were combined, the gradients became more strongly negative, at minus 6.5 degrees per kilometre in the pre-monsoon, minus 5.7 in the monsoon and minus 6.0 annually. The authors caution that the temperature-altitude relationship in the Khumbu region is fundamentally non-linear, which limits direct comparison with lapse rates calculated across different elevational ranges, and they therefore provide a non-linear equation for the gradient between Phortse and South Col. Radiation measurements add further nuance: maximum daily incoming longwave radiation during the monsoon exceeds 350 watts per square metre at Phortse but stays below about 250 at South Col, likely reflecting the colder, drier, less cloudy atmosphere aloft.</p>
<p>A central contribution of the study is a rigorous comparison with ERA5, the fifth-generation global reanalysis produced by the European Centre for Medium-Range Weather Forecasts that many researchers use as a stand-in for sparse mountain observations. Extracted from the nearest grid point at the 350 hectopascal pressure level, ERA5 temperatures tracked observed variability at South Col reasonably well, with coefficients of determination above 0.6 in both 2019 and 2022, and even higher agreement of 0.85 with the short Balcony record in 2019. But the reanalysis systematically underestimated air temperature, showed far less diurnal variability because the pressure level represents the free atmosphere rather than the surface, and consistently overestimated mean wind speeds, with mean absolute errors of 4.8 and 5.3 metres per second in 2019 and 2022. Relative humidity comparisons carried mean absolute errors above 20 percent, though 10-day running means from both datasets clearly captured the monsoon onset at South Col on 1 July 2019 and 14 June 2022.</p>
<p>The practical implications reach well beyond atmospheric science. Observations from South Col reveal mean winds that can exceed 30 metres per second and gusts above 60 metres per second, conditions capable of blowing mountaineers off their feet and inducing cold injuries. Pressure data from the network have already shown that oxygen availability at the summit varies on synoptic timescales, meaning the apparent elevation of Everest, how high the mountain would feel without supplemental oxygen, can shift by almost 750 metres, and a winter ascent without bottled oxygen may at times be impossible. The team anticipates that ERA5, once its time-varying biases are corrected with empirical-statistical or machine-learning approaches, could be used to gap-fill and extend the intermittent records from Everest&#8217;s upper slopes, enabling hyper-local forecasts that help expeditions identify optimal climbing windows.</p>
<p>The archive also underpins research on the region&#8217;s fragile cryosphere and its role as a freshwater source for downstream communities. It has already been used to estimate surface energy balances at the summit and at South Col Glacier, revealing a high-altitude ice system acutely sensitive to changes in effective precipitation because of extremely high insolation and its responsiveness to albedo variations. Combined with the longer-running EvK2CNR and GLACIOCLIM networks at lower elevations, the new data allow quantification of elevational gradients in key meteorological variables across roughly five vertical kilometres, covering the entire glacierized range of the Khumbu region, information essential for distributed glacier and hydrological modelling. And because meteorological measurements from the highest reaches only began in 2022, the authors note that scientists are still at the very beginnings of exploring the weather of this extreme environment, with questions the data might answer likely to grow rapidly across disciplines far beyond the climate sciences.</p>
<p><strong>Subject of Research:</strong> High-altitude weather station observations across the Mount Everest region of Nepal</p>
<p><strong>Article Title:</strong> Weather station data from the Mount Everest region, Nepal: 3810–8810 m above sea level</p>
<p><strong>Article References:</strong> Khadka, A., Perry, L. B., Matthews, T., Sherpa, T. G., Shrestha, C. B., Shrestha, D., Aryal, D., Tuladhar, S., Pradhananga, N., Kayastha, D., Raichle, B., Athans, P., Sherpa, D. Y., Garrett, K., Wheeler, G., Young, T., &amp; Elmore, A. (2026). Weather station data from the Mount Everest region, Nepal: 3810–8810 m above sea level. <em>Earth System Science Data, 18</em>(10), 7253-7267. <a href="https://doi.org/10.5194/essd-18-7253-2026" rel="noopener noreferrer">https://doi.org/10.5194/essd-18-7253-2026</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.5194/essd-18-7253-2026" rel="noopener noreferrer">10.5194/essd-18-7253-2026</a></p>
<p><strong>Keywords:</strong> Mount Everest, automatic weather stations, Himalaya, climate data, ERA5 reanalysis, temperature lapse rate, monsoon, glaciology, high-altitude meteorology, Khumbu, quality control, open data</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">257602</post-id>	</item>
		<item>
		<title>Scientists Unveil Massive Global Database of Rock Properties Spanning 70 Countries</title>
		<link>https://scienmag.com/scientists-unveil-massive-global-database-of-rock-properties-spanning-70-countries/</link>
		
		<dc:creator><![CDATA[Violet Maxwell]]></dc:creator>
		<pubDate>Fri, 09 Oct 2026 13:30:10 +0000</pubDate>
				<category><![CDATA[Earth Science]]></category>
		<category><![CDATA[carbon dioxide sequestration]]></category>
		<category><![CDATA[carbon sequestration]]></category>
		<category><![CDATA[CSIRO]]></category>
		<category><![CDATA[data standardization in geoscience]]></category>
		<category><![CDATA[database]]></category>
		<category><![CDATA[geoscience data consolidation]]></category>
		<category><![CDATA[geothermal energy]]></category>
		<category><![CDATA[global petrophysical data]]></category>
		<category><![CDATA[groundwater]]></category>
		<category><![CDATA[groundwater flow prediction]]></category>
		<category><![CDATA[open data]]></category>
		<category><![CDATA[permeability]]></category>
		<category><![CDATA[permeability and porosity datasets]]></category>
		<category><![CDATA[petrophysics]]></category>
		<category><![CDATA[planetary-scale rock property measurements]]></category>
		<category><![CDATA[porosity]]></category>
		<category><![CDATA[radioactive waste repository analysis]]></category>
		<category><![CDATA[rock properties]]></category>
		<category><![CDATA[Rock property database]]></category>
		<category><![CDATA[subsurface modeling inputs]]></category>
		<category><![CDATA[subsurface modelling]]></category>
		<category><![CDATA[thermal conductivity]]></category>
		<category><![CDATA[thermal conductivity of rocks]]></category>
		<category><![CDATA[underground heat transfer modeling]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=254109</guid>

					<description><![CDATA[Researchers at CSIRO have compiled more than 112,000 rock property measurements from over 600 publications across 70 countries into a freely accessible Global Petrophysical Database designed to support groundwater, carbon storage, resource exploration and waste disposal modelling.]]></description>
										<content:encoded><![CDATA[<p>Every model of the subsurface, whether it is designed to predict how groundwater will move beneath a farming region, how carbon dioxide will behave after injection into a deep saline aquifer, or how heat will migrate around a repository for radioactive waste, depends on one deceptively simple set of inputs: the physical properties of the rocks themselves. Porosity determines how much fluid a rock can store, permeability governs how easily that fluid can move through it, and thermal conductivity controls how heat is transported through the solid framework. Yet for decades, researchers who needed these numbers have faced a frustrating reality. The measurements exist, scattered across thousands of journal papers, government reports and legacy datasets, but they are inconsistent, incompletely documented and often buried in formats that resist modern computational analysis. A team of Australian researchers now believes it has solved that problem on a planetary scale.</p>
<p>In a preprint currently under review for the data-focused journal Earth System Science Data, Heather Anne Sheldon of CSIRO Mineral Resources in Canberra and her colleagues across CSIRO&#8217;s energy, environment and information management divisions describe the creation of a Global Petrophysical Database, or GPD, that consolidates more than 112,000 individual rock property measurements drawn from over 600 publications. The compilation spans data from 70 countries and incorporates ten existing databases that had previously been maintained in isolation. Crucially, the database includes both in-situ measurements, taken where the rock actually sits in the subsurface, and laboratory measurements performed on samples brought to the surface, allowing users to compare how properties behave under natural conditions and under controlled experimental ones.</p>
<p>The technical scope of the resource is what sets it apart from earlier attempts. Previous petrophysical databases have tended to be limited in geographic extent, or restricted in the range of properties, rock types or measurement techniques they cover. The GPD deliberately cuts across those boundaries, encompassing igneous, metamorphic and sedimentary rocks alike, and focusing primarily on the properties that control subsurface fluid flow, mass transport and heat transport. Each measurement is accompanied by associated metadata, including rock type, geographic location and depth, the contextual information that transforms a bare number into something a modeller can actually use. Without that context, a permeability value is nearly meaningless, because the same rock type can behave radically differently at two kilometres depth than it does in a surface outcrop.</p>
<p>The practical motivation behind the project is easy to understand for anyone who has attempted a literature synthesis in this field. As the authors note, collating petrophysical property data from the literature is time consuming and is frequently hampered by inconsistent or incomplete metadata. A hydrogeologist modelling aquifer recharge might spend weeks tracking down porosity measurements for a particular sandstone formation, only to find that half the relevant papers omit the depth at which samples were collected or fail to specify the measurement technique. Those omissions matter enormously. Permeability measured on a plug in a laboratory at ambient conditions can differ by orders of magnitude from the effective permeability of the same formation in the ground, where confining stress, temperature and natural fracturing all play their part.</p>
<p>The applications the authors identify span some of the most consequential geoscience challenges of the coming decades. Groundwater management depends on accurate representations of how water moves through aquifers and the low-permeability layers that confine them. Carbon geo-sequestration, the injection of carbon dioxide into deep geological formations as a climate mitigation strategy, requires confident knowledge of reservoir porosity and caprock integrity. Resource exploration, whether for minerals, hydrocarbons or geothermal energy, relies on petrophysical properties to interpret geophysical surveys and to build predictive models of ore-forming fluid flow. And subsurface waste disposal, including the long-term storage of hazardous materials, demands an exceptionally rigorous understanding of how slowly, or quickly, fluids and dissolved species can migrate through host rocks over thousands of years.</p>
<p>What makes the database genuinely powerful for the modern era is its accessibility and its machine-readability. The GPD is publicly available through a persistent digital object identifier hosted by CSIRO, and it is supported by a graphical web interface called the Global Petrophysical Database Explorer, which is designed to facilitate rapid exploration and visualisation of the data. A researcher can query the collection, filter by rock type or property, and visualise distributions without writing a single line of code, while those building numerical models can ingest the underlying data directly. In an age when machine learning is increasingly applied to geoscience problems, a large, well-documented, consistently structured dataset of this kind is exactly the raw material such methods require, and its absence has been a persistent bottleneck.</p>
<p>The scale of the underlying effort should not be understated. Harmonising data from more than 600 publications means confronting a babel of units, measurement standards, rock classification schemes and reporting conventions that have evolved independently across subdisciplines and decades. A permeability reported in millidarcies by a petroleum engineer, one reported in square metres by a hydrogeologist, and a third reported only qualitatively in an old mining report must all be reconciled into a coherent framework. The decision to fold in ten pre-existing databases, rather than starting from scratch, reflects a pragmatic philosophy: the community has already generated an enormous volume of valuable measurements, and the highest-value contribution is to make them findable, comparable and reusable, in line with the open-data principles that journals such as Earth System Science Data were created to champion.</p>
<p>The timing is significant. As governments accelerate investment in carbon capture and storage, geothermal energy and managed aquifer recharge, demand for reliable subsurface property data is rising sharply, and the cost of poor data is measured not just in wasted research time but in flawed decisions about where to inject carbon dioxide or how to site critical infrastructure. A single, globally scoped reference dataset allows modellers to benchmark their assumptions against real measurements from analogous rock types and settings worldwide, rather than defaulting to textbook values that may bear little resemblance to the formations under study. It also exposes gaps: where the database is thin, it points directly to where new field and laboratory campaigns would deliver the greatest scientific return.</p>
<p>The work is currently at the preprint stage, published as a discussion paper with peer review open and ongoing, so the details may evolve as referees and community members weigh in. But the core contribution is already tangible and freely available. For the researchers who built it, the goal was straightforward: to save others the time they themselves had spent hunting for rock property values, and to give the geoscience community a shared foundation for modelling and decision making. If the database achieves the uptake its designers hope, the era of rebuilding the same scattered literature review for every new subsurface project may finally be drawing to a close, replaced by a single, searchable, global record of what we know about the rocks beneath our feet.</p>
<p><strong>Subject of Research:</strong> A global petrophysical database compiling rock property measurements for subsurface modelling</p>
<p><strong>Article Title:</strong> A global petrophysical database for igneous, metamorphic and sedimentary rocks</p>
<p><strong>Article References:</strong> Sheldon, H. A., Dewhurst, D. N., Bekele, E., Raiber, M., Mallants, D., Turnadge, C., &amp; Williams, G. (2026). A global petrophysical database for igneous, metamorphic and sedimentary rocks. <a href="https://doi.org/10.5194/essd-2026-556" rel="noopener noreferrer">https://doi.org/10.5194/essd-2026-556</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.5194/essd-2026-556" rel="noopener noreferrer">10.5194/essd-2026-556</a></p>
<p><strong>Keywords:</strong> petrophysics, rock properties, porosity, permeability, thermal conductivity, database, groundwater, carbon sequestration, geothermal energy, subsurface modelling, CSIRO, open data</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">254109</post-id>	</item>
		<item>
		<title>Massive Open Dataset Captures Shaking from 667 Colombian Earthquakes</title>
		<link>https://scienmag.com/massive-open-dataset-captures-shaking-from-667-colombian-earthquakes/</link>
		
		<dc:creator><![CDATA[Violet Maxwell]]></dc:creator>
		<pubDate>Fri, 09 Oct 2026 07:18:59 +0000</pubDate>
				<category><![CDATA[Earth Science]]></category>
		<category><![CDATA[Colombia]]></category>
		<category><![CDATA[Colombia earthquake history]]></category>
		<category><![CDATA[Colombia seismic risk assessment]]></category>
		<category><![CDATA[Colombian earthquake ground motion dataset]]></category>
		<category><![CDATA[earthquake catalog Colombia 667 events]]></category>
		<category><![CDATA[Earthquake engineering]]></category>
		<category><![CDATA[earthquake monitoring in Colombia]]></category>
		<category><![CDATA[flatfile]]></category>
		<category><![CDATA[Fourier amplitude spectrum]]></category>
		<category><![CDATA[ground motion recordings Colombia]]></category>
		<category><![CDATA[ground-motion dataset]]></category>
		<category><![CDATA[moment magnitude]]></category>
		<category><![CDATA[open data]]></category>
		<category><![CDATA[open-access seismic data Colombia]]></category>
		<category><![CDATA[seismic data processing standards]]></category>
		<category><![CDATA[seismic hazard]]></category>
		<category><![CDATA[shallow earthquake ground motions]]></category>
		<category><![CDATA[spectral acceleration]]></category>
		<category><![CDATA[strong-motion records]]></category>
		<category><![CDATA[subduction]]></category>
		<category><![CDATA[subduction zone earthquakes South America]]></category>
		<category><![CDATA[tectonic plate collision Colombia]]></category>
		<category><![CDATA[Vs30]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=252501</guid>

					<description><![CDATA[An open-access compilation of 7550 uniformly processed recordings from 667 shallow Colombian earthquakes promises to sharpen seismic hazard maps and building design across northwestern South America.]]></description>
										<content:encoded><![CDATA[<p>Colombia sits at one of the most complicated tectonic crossroads on Earth. Beneath its northwestern corner, the Nazca and Caribbean plates grind beneath the South American continent, while the Panamá-Chocó block pushes in from the northwest, producing a tangled mixture of subduction megathrusts, deep intraslab ruptures, and shallow crustal faults that slice through the Andes themselves. This collision zone has repeatedly unleashed devastating earthquakes, from the catastrophic magnitude 8.8 Colombia-Ecuador event of 1906, one of the largest ever recorded worldwide, to the 1999 Eje Cafetero earthquake that killed more than 2000 people. Yet despite this well-documented seismic menace, the country has never had a comprehensive, uniformly processed catalogue of the actual ground motions its earthquakes generate. A new study published in Earth System Science Data changes that, presenting an open-access dataset of 7550 three-component acceleration recordings from 667 shallow earthquakes, compiled and processed to the same rigorous standards as the flagship ground-motion databases of North America, Europe, and Japan.</p>
<p>The work was led by Daniel Martinez-Jaramillo, a doctoral researcher at the Universidad Nacional Autónoma de México, together with Sreeram-Reddy Kotha of ISTerre in Grenoble, F. Ramón Zúñiga, and Pierre Lacan. Their starting point was the national seismic network of the Colombian Geological Survey, known by its FDSN code CM, which has operated since 1993 and currently comprises 506 seismological and accelerometric stations distributed across the country. Local networks monitoring volcanoes, mining districts, and oil and gas fields contributed additional stations. Upon request, the survey supplied more than 10,000 quality-checked acceleration time series. The team applied strict selection criteria: events between 0 and 14 degrees north latitude and 69 to 82 degrees west longitude, at depths shallower than 50 kilometres, with magnitudes above 4 and stations within 400 kilometres of the epicentre. After discarding volcanic signals and spurious spikes, 667 earthquakes and 7550 records survived the cut, forming what seismologists call a flatfile: a single table describing every earthquake, every recording site, and every measured level of shaking.</p>
<p>The technical care invested in processing these waveforms is what elevates the dataset from a raw archive to a scientific instrument. Each acceleration time series underwent the same sequence of corrections, following protocols established for the Italian strong-motion databases: baseline correction, cosine tapering, a second-order acausal bandpass Butterworth filter, double integration to obtain displacement, linear detrending of that displacement, and double differentiation back to corrected acceleration. The high-pass corner of the filter was made magnitude-dependent, ranging from 0.15 hertz for the smallest events down to 0.05 hertz for the largest, while the low-pass corner was fixed at a median value of 32.26 hertz. A sensitivity analysis showed that this fixed low-pass choice barely affects most records, shifting the orientation-independent RotD50 peak ground acceleration by an average of just 2.31 percent, though 6.4 percent of recordings, concentrated around magnitude 4.5 at distances beyond 250 kilometres, showed variations above 10 percent, likely reflecting high-frequency noise near the filter corner.</p>
<p>The intensity measures extracted from the corrected records are comprehensive. The dataset provides peak ground acceleration, velocity, and displacement for all three components, along with 5 percent-damped spectral accelerations at 31 oscillator periods spanning 0.01 to 8 seconds, computed as RotD50 values, the median over all non-redundant horizontal orientations, a measure now standard in modern ground-motion modelling. Fourier amplitude spectra for each component cover frequencies from 0.04 to 50 hertz, smoothed with the Konno-Ohmachi function, and the effective amplitude spectrum, a geometric combination of the two horizontal components, is also included. Crucially, the authors report the lowest and highest usable frequencies for every record, bounded by a safety factor of 1.25 around the filter corners so that users know exactly which spectral values are trustworthy. This usable-bandwidth information, often omitted from older compilations, allows engineers and seismologists to weight each record appropriately in downstream analyses.</p>
<p>Perhaps the most technically inventive element concerns magnitudes. Moment magnitude is the preferred scale for hazard analysis, but reliable agency-reported values were not available for every event. Where they were missing, the team estimated magnitudes from the recordings themselves by fitting a single-corner Brune omega-squared source model to each corrected spectrum, using a coarse grid search followed by non-linear least-squares refinement in log-log space. From the resulting corner frequency, the source radius was computed assuming a shear-wave velocity of 3.5 kilometres per second, and the seismic moment followed from the Eshelby circular-crack relation under an assumed constant stress drop. Testing stress drops of 1, 3, and 5 megapascals against independently reported moment magnitudes from the Global CMT catalogue and other agencies, the authors found that 5 megapascals reproduced catalogued values best. This procedure homogenised the entire catalogue, ultimately widening the magnitude range to 3.5 through 7.2, with 24 percent of events falling below magnitude 4 once recalculated.</p>
<p>Site characterisation received equal attention. The 227 recording stations are described by their time-averaged shear-wave velocity in the upper 30 metres, the widely used Vs30 parameter, together with horizontal-to-vertical spectral ratios and predominant site periods derived from seismograms. For 154 sites, these parameters come from a recent northwestern South America database, and 28 of them have in-situ, microtremor-based Vs30 measurements. The remaining stations rely on Vs30 values inferred from a topographic-slope-based map of Colombia, a pragmatic but imperfect proxy. The authors are candid about this limitation, noting that proxy-based site parameters introduce additional epistemic uncertainty that may inflate the site-to-site variability observed in their validation, and pointing to H/V spectral ratio classification schemes as a promising complement for future work. Distance metrics are equally thorough: epicentral and hypocentral distances are reported for all events, while finite-fault measures such as Joyner-Boore distance are computed for the 42 earthquakes larger than magnitude 5.5, with rupture dimensions scaled from published empirical relations.</p>
<p>Validation came through residual analysis, the standard test of whether a new dataset behaves consistently with established ground-motion prediction models. The team compared observed peak ground acceleration, spectral accelerations, and effective amplitude spectra against the global NGA-West2 model of Abrahamson and colleagues from 2014, its regional adaptation for northern South America published by Arteta and colleagues in 2023, and the Bayless-Abrahamson 2019 Fourier-spectrum model. The residuals, decomposed into between-event, between-site, and leftover components, showed no significant biases and followed Gaussian-like distributions confirmed by Shapiro-Wilk tests. As expected, the regional model produced lower overall variability than the global one, and the between-event trends support the internal consistency of the compilation. One instructive exception emerged: within the magnitude range of roughly 4.2 to 4.7, the regional model showed a negative trend in between-event residuals, suggesting that magnitude calibration at these moderate levels could benefit from future refinement.</p>
<p>The statistical footprint of the dataset reveals both its strengths and its character. About 85 percent of records lie beyond 100 kilometres from the epicentre, while only 5.3 percent come from within 50 kilometres, meaning near-source shaking remains comparatively sparse. Roughly 65 percent of the records originate from events shallower than 20 kilometres, confirming the dataset&#8217;s focus on shallow crustal seismogenic sources, with the remainder spanning depths down to about 55 kilometres. Strong events are well represented: 716 records, or 9.5 percent of the total, come from earthquakes above magnitude 5.5, including the magnitude 7.2 Pizarro earthquake of 2004 on the Pacific coast and the magnitude 6.1 San Juanito earthquake of 2023 in eastern Colombia. Epicentres of 488 events fall on continental territory and 179 offshore, and the authors deliberately leave tectonic classification, whether subduction interface, intraslab, or crustal, to users, since appropriate assignments depend on the slab models each analyst chooses.</p>
<p>The practical payoff could be substantial. Ground-motion datasets of this kind are the raw material for probabilistic seismic hazard assessment, the framework that underlies building codes, and for the emerging generation of partially non-ergodic ground-motion models, which replace globally averaged predictions with region-specific ones by exploiting dense, well-characterised observations. Colombian engineers designing buildings, and officials updating national hazard maps, will gain a resource calibrated to the actual faults and soils beneath their feet rather than to California or Japan. The Fourier spectra additionally open a window onto source physics, allowing researchers to constrain corner frequencies, seismic moments, and stress drops for Colombian earthquakes directly. Released openly under a Creative Commons Attribution 4.0 licence on Zenodo, with the underlying time series available through the Colombian Geological Survey&#8217;s acceleration catalogue and the FDSN network CM, the dataset invites reuse far beyond Colombia&#8217;s borders, offering a template for other seismically active but data-poor regions of the world where the next damaging earthquake is only a matter of time.</p>
<p><strong>Subject of Research:</strong> A uniformly processed open dataset of ground-motion recordings from shallow earthquakes in Colombia for seismic hazard analysis</p>
<p><strong>Article Title:</strong> Ground-motion dataset for shallow earthquakes in Colombia</p>
<p><strong>Article References:</strong> Martinez-Jaramillo, D., Kotha, S.-R., Zúñiga, F. R., &amp; Lacan, P. (2026). Ground-motion dataset for shallow earthquakes in Colombia. <em>Earth System Science Data, 18</em>(10), 7403-7415. <a href="https://doi.org/10.5194/essd-18-7403-2026" rel="noopener noreferrer">https://doi.org/10.5194/essd-18-7403-2026</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.5194/essd-18-7403-2026" rel="noopener noreferrer">10.5194/essd-18-7403-2026</a></p>
<p><strong>Keywords:</strong> Colombia, ground-motion dataset, seismic hazard, earthquake engineering, strong-motion records, moment magnitude, spectral acceleration, Fourier amplitude spectrum, subduction, Vs30, flatfile, open data</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">252501</post-id>	</item>
		<item>
		<title>Unclear data licences are stalling the world&#8217;s first global ecosystem map</title>
		<link>https://scienmag.com/unclear-data-licences-are-stalling-the-worlds-first-global-ecosystem-map/</link>
		
		<dc:creator><![CDATA[Gavin Prescott]]></dc:creator>
		<pubDate>Thu, 08 Oct 2026 18:02:32 +0000</pubDate>
				<category><![CDATA[Biology]]></category>
		<category><![CDATA[Earth Science]]></category>
		<category><![CDATA[barriers to global environmental data integration]]></category>
		<category><![CDATA[biodiversity]]></category>
		<category><![CDATA[conservation]]></category>
		<category><![CDATA[copyright]]></category>
		<category><![CDATA[Creative Commons]]></category>
		<category><![CDATA[data licensing]]></category>
		<category><![CDATA[data licensing issues in environmental science]]></category>
		<category><![CDATA[data sharing policies in environmental monitoring]]></category>
		<category><![CDATA[data sovereignty]]></category>
		<category><![CDATA[development of comprehensive ecosystem maps]]></category>
		<category><![CDATA[ecosystem mapping]]></category>
		<category><![CDATA[Ecosystem mapping challenges]]></category>
		<category><![CDATA[geospatial data]]></category>
		<category><![CDATA[global ecosystem dataset availability]]></category>
		<category><![CDATA[Global Ecosystems Atlas]]></category>
		<category><![CDATA[Group on Earth Observations]]></category>
		<category><![CDATA[impact of ambiguous data licenses]]></category>
		<category><![CDATA[legal barriers to satellite imagery data]]></category>
		<category><![CDATA[open data]]></category>
		<category><![CDATA[open data for ecological research]]></category>
		<category><![CDATA[PLOS Ecosystems]]></category>
		<category><![CDATA[role of law in ecological data collection]]></category>
		<category><![CDATA[scientific data reuse restrictions]]></category>
		<category><![CDATA[standardization of ecosystem data licensing]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=248777</guid>

					<description><![CDATA[A new PLOS Ecosystems analysis of 343 datasets finds that nearly a third carry no licence information, legally blocking efforts to build the first comprehensive global map of Earth's ecosystems.]]></description>
										<content:encoded><![CDATA[<p>A quarter of the way through the twenty-first century, humanity still lacks a comprehensive, authoritative, and scientifically sound map of the world&#8217;s ecosystems. That may sound surprising in an age of ubiquitous satellite imagery and increasingly powerful machine-learning classification algorithms, but the bottleneck is not technology. It is law. According to a new opinion article published in PLOS Ecosystems, the greatest obstacle to building a truly global picture of Earth&#8217;s ecosystems is something as mundane as ambiguous data licensing—and the scale of the problem is now quantified for the first time.</p>
<p>The article, written by Falko Buschke, David Patterson, and colleagues at the Global Ecosystems Atlas initiative, convened by the Group on Earth Observations (GEO) Secretariat in Geneva, together with collaborators at James Cook University in Australia, draws on an extensive audit of candidate datasets for the Atlas. As of January 2026, the team had compiled a catalogue of 343 ecosystem datasets that could, in principle, be combined—with full attribution to their owners—into the first comprehensive, standardised, and open global map of ecosystems. Yet nearly one-third of those datasets, 29.7 percent, provide no explicit guidance whatsoever on how the data may be reused.</p>
<p>That figure matters because of a legal default that most people never think about. Like all creative works, ecosystem maps remain under exclusive copyright unless clearly stated otherwise. In practical terms, this means that even when a dataset is freely downloadable from a public website, users must assume it is protected and cannot be reused without prior permission. Permission takes the form of either an explicit data licence or direct authorisation from the provider. A dataset that sits openly on the internet but carries no licence statement is, legally speaking, locked.</p>
<p>The consequences ripple far beyond academic inconvenience. The Global Ecosystems Atlas is designed to bring together existing ecosystem information from trusted sources—national authorities, nongovernmental organisations, and the peer-reviewed scientific literature—and to fill remaining gaps using classification models trained on expert-evaluated, point-based data. Local expertise remains fundamental to this approach: field surveys of species composition, annotations of ecosystem-class records, and manual refinements of historical maps all feed into the final product. Top-down global approaches, however sophisticated, have repeatedly proven resistant to producing accurate maps without this local validation. When a third of the underlying information cannot legally be integrated, decades of preceding investment in local knowledge are effectively stranded.</p>
<p>The audit revealed a licensing landscape that is fragmented in ways that impose real technical and legal costs. Beyond the unlicensed third of the catalogue, a further 20.4 percent of datasets are governed by custom licences—legal terms tailored to a specific dataset or institution. These range from institutional policies of national governments and large research organisations to personalised modifications of standard licences. Some data owners even offer different conditions to different users, permitting free use by individuals and small businesses while restricting large corporations above a certain size. For nonexperts, navigating such bespoke terms is daunting, and for project managers, custom licences introduce legal overheads to verify compliance. Worse, when custom licences embed copyleft conditions—requiring that any derivative work be shared under identical terms—they severely limit uptake in composite products.</p>
<p>Only 47.8 percent of the datasets in the catalogue are shared under Creative Commons licences, the standardised public instruments designed precisely to make reuse conditions clear. Within that open minority, the majority—39.6 percent of all datasets—require simple attribution (CC-BY), while just 4 percent restrict reuse to noncommercial purposes (CC-BY-NC). On the surface, that seems like a healthy majority of usable data. But the mathematics of aggregation is unforgiving: any derived product assembled from multiple datasets must be licensed according to the most stringent conditions of its constituent parts.</p>
<p>This &#8216;weakest link&#8217; principle means that even a single restrictive dataset can contaminate an entire global product. Include one dataset carrying a noncommercial condition, and the whole composite is barred from commercial applications such as corporate biodiversity reporting—a rapidly growing use case as companies come under pressure to disclose nature-related risks. Add a no-derivatives condition, and the product cannot be combined with any other dataset at all. A share-alike condition constrains the licence of everything downstream and can conflict with datasets that carry slightly different terms. In a synthesis that might draw on hundreds of sources, a handful of ambiguous or restrictive licences can quietly determine what the final map is legally allowed to do.</p>
<p>It is tempting to conclude that all ecosystem data should simply be released under maximally open licences. The authors resist that blanket prescription, and for good reason. Biodiversity data can be regarded as social infrastructure, and respecting data sovereignty—particularly that of Indigenous peoples and marginalised communities—is essential for equitable and effective conservation outcomes. Selling commercial access to biodiversity information is also a legitimate way for custodians to recoup the costs of producing and maintaining it, which is only possible when commercial rights are protected. Open data advocacy that ignores these realities risks repeating the extractive patterns that conservation is trying to move beyond.</p>
<p>In their discussions with data owners—scientists, government officials, and NGO researchers—the team found that licensing is often an afterthought, with stakeholders receiving little guidance on how to select and apply licences that support their intended uses. Their response is a set of three practical recommendations. First, license ecosystem data explicitly: choosing an appropriate Creative Commons licence with a user-friendly selection tool and copying the licence information alongside the data—on the access website, or in a ReadMe or metadata file accompanying GIS files—can take minutes. Second, use standardised licences wherever possible, since Creative Commons terms are translated into many languages and widely understood, whereas custom licences often impose the same restrictions in inaccessible legal jargon. Third, when restrictions are genuinely necessary, build in mechanisms to navigate them legally. The World Database on Protected Areas offers a prominent model: it maintains a public version for noncommercial use alongside a licensed version, accessible through the Integrated Biodiversity Assessment Tool, for commercial purposes.</p>
<p>The authors emphasise that initiatives like the Global Ecosystems Atlas are not intended to replace existing ecosystem information but to build on it, adding value by bringing datasets together under consistent standards. That vision depends on recognising the effort and investment behind every existing map, and on making the legal terms of reuse as clear as the data themselves. Clear data licences, the article argues, are an essential though often overlooked requirement for bottom-up global ecosystem mapping. The alternative is a world where the raw material for understanding Earth&#8217;s ecosystems exists in abundance—yet remains legally invisible to the very efforts trying to assemble it.</p>
<p><strong>Subject of Research:</strong> Data licensing barriers to integrating local ecosystem datasets into global ecosystem maps</p>
<p><strong>Article Title:</strong> Ambiguous data licences undermine global ecosystem mapping efforts</p>
<p><strong>Article References:</strong> Buschke, F., Patterson, D., Cresswell, B. J., Gros-Dubois, N., Lloyd, T. J., Gomersall, L., Young, A. R., Gevorgyan, Y., &amp; Murray, N. J. (2026). Ambiguous data licences undermine global ecosystem mapping efforts. <em>PLOS Ecosystems, 1</em>(1), e0000017. <a href="https://doi.org/10.1371/journal.pesy.0000017" rel="noopener noreferrer">https://doi.org/10.1371/journal.pesy.0000017</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1371/journal.pesy.0000017" rel="noopener noreferrer">10.1371/journal.pesy.0000017</a></p>
<p><strong>Keywords:</strong> ecosystem mapping, data licensing, Global Ecosystems Atlas, Creative Commons, open data, biodiversity, data sovereignty, copyright, Group on Earth Observations, conservation, geospatial data, PLOS Ecosystems</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">248777</post-id>	</item>
		<item>
		<title>Global coalition commits $1.8 billion to build open data for AI models of biology</title>
		<link>https://scienmag.com/global-coalition-commits-1-8-billion-to-build-open-data-for-ai-models-of-biology/</link>
		
		<dc:creator><![CDATA[Denise Maddox]]></dc:creator>
		<pubDate>Wed, 07 Oct 2026 13:36:35 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[AI and genomics]]></category>
		<category><![CDATA[AI in biology]]></category>
		<category><![CDATA[artificial intelligence in biology]]></category>
		<category><![CDATA[Biohub]]></category>
		<category><![CDATA[biological data measurement technology]]></category>
		<category><![CDATA[collaborative biological data initiatives]]></category>
		<category><![CDATA[cryo-electron tomography]]></category>
		<category><![CDATA[Department of Energy]]></category>
		<category><![CDATA[digital experimentation in biology]]></category>
		<category><![CDATA[disease mechanism modeling]]></category>
		<category><![CDATA[Google DeepMind]]></category>
		<category><![CDATA[Isomorphic Labs]]></category>
		<category><![CDATA[large-scale biological data funding]]></category>
		<category><![CDATA[NIH]]></category>
		<category><![CDATA[open data]]></category>
		<category><![CDATA[Open data for AI-driven biological modeling]]></category>
		<category><![CDATA[open resource for biological research]]></category>
		<category><![CDATA[predictive models]]></category>
		<category><![CDATA[predictive models of living systems]]></category>
		<category><![CDATA[single-cell biology]]></category>
		<category><![CDATA[transforming biology into a predictive science]]></category>
		<category><![CDATA[Virtual Biology Initiative]]></category>
		<category><![CDATA[virtual cell]]></category>
		<category><![CDATA[virtual cell simulation]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=244605</guid>

					<description><![CDATA[Biohub, the Department of Energy, NIH, Google DeepMind, Isomorphic Labs, Meta, and other partners have committed $1.8 billion to generate open, AI-ready biological data for building predictive models of cells and disease.]]></description>
										<content:encoded><![CDATA[<p>In what organizers describe as the largest coordinated investment ever made in generating artificial intelligence-ready biological data, Biohub, the U.S. Department of Energy, the National Institutes of Health, and a roster of new funding partners announced on October 7, 2026 a combined commitment of $1.8 billion in funding, data, computation, and new measurement technology. The effort, an expansion of the Virtual Biology Initiative first unveiled in April 2026, aims to produce an open resource that the global research community can use to train predictive models of living systems. The stated ambition is nothing less than a virtual cell: a computational model accurate enough that scientists could perform experiments digitally, predicting how any cell responds to an intervention before ever touching a pipette. If successful, the initiative could compress the timelines for understanding disease mechanisms and developing new therapies, transforming biology from a largely descriptive science into a predictive one.</p>
<p>The scale of the challenge explains why no single institution is attempting it alone. Modern AI models in biology, from protein structure prediction to whole-cell simulators, are ultimately limited by the quality and breadth of the experimental data used to train them. Today&#8217;s datasets capture cell responses to interventions across only a small fraction of the cell types and conditions that matter in human health. The Virtual Biology Initiative is designed to close that gap by coordinating data generation across institutions and disciplines, expanding cell response measurements to far more cell types and conditions than have yet been studied, and building and validating technologies capable of studying cells and their interactions at greater scale, speed, and accuracy. The result, organizers say, will be a foundational, openly accessible dataset that no laboratory, company, or government agency could produce on its own.</p>
<p>The federal contribution anchors the effort on two fronts. The Department of Energy&#8217;s Office of Science will invest more than $500 million over five years in laboratory measurement, modeling, and computation toward building the AI-ready open data resource. That investment flows through the Genesis Mission, a cross-agency initiative led by DOE, and draws on some of the most powerful scientific infrastructure in the world: exascale supercomputing, X-ray and neutron scattering facilities, cryo-electron microscopy and tomography, and autonomous laboratories distributed across the National Laboratory system. Darío Gil, DOE&#8217;s Under Secretary for Science, framed the partnership as a new standard for open science, pointing to the combination of DOE&#8217;s computing and measurement assets — including user facilities at the Joint Genome Institute, the Environmental Molecular Sciences Laboratory, and advanced structural beamlines — with Biohub&#8217;s AI models, tool development, and biological data capabilities.</p>
<p>The National Institutes of Health, for its part, will coordinate the contribution of relevant datasets, repositories, and knowledge bases developed through more than $500 million in prior federal investment aligned to the initiative. Through its Bio Genesis Mission, NIH plans to bring together existing biomedical datasets, national data infrastructure, and research programs to help build AI-ready resources for the broader scientific community. The resources in play are substantial: national biomedical repositories catalogued by NIH&#8217;s National Library of Medicine and the National Center for Biotechnology Information, as well as NIH Common Fund programs that are already developing coordinated biological atlases, shared data standards, and AI-ready biomedical datasets. Biohub will work with NIH to standardize these datasets for AI model training, a step that many researchers consider as important as the raw data itself, since heterogeneous formats and inconsistent annotations have long hampered large-scale machine learning in biology.</p>
<p>Industry is contributing at significant scale as well. Google DeepMind, Isomorphic Labs, and Meta are collectively investing $300 million in the Virtual Biology Initiative to create the technologies and multi-modal datasets needed to build predictive models of life. Max Jaderberg, President of Isomorphic Labs, emphasized that generating the data required for predictive systems biology means scaling past the limits of what any single organization can produce today, and described the initiative as building a massive, multimodal data foundation intended to push the industry toward the next significant breakthrough in biology. Pushmeet Kohli, Vice President of AI for Science at Google DeepMind and Google Cloud&#8217;s Chief Scientist, argued that the quest to build a virtual cell is one of the great collective scientific challenges and cannot be solved without open, experimental biological data at an unprecedented scale showing how living cells behave and respond to change. In his view, the investment will help create an open, standardized data commons that lays the foundation researchers worldwide need to better model biology.</p>
<p>Biohub&#8217;s own founding commitment of $500 million anchors the scientific core of the program. Of that sum, $400 million supports new technologies that expand what biologists can measure. Among them is cryo-electron tomography, an imaging technique that resolves structures at near-atomic detail inside the cell, allowing researchers to observe molecular machines in their native cellular environment rather than in isolation. Another is advanced microscopy capable of imaging millions to billions of cells in living tissue, which would make it possible to capture how entire organs respond to perturbations at single-cell resolution. The remainder funds engineering tools to build and perturb biology at molecular, cellular, tissue, and whole-organism levels — the experimental manipulations that give AI models the cause-and-effect data they need to learn how biological systems respond to change. A further $100 million funds research outside Biohub, extending the reach of the program across the wider scientific community.</p>
<p>The initiative has also drawn in leading scientific institutions and consortia with experience organizing transformative international collaborations, from the Human Genome Project onward. The Allen Institute, Broad Institute, Gladstone Institutes, the Human Cell Atlas, the Human Protein Atlas, and the Wellcome Sanger Institute have come together to help nucleate the scientific community across academia and industry around developing effective strategies to maximize the impact of virtual biology. These groups are committed to working together as part of the Virtual Biology Initiative as well as through independent efforts toward the shared goal. NVIDIA will support the initiative by leveraging accelerated computing infrastructure, domain-specific software, and technical expertise, while Renaissance Philanthropy is helping to expand funding for data generation. The breadth of the coalition reflects a growing consensus that the bottleneck in computational biology is no longer algorithms but data.</p>
<p>Biohub&#8217;s role extends beyond generating measurements to building the connective tissue that lets disparate datasets work in a unified fashion: shared standards, common identifiers, and a single point of access. Equally important, the organization is working to build the scientific community around these resources, convening researchers across institutions and disciplines, connecting complementary expertise and capabilities, and creating opportunities to define and pursue ambitious scientific questions together. The approach builds on a track record from the past decade, during which Biohub has expanded the reach and impact of measurement technologies and open datasets through projects such as Tabula Sapiens, a cross-tissue map of human cell types; OpenCell, an atlas of protein localization and interactions; and Zebrahub, a developmental atlas of the zebrafish. It has also built and maintained community data infrastructure, including CELLxGENE and the CryoET Data Portal, platforms that have become widely used resources for the single-cell and structural biology communities.</p>
<p>For the researchers involved, the payoff of a validated virtual cell could be broad and profound. Nicole Kleinstreuer, NIH Deputy Director for Program Coordination, Planning, and Strategic Initiatives, noted that by combining resources and expertise, the partnership can accelerate the development of universal cell models with sufficient biological complexity to predict how any cell responds to an intervention. The return, she suggested, could be substantially faster timelines for medical breakthroughs compared with attempting to attain the same results through laboratory experiments alone. Alex Rives, Biohub&#8217;s Head of Science, put the vision in even starker terms: an accurate predictive model of biology could dramatically accelerate scientific discovery by enabling scientists to perform experiments digitally, unlocking a far greater understanding of disease and opening completely new paths for cures. He called the creation of a virtual cell one of the most important challenges for the next era of science, one that will require coordinated data generation at national and international scale — and he extended an open invitation to the worldwide scientific community to join the project.</p>
<p>The announcement lands at a moment when AI-driven biology is moving from demonstration to application, with structure prediction and molecular design already reshaping drug discovery. What has been missing, experts in the field broadly agree, is the kind of systematic, standardized, openly available measurement data that powered the revolution in language models and, more recently, in protein modeling. By committing $1.8 billion across government, industry, philanthropy, and academia — spanning funding, data, computation, and measurement technology — the Virtual Biology Initiative is betting that the next leap in medicine will come not from a single algorithmic breakthrough but from the patient, coordinated work of measuring life at scale and making the results universally available. Whether the bet pays off will depend on execution across dozens of institutions, but the scale of the commitment signals that the race to simulate biology has officially begun.</p>
<p><strong>Subject of Research:</strong> A $1.8 billion international initiative to generate open, AI-ready biological data for predictive models of cells and disease</p>
<p><strong>Article Title:</strong> International, cross-sector collaboration commits nearly $2 billion to build foundational data for AI models to predict and treat disease</p>
<p><strong>Article References:</strong> International, cross-sector collaboration commits nearly $2 billion to build foundational data for AI models to predict and treat disease. (n.d.). <a href="https://www.eurekalert.org/news-releases/1146765" rel="noopener noreferrer">Original publication</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> Not provided</p>
<p><strong>Keywords:</strong> Biohub, Virtual Biology Initiative, virtual cell, AI in biology, Department of Energy, NIH, Google DeepMind, Isomorphic Labs, cryo-electron tomography, open data, single-cell biology, predictive models</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">244605</post-id>	</item>
		<item>
		<title>Massive Multi-Omics Atlas MOSAIC Maps the Hidden Diversity Inside Tumors</title>
		<link>https://scienmag.com/massive-multi-omics-atlas-mosaic-maps-the-hidden-diversity-inside-tumors/</link>
		
		<dc:creator><![CDATA[Nathaniel Bowman]]></dc:creator>
		<pubDate>Wed, 07 Oct 2026 06:33:25 +0000</pubDate>
				<category><![CDATA[Medicine]]></category>
		<category><![CDATA[advances in cancer genomics]]></category>
		<category><![CDATA[Artificial Intelligence]]></category>
		<category><![CDATA[cancer biomarkers]]></category>
		<category><![CDATA[cancer subpopulation characterization]]></category>
		<category><![CDATA[Genome Medicine]]></category>
		<category><![CDATA[intra-tumoral genetic diversity]]></category>
		<category><![CDATA[intra-tumoral heterogeneity]]></category>
		<category><![CDATA[large-scale cancer datasets]]></category>
		<category><![CDATA[MOSAIC]]></category>
		<category><![CDATA[multi-center cancer profiling]]></category>
		<category><![CDATA[multi-omics]]></category>
		<category><![CDATA[multi-omics cancer atlas]]></category>
		<category><![CDATA[multi-omics data integration in cancer]]></category>
		<category><![CDATA[open data]]></category>
		<category><![CDATA[precision oncology]]></category>
		<category><![CDATA[precision oncology challenges]]></category>
		<category><![CDATA[resistance mechanisms in tumors]]></category>
		<category><![CDATA[single-nuclei RNA-seq]]></category>
		<category><![CDATA[Spatial transcriptomics]]></category>
		<category><![CDATA[spatial tumor mapping]]></category>
		<category><![CDATA[tumor heterogeneity analysis]]></category>
		<category><![CDATA[tumor microenvironment]]></category>
		<category><![CDATA[tumor microenvironment profiling]]></category>
		<category><![CDATA[whole exome sequencing]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=243531</guid>

					<description><![CDATA[The MOSAIC consortium has profiled more than 2,700 tumor samples with spatial and single-cell multi-omics to build a standardized public atlas for studying intra-tumoral heterogeneity and discovering clinically relevant cancer subtypes.]]></description>
										<content:encoded><![CDATA[<p>Cancer has never been a single disease, and increasingly, oncologists are realizing that a single tumor is not a single entity either. Within one tumor mass, malignant cells can carry different genetic mutations, activate different signaling pathways, and recruit entirely different neighborhoods of immune and stromal cells. This phenomenon, known as intra-tumoral heterogeneity, is one of the central reasons why precision oncology so often falls short of its promise: a biopsy sampled from one region of a tumor may reveal a molecular profile that is dramatically different from another region just millimeters away, and a therapy designed against one dominant clone may leave resistant subpopulations untouched. A large international consortium now reports a systematic effort to confront this problem at unprecedented scale, publishing its design, early results, and first public dataset in the journal Genome Medicine.</p>
<p>The initiative, called MOSAIC (Multi-Omics Spatial Atlas in Cancer), is a multi-center clinical omics study that has profiled more than 2,700 cancer samples spanning multiple tumor types. The consortium brings together researchers from Owkin and clinical and academic partners including Gustave Roussy in France, Charité Universitätsmedizin Berlin and University Hospital Erlangen in Germany, Lausanne University Hospital in Switzerland, and the University of Pittsburgh in the United States. The study was funded by Owkin Inc. and conducted as a non-interventional, multicenter clinical protocol, registered as NCT06625203, with ethics approval at each participating site. What distinguishes MOSAIC from many previous atlas projects is not just its size but its deliberate integration of data modalities that are usually collected and analyzed in isolation.</p>
<p>Technically, each tumor sample in MOSAIC is characterized through several complementary lenses. Spatial transcriptomics captures gene expression across intact tissue sections, preserving the two-dimensional geography of the tumor and allowing researchers to measure which genes are active at which coordinates on a slide. Single-nuclei transcriptomics resolves expression at the level of individual cells, distinguishing malignant subclones from the diverse non-malignant cells of the tumor microenvironment. These are layered alongside digitized histology scans, bulk whole-transcriptome sequencing, and whole-exome sequencing, which together provide the genomic mutations, the aggregate expression landscape, and the classical pathological view of the tissue. Extensive curated clinical information, including treatment regimens and outcomes, is attached to each sample, so that molecular patterns can eventually be linked to how patients actually responded.</p>
<p>The rationale for this multi-modal design is that no single technology can capture the full complexity of a tumor. Histology reveals architecture but not molecular identity; bulk sequencing averages away the very heterogeneity researchers want to measure; single-cell methods lose spatial context; and spatial methods, until recently, lacked single-cell resolution. By generating all of these representations from the same samples under standardized protocols, MOSAIC aims to create a resource in which computational methods, including artificial intelligence, can learn to connect what a pathologist sees under the microscope with what the genome and transcriptome reveal, and ultimately to identify clinically relevant biomarkers and cancer subtypes that no single modality alone could expose.</p>
<p>Alongside the full study design, the consortium has released an initial public dataset called the MOSAIC Window, now available through the European Genome-phenome Archive under study ID EGAS50000000689. This first release covers 60 patients across five tumor types and includes quality-controlled data from spatial transcriptomics, single-nuclei RNA sequencing, bulk RNA sequencing, and whole-exome sequencing, together with curated clinical descriptors such as age, smoking status, disease stage, and survival. The consortium describes this release as a glimpse into the project&#8217;s potential, with further data releases planned as the study matures. The cohort design spans eleven major cancer types overall, including bladder, breast, colorectal, gastric, glioblastoma, head and neck squamous cell carcinoma, mesothelioma, non-small cell lung cancer, ovarian, and pancreatic cancers, sampled at baseline diagnosis, after neoadjuvant therapy, and at recurrence or progression.</p>
<p>The early analyses reported in the paper demonstrate the analytical power that comes from combining these modalities. Using the MOSAIC Window data, the researchers quantified both intra-patient and inter-patient heterogeneity, uncovering distinct malignant cell subsets within individual tumors. They then examined how these malignant subpopulations relate to their surroundings and found a correlation between the intrinsic oncogenic signaling activity of malignant cells and the colocalization of specific tumor microenvironment cell types. In other words, the signaling programs running inside cancer cells appear to be associated with which immune and stromal cells gather around them, a link that could help explain why some tumor regions are immunologically cold and others inflamed, and why immunotherapy can succeed in one part of a tumor while failing elsewhere.</p>
<p>To make these concepts concrete, the paper presents four diverse case studies, each illustrating a different way that integrated multi-omics data can decipher heterogeneity. In these examples, the team used computational tools such as differential expression analysis and gene set enrichment analysis to characterize malignant cell clusters identified from single-nuclei data, and quantified pathway activities with methods like PROGENy and GSVA to compare signaling states across subpopulations. Statistical measures, including variance analyses and Pearson correlations between spot-level cellular fractions, were used to test whether the associations between malignant signaling and microenvironment composition were robust across samples and indications. The case studies span different tumor types and different biological questions, collectively highlighting how the same standardized data framework can support a wide range of investigations.</p>
<p>The technical rigor behind the resource is considerable. Supplementary documentation accompanying the paper details quality metrics for every modality, from median gene counts per spatial spot and read-mapping percentages in transcriptomics to sequencing depth, contamination estimates, and GC content in whole-exome data. A detailed data management plan describes how samples are pseudonymized, how data flows across modalities are linked to a common study identifier, and how the resource adheres to FAIR principles, making the data findable, accessible, interoperable, and reusable. The clinical study protocol, covering objectives, endpoints, eligibility criteria, and the statistical analysis plan, is also published, giving the community full visibility into how the resource was constructed.</p>
<p>For the field of computational oncology, the significance of MOSAIC lies partly in addressing a persistent bottleneck: the scarcity of large, consistent, well-annotated datasets. Machine learning models for histology and multi-omics integration have advanced rapidly, but they are often trained on small or inconsistently processed cohorts, limiting their generalizability and clinical translation. A standardized atlas of thousands of tumors, each with matched spatial, single-cell, genomic, histological, and clinical data, provides exactly the kind of training substrate that multimodal artificial intelligence approaches require. The consortium explicitly states its aim to use artificial intelligence and other computational approaches to integrate all data modalities and extract clinically actionable insights, positioning MOSAIC as infrastructure for the next generation of biomarker discovery.</p>
<p>There are, of course, caveats worth noting. The study is funded entirely by a single company, Owkin, and several authors are current or former employees holding shares or stock options in the company, alongside a range of disclosed consulting relationships with pharmaceutical and diagnostics firms at the academic sites. The paper is published as a shared early version, citable with a permanent DOI and subject to further editorial updates. And while the MOSAIC Window release demonstrates feasibility and analytical promise, the clinical value of the resource will ultimately depend on whether the biomarkers and subtypes it reveals can be validated prospectively and shown to improve patient outcomes. Still, with more than 2,700 samples profiled, an open first dataset already in researchers&#8217; hands, and a standardized framework linking molecular, spatial, and clinical dimensions of cancer, MOSAIC represents one of the most ambitious attempts yet to map the internal diversity of tumors at scale, and a foundation on which the precision oncology community can build for years to come.</p>
<p><strong>Subject of Research:</strong> Multi-omics characterization of intra-tumoral heterogeneity using spatial and single-cell profiling in cancer</p>
<p><strong>Article Title:</strong> MOSAIC: Intra-tumoral heterogeneity characterization through large-scale spatial and cell-resolved multi-omics profiling</p>
<p><strong>Article References:</strong> MOSAIC Consortium, Cornish, A. J., Bayard, Q., Karabajakian, A., Madissoon, E., Ferrarini, G., Youssef, A., Badoual, C., de Leval, L., Dressman, D., Durand, E. Y., Erber, R., Florian, S., Garberis, I., Haignere, C., Homicsko, K., Keilholz, U., Lee, A. V., Lehar, J., &#8230; Hoffmann, C. (2026). MOSAIC: Intra-tumoral heterogeneity characterization through large-scale spatial and cell-resolved multi-omics profiling. <em>Genome Medicine</em>. <a href="https://doi.org/10.1186/s13073-026-01791-y" rel="noopener noreferrer">https://doi.org/10.1186/s13073-026-01791-y</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1186/s13073-026-01791-y" rel="noopener noreferrer">10.1186/s13073-026-01791-y</a></p>
<p><strong>Keywords:</strong> MOSAIC, intra-tumoral heterogeneity, spatial transcriptomics, single-nuclei RNA-seq, multi-omics, tumor microenvironment, precision oncology, whole-exome sequencing, cancer biomarkers, artificial intelligence, Genome Medicine, open data</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">243531</post-id>	</item>
		<item>
		<title>Decade After Conservation Thinning, Italian Beech Forests Reveal Their Biodiversity Secrets</title>
		<link>https://scienmag.com/decade-after-conservation-thinning-italian-beech-forests-reveal-their-biodiversity-secrets/</link>
		
		<dc:creator><![CDATA[Margaret Porter]]></dc:creator>
		<pubDate>Sun, 04 Oct 2026 09:59:17 +0000</pubDate>
				<category><![CDATA[Agriculture]]></category>
		<category><![CDATA[and fungi data in forest studies]]></category>
		<category><![CDATA[Apennine mountain forests in Italy]]></category>
		<category><![CDATA[Apennines]]></category>
		<category><![CDATA[beech forests]]></category>
		<category><![CDATA[biodiversity]]></category>
		<category><![CDATA[Biodiversity assessment in European beech forests]]></category>
		<category><![CDATA[Conservation thinning impact on forest ecosystems]]></category>
		<category><![CDATA[deadwood]]></category>
		<category><![CDATA[epiphytic lichens]]></category>
		<category><![CDATA[European forest structure and species diversity]]></category>
		<category><![CDATA[Forest habitat complexity and species richness]]></category>
		<category><![CDATA[forest management]]></category>
		<category><![CDATA[Forest management strategies for biodiversity conservation]]></category>
		<category><![CDATA[forest structure]]></category>
		<category><![CDATA[Gran Sasso]]></category>
		<category><![CDATA[Integration of plant]]></category>
		<category><![CDATA[lichen]]></category>
		<category><![CDATA[LIFE+ FAGUS project ecological outcomes]]></category>
		<category><![CDATA[Long-term effects of forest management on biodiversity]]></category>
		<category><![CDATA[Monitoring forest biodiversity after conservation interventions]]></category>
		<category><![CDATA[Natura 2000]]></category>
		<category><![CDATA[open data]]></category>
		<category><![CDATA[role]]></category>
		<category><![CDATA[saproxylic fungi]]></category>
		<category><![CDATA[Saproxylic fungi and epiphytic lichens in managed forests]]></category>
		<category><![CDATA[vascular plants]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=234566</guid>

					<description><![CDATA[An open dataset from Italy's Gran Sasso and Monti della Laga National Park simultaneously documents forest structure, plants, lichens, and deadwood fungi in beech forests a decade after conservation-oriented management.]]></description>
										<content:encoded><![CDATA[<p>Deep in the Apennine mountains of central Italy, European beech forests that were carefully reshaped by conservationists a decade ago are now offering one of the most detailed portraits yet of how forest management can steer biodiversity. A large team of Italian researchers has published an open dataset that simultaneously documents forest structure and the diversity of three very different groups of organisms—vascular plants, epiphytic lichens, and saproxylic fungi—across nineteen sampling units within the Gran Sasso and Monti della Laga National Park. The work, published as a data paper in the journal Plant Biosystems, captures the state of these forests in July 2024, roughly ten years after conservation-oriented interventions were carried out there under the LIFE+ FAGUS project.</p>
<p>The significance of the study lies in its integrated design. European forests bear the imprint of centuries of human use, from timber harvesting to livestock grazing, and this long history has often homogenized stand structure, simplifying habitats and eroding the diversity of the biological communities they support. In recent decades, forest management has shifted toward imitating natural dynamics and promoting structural complexity, with the aim of sustaining both biodiversity and ecosystem services. Yet most previous studies have examined only trees or vascular plants, overlooking groups such as lichens, fungi, and invertebrates that may respond very differently to the same silvicultural choices. By sampling multiple taxonomic groups in the same plots at the same time, alongside detailed measurements of trees and deadwood, the new dataset directly addresses that gap.</p>
<p>The interventions themselves, performed in 2013 and 2014, were designed to increase structural complexity in beech-dominated stands. They included the creation of canopy gaps, thinning of dense regeneration, retention of deadwood, and the enhancement of habitat trees—features that collectively alter light availability, humidity, and substrate diversity within the forest. Similar sampling of the same taxonomic groups had been carried out at the same sites in 2013, before the interventions, which makes the new data a valuable temporal reference point for researchers seeking to quantify how forests change after management for conservation.</p>
<p>The fieldwork was conducted over four days in July 2024 by the Ecology, Lichenology and Mycology Group of the Società Botanica Italiana, spanning three sites: Prati di Tivo and Venaquaro, both belonging to the priority habitat type 9210* of Apennine beech forests with Taxus and Ilex, and Incodaro, representing habitat 9220*, the Apennine beech forests associated with silver fir. These habitat types, protected under the European Habitats Directive, are fragmented along the Apennine chain and are considered conservation priorities, making the park&#8217;s Natura 2000 designation an important backdrop for the work.</p>
<p>Each sampling unit followed a rigorous, nested protocol. Vascular plants were surveyed within circular plots of twenty meters radius, divided into four quadrants along the cardinal directions, with the cover of every species estimated separately in the herbaceous, shrub, and tree layers. Epiphytic lichens were recorded on the three suitable living beech trees nearest to each plot center, using standardized frames of five vertically arranged quadrats placed on each trunk at one meter height in the four cardinal directions—a method derived from the European guideline for mapping lichen diversity as an indicator of environmental stress. Saproxylic fungi, those that live on dead wood, were surveyed within a nested thirteen-meter-radius subunit, with every sporocarp larger than one millimeter recorded on all standing and lying deadwood fragments exceeding ten centimeters in diameter.</p>
<p>Forest structure was measured with equal care. Trees were inventoried within three concentric subunits with different diameter thresholds, and for each tree the researchers recorded species, diameter at breast height, height, vitality, regeneration origin, and social position within the canopy. Lying deadwood was classified as logs or stumps, measured for dimensions, and assigned to one of five decay classes ranging from freshly fallen wood with intact bark to soft, powdery remnants fully in contact with the ground. Across the dataset, tree diameters ranged from 2.5 to 112 centimeters with a median of 20 centimeters, while tree heights spanned 2 to 42 meters. Deadwood diameters reached up to 110 centimeters, and the longest log measured an impressive 18 meters.</p>
<p>The biological results reveal a clear hierarchy among the taxonomic groups. Vascular plants were by far the richest, with 191 entities belonging to 132 genera recorded across all plots, concentrated overwhelmingly in the understorey. Species richness was highest at Venaquaro with 133 species, followed closely by Prati di Tivo with 132, while Incodaro hosted 94. European beech dominated the tree and shrub layers, accompanied by silver fir, yew, and holly, while the herb layer featured widespread species such as Viola reichenbachiana, Rubus hirtus, Daphne laureola, Sanicula europaea, and the ghostly chlorophyll-free orchid Neottia nidus-avis. Epiphytic lichens came second, with 49 entities in 31 genera, peaking at Prati di Tivo with 34 species; the most frequent were Lecidella elaeochroma and Glaucomaria carpinea, found in every sampling unit. Saproxylic fungi contributed 46 entities in 33 genera, with Incodaro the richest site at 26 species, and wood-inhabiting ascomycetes such as Xylaria hypoxylon, Nemania serpens, and Biscogniauxia nummularia among the most common.</p>
<p>Notably, the three sites show contrasting deadwood profiles that hint at different management legacies. Logs outnumber stumps roughly two to one across the whole dataset, but the pattern varies sharply: Incodaro holds about five times as many logs as stumps, Prati di Tivo shows an even balance, and Venaquaro has roughly twice as many stumps as logs. Because saproxylic fungi depend entirely on dead wood for their life cycles, and lichens respond to bark texture, light, and humidity shaped by canopy structure, such structural differences are precisely the kind of variables that can explain why different groups flourish in different places. The dataset makes it possible to test these relationships quantitatively rather than anecdotally.</p>
<p>The authors are candid about the limitations of their design. July is not the optimal season for surveying saproxylic fungi in the Mediterranean region, since many fruiting bodies appear in autumn, but the choice preserved simultaneity with the plant and lichen sampling and reflected practical constraints of relocating the units in mountainous terrain. A small number of specimens across all three groups could not be identified to species level and are retained in the dataset as undetermined records at genus or family level, with the morphological reasoning for each decision documented. This transparency, the researchers argue, is essential for anyone hoping to reuse the data responsibly.</p>
<p>What elevates the work beyond a regional inventory is its commitment to open science and comparability. The full dataset is freely available on Zenodo, complete with example R scripts, and the sampling design adheres to a published European handbook for multi-taxon biodiversity studies in forests. The authors emphasize that sharing harmonized datasets and protocols is critical for advancing knowledge of how forest management affects biodiversity, for running comparative studies across space and time, and for closing persistent gaps in our understanding of the structure–biodiversity relationship. As European policy pushes forests toward multifunctionality—balancing carbon storage, timber production, and conservation—datasets like this one, born from a decade of patient stewardship in the Apennines, provide the empirical scaffolding needed to decide which silvicultural interventions genuinely deliver biodiversity gains, and for which organisms.</p>
<p><strong>Subject of Research:</strong> Multi-taxon biodiversity and forest structure in conservation-managed Apennine beech forests</p>
<p><strong>Article Title:</strong> Forest structure and multi-taxonomic diversity in Apennine beech forests subjected to conservation-oriented interventions</p>
<p><strong>Article References:</strong> Legnaro Diamanti, M., Cangelmi, G., De Benedictis, L. L. M., Balducci, L., Benesperi, R., Bianchi, E., Carli, E., Ceci, A., Chianucci, F., Ciaschetti, G., De Simone, L., De Toma, A., Di Pietro, F., Fanfarillo, E., Fiaschi, T., Francioni, M., Girometta, C. E., Leonardi, M., Loppi, S., &#8230; Burrascano, S. (2026). Forest structure and multi-taxonomic diversity in Apennine beech forests subjected to conservation-oriented interventions. <em>Plant Biosystems, 160</em>(4), Article 224. <a href="https://doi.org/10.1007/s44473-026-00238-x" rel="noopener noreferrer">https://doi.org/10.1007/s44473-026-00238-x</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s44473-026-00238-x" rel="noopener noreferrer">10.1007/s44473-026-00238-x</a></p>
<p><strong>Keywords:</strong> beech forests, forest structure, biodiversity, vascular plants, epiphytic lichens, saproxylic fungi, deadwood, forest management, Apennines, Gran Sasso, Natura 2000, open data</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">234566</post-id>	</item>
		<item>
		<title>Massive New Database Reveals When We Learn Every Word of Russian</title>
		<link>https://scienmag.com/massive-new-database-reveals-when-we-learn-every-word-of-russian/</link>
		
		<dc:creator><![CDATA[Glenn Wilkins]]></dc:creator>
		<pubDate>Fri, 02 Oct 2026 04:22:15 +0000</pubDate>
				<category><![CDATA[Psychology & Psychiatry]]></category>
		<category><![CDATA[age of acquisition]]></category>
		<category><![CDATA[Behavior Research Methods]]></category>
		<category><![CDATA[childhood language acquisition]]></category>
		<category><![CDATA[cognitive psychology]]></category>
		<category><![CDATA[cognitive psychology of language]]></category>
		<category><![CDATA[comprehensive language acquisition resources]]></category>
		<category><![CDATA[impact of early word learning]]></category>
		<category><![CDATA[language acquisition]]></category>
		<category><![CDATA[language processing and reaction times]]></category>
		<category><![CDATA[large-scale Russian vocabulary database]]></category>
		<category><![CDATA[lexical decision task research]]></category>
		<category><![CDATA[lexical norms]]></category>
		<category><![CDATA[lexical processing]]></category>
		<category><![CDATA[longitudinal language learning studies]]></category>
		<category><![CDATA[open data]]></category>
		<category><![CDATA[psycholinguistics]]></category>
		<category><![CDATA[retrospective memory accuracy]]></category>
		<category><![CDATA[Russian language]]></category>
		<category><![CDATA[Russian language vocabulary development]]></category>
		<category><![CDATA[vocabulary development]]></category>
		<category><![CDATA[vocabulary validation in children]]></category>
		<category><![CDATA[word frequency]]></category>
		<category><![CDATA[word recognition]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=225650</guid>

					<description><![CDATA[Researchers have collected and validated age-of-acquisition ratings for 30,849 Russian words from over 2,200 adults, creating the largest such database for any Slavic language and showing that adults' memories of when they learned words closely match children's actual vocabulary development.]]></description>
										<content:encoded><![CDATA[<p>When did you first understand the word &#8220;dragon&#8221;? For most people the answer comes easily, even decades later, and it turns out that these retrospective memories are far more accurate than they might seem. A team of Russian psychologists has now harnessed this phenomenon on an unprecedented scale, collecting age-of-acquisition estimates for 30,849 Russian words from more than 2,200 adult speakers and validating them against the actual vocabulary knowledge of schoolchildren. The result, published in Behavior Research Methods, is the largest and most comprehensive age-of-acquisition resource ever assembled for the Russian language, and one of the largest for any language in the world.</p>
<p>Age of acquisition, usually abbreviated AoA, refers to the age at which a person first learns a word in their native language. For half a century, cognitive psychologists have documented that this seemingly simple variable exerts a powerful influence on how the mind processes language. Pictures of objects whose names were learned early in life are named faster, words acquired early are read faster, and in the lexical decision task, where participants judge whether a letter string is a real word, both reaction times and error rates depend on when the word entered a person&#8217;s vocabulary. The effect even extends to semantic tasks: words learned early generate more associates in word-association experiments and are categorized more quickly, although interestingly the effect does not appear in picture categorization. Because AoA shapes performance across such a broad range of tasks, researchers treat it as one of the essential psycholinguistic variables, alongside word frequency, length, and concreteness, that must be controlled when designing experiments on reading and lexical processing.</p>
<p>The standard way to measure AoA is simply to ask adults. Participants estimate, in years, the age at which they believe they learned each word, producing what researchers call subjective AoA ratings. Such norms now exist for Chinese, Croatian, Dutch, English, French, German, Icelandic, Italian, Japanese, Portuguese, Spanish, Turkish, and many other languages. Methodological preferences have shifted over time: early studies used coarse rating scales, such as a seven-point scale in which the lowest point covered ages zero to two and the highest covered thirteen and older. More recent work favors asking respondents to report a specific age in years, a technique that avoids artificially restricting the response range and is easier for participants to use. Cross-language comparisons show that the order in which words are learned is remarkably consistent across cultures, with reliability-adjusted correlations between languages reaching as high as .96 between Polish and Slovak.</p>
<p>Until now, Russian researchers have worked with a patchwork of small datasets. Existing norms covered only a few hundred items each: 260 nouns denoting pictured objects, 190 and 375 verbs in separate studies, 696 nouns, 414 verbs, 475 adjectives, and 506 concrete and abstract nouns. Most of these ratings were never validated against children&#8217;s actual word knowledge. The new dataset changes that picture dramatically, expanding the lexical coverage of Russian AoA norms by two orders of magnitude and providing, for the first time, a resource comparable in scope to the major English and Dutch norms that each cover roughly 30,000 words.</p>
<p>The scale of the data collection effort was considerable. Approximately 3,000 participants took part, recruited both in person, primarily first-year university students and faculty members, and online through the Yandex.Toloka crowdsourcing platform, where respondents received about $2.50 in compensation. After a battery of quality checks, 2,201 protocols were retained, roughly 73 percent of the total. The respondents, whose mean age was 27.2 years, rated words organized into 103 lists of 300 target words each. Each list also contained ten calibrator words, included to show respondents the possible range of ratings, and thirty control words with previously established ratings, used to verify the validity of each person&#8217;s responses. Data collection proceeded in three annual phases between 2022 and 2024, with ratings gathered for 7,500 words in the first year, 16,500 in the second, and 6,900 in the third.</p>
<p>The stimulus set was drawn mainly from a standard frequency dictionary of contemporary Russian, supplemented by words from category norms and other items familiar to modern speakers. The researchers deliberately sampled across grammatical categories: within each list, 134 items were nouns, 76 verbs, 70 adjectives, 17 adverbs, and 3 belonged to other parts of speech, mirroring the distribution in the source dictionary. Words from the same derivational family, such as &#8220;cunning&#8221; and &#8220;craftiness,&#8221; were generally kept in different lists to avoid inflating similarity between ratings. The final set ranged from one to 24 letters in length, with words presented in their lemmatized form except where a plural is more common in actual usage.</p>
<p>Ensuring data quality required elaborate screening. Each protocol was checked for autocorrelation, meaning suspiciously smooth sequences of ratings, and for its correlation with the known ratings of the control words. Protocols were excluded if a respondent&#8217;s mean rating deviated more than two standard deviations from the list average, if the respondent claimed to have learned every word after age four, or if more than sixty words were marked as unknown. In the online sessions, individual responses with latencies shorter than 800 milliseconds were also discarded, and all data were additionally inspected manually. The result was remarkably consistent: a bootstrap split-half analysis, in which respondents were repeatedly divided into two random halves and their mean ratings correlated, yielded a mean raw correlation of .895 and a Spearman-Brown-corrected reliability of .944, indicating that adults provide highly stable estimates even across thousands of unfamiliar-seeming items.</p>
<p>The most striking aspect of the study is its validation strategy. Rather than relying solely on internal consistency, the researchers tested whether the adult ratings actually predict when children learn words. They constructed multiple-choice vocabulary tests in which students had to select, from five options, the word or phrase best matching a target word, following careful design principles such as ensuring the correct answer was simpler than the target and that distractors were semantically balanced. In the grade-level validation, a 300-item test was administered to students in grades 2, 4, 6, 8, and 10 in schools in the Moscow region and the city of Kursk. A word was considered known at a given grade if at least 75 percent of students answered correctly and students in all higher grades did as well. The correlation between these objective, grade-based estimates and the adult subjective ratings was .805, a remarkably strong correspondence showing that adults&#8217; retrospective judgments closely track the real developmental timeline of word learning.</p>
<p>A second procedure, within-grade validation, examined whether the ratings predict accuracy among children of the same age. Two rounds of testing, eight to ten months apart, involved 700 schoolchildren in total. Across every grade level, the correlations between subjective AoA and the proportion of correct answers were consistently negative: the later a word was estimated to be acquired, the lower the accuracy on its test item, both among second-graders and among eighth-graders. After statistical corrections for measurement error and range restriction, these correlations increased further. The ratings also correlated .698 with objective AoA values obtained by asking children of different ages to name pictured objects, and correlations with the seven previous Russian datasets ranged from .68 to .91, providing converging evidence from every available angle.</p>
<p>The norms also reproduce the expected relationships with other psycholinguistic variables. Later-acquired words tend to be longer, with correlations of about .28 with word length, and less frequent, with correlations between -.38 and -.45 depending on the frequency source. They are also less concrete, less imageable, and less familiar, with the strongest association observed for imageability. Intriguingly, the study also revealed respondent-level patterns: older adults gave slightly later AoA estimates and marked far fewer words as unknown, suggesting that vocabulary knowledge continues to expand across the lifespan and that our memories of word learning may shift with our own age. With the full dataset now freely available on the Open Science Framework, researchers in psycholinguistics, reading development, computational modeling, and language education gain a powerful new tool, and cross-linguistic comparisons of how children build their lexicons become possible for Russian at a scale never before achievable.</p>
<p><strong>Subject of Research:</strong> Subjective age-of-acquisition norms for 30,849 Russian words and their validation against schoolchildren&#x27;s vocabulary knowledge</p>
<p><strong>Article Title:</strong> Subjective age of acquisition norms for 30,849 Russian words</p>
<p><strong>Article References:</strong> Subjective age of acquisition norms for 30,849 Russian words. (n.d.). <a href="https://doi.org/10.3758/s13428-026-03138-2" rel="noopener noreferrer">https://doi.org/10.3758/s13428-026-03138-2</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.3758/s13428-026-03138-2" rel="noopener noreferrer">10.3758/s13428-026-03138-2</a></p>
<p><strong>Keywords:</strong> age of acquisition, psycholinguistics, Russian language, lexical norms, vocabulary development, word frequency, lexical processing, language acquisition, Behavior Research Methods, open data, word recognition, cognitive psychology</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">225650</post-id>	</item>
		<item>
		<title>Kenya&#8217;s 2019 Census Gets an Open-Data Makeover That Other Nations Can Copy</title>
		<link>https://scienmag.com/kenyas-2019-census-gets-an-open-data-makeover-that-other-nations-can-copy/</link>
		
		<dc:creator><![CDATA[Courtney Benton]]></dc:creator>
		<pubDate>Sat, 26 Sep 2026 11:43:42 +0000</pubDate>
				<category><![CDATA[Social Science]]></category>
		<category><![CDATA[API-accessible census data]]></category>
		<category><![CDATA[capacity building]]></category>
		<category><![CDATA[Census data visualization and exploration]]></category>
		<category><![CDATA[CSPro]]></category>
		<category><![CDATA[data APIs]]></category>
		<category><![CDATA[Data democratization in low-income countries]]></category>
		<category><![CDATA[Data transparency in national statistics]]></category>
		<category><![CDATA[Evidence-based policymaking tools]]></category>
		<category><![CDATA[International collaboration in statistical projects]]></category>
		<category><![CDATA[Istat]]></category>
		<category><![CDATA[JSON-stat]]></category>
		<category><![CDATA[Kenya 2019 census]]></category>
		<category><![CDATA[KNBS]]></category>
		<category><![CDATA[National statistical data dissemination]]></category>
		<category><![CDATA[official statistics]]></category>
		<category><![CDATA[open data]]></category>
		<category><![CDATA[Open data in African countries]]></category>
		<category><![CDATA[Open data models for government data]]></category>
		<category><![CDATA[open-source software]]></category>
		<category><![CDATA[Open-source tools for data sharing]]></category>
		<category><![CDATA[SDGs]]></category>
		<category><![CDATA[statistical dissemination]]></category>
		<category><![CDATA[Sustainable development data]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=216247</guid>

					<description><![CDATA[A Kenya-Italy cooperation project transformed the 2019 census into an open, API-accessible data resource built entirely on free, sustainable open-source tools.]]></description>
										<content:encoded><![CDATA[<p>When Kenya counted its people in 2019, the exercise produced one of the most detailed portraits of any African nation ever assembled: every household, every settlement, every age cohort captured in a single national snapshot. Yet for years afterward, much of that information sat locked inside formats that only specialists could open. A new study published in Social Indicators Research describes how a cooperation project between the Kenya National Bureau of Statistics (KNBS) and the Italian National Institute of Statistics (Istat), funded by the Italian Agency for Development Cooperation (AICS), dismantled those barriers. The result is a structured, API-accessible, interactively explorable dissemination pipeline for the 2019 Population and Housing Census, built entirely on open-source tools, delivered at no licensing cost, and owned outright by the Kenyan institution that produced the data. The authors, a team of Istat statisticians led by Mauro Bruno, argue that the approach offers a replicable and sustainable model for national statistical offices across the low- and middle-income world.</p>
<p>The problem the project tackled is deceptively simple to state and notoriously hard to solve. Timely, disaggregated, and accessible statistical data are essential for evidence-based policymaking, from allocating health budgets to planning school construction and tracking the Sustainable Development Goals. But many national statistical offices, particularly in resource-constrained settings, still rely on static dissemination: PDF tables, spreadsheets buried on websites, or summary publications that resist reuse. Open Data Watch&#8217;s 2021 review of data portals in IDA-eligible countries documented how widespread this limitation remains, and the United Nations Statistics Division&#8217;s handbook on dissemination makes clear that the gap between collecting data and making it genuinely usable is one of the persistent weak links in official statistics. A census that cannot be queried by a district health officer, a journalist, or a researcher is a census whose value decays with every year it remains locked away.</p>
<p>What distinguishes the Kenyan project is its deliberate philosophy of restraint. Rather than proposing to replace comprehensive, enterprise-grade statistical platforms, the Istat team set out to close specific information gaps using lightweight, open-source tools that fit within the skills and resources already available inside KNBS. This is a methodological stance as much as a technical one. Large commercial dissemination systems can cost substantial sums in licensing and maintenance, and when donor-funded deployments outstrip local capacity to operate them, they tend to fall silent after the consultants leave. The Kenyan model inverts that logic: it starts from the tools KNBS staff already used daily and builds outward, ensuring that every component added to the pipeline could be understood, maintained, and extended by the institution itself. Full institutional ownership, the authors emphasize, was not an afterthought but a design requirement.</p>
<p>The technical architecture rewards a closer look, because its components are individually modest and collectively powerful. Kenya&#8217;s census processing, like that of many countries, relies on CSPro, the census and survey processing system developed by the U.S. Census Bureau. Istat&#8217;s cooperation teams had previously built two open-source bridges for this ecosystem: CSPro2csv, which transforms CSPro tabulations into structured data, and CSPro2sql, which migrates CSPro microdata into a relational database. Both are publicly available on GitHub under the IstatCooperation organization. In the Kenyan pipeline, these tools convert census outputs into machine-readable structures that can then be exposed through modern dissemination interfaces. The project also drew on JSON-stat, a lightweight standard for statistical data dissemination designed to make multidimensional data easy to transmit and consume, alongside established frameworks such as the W3C&#8217;s RDF Data Cube vocabulary and the SDMX technical standards that underpin much of international statistical exchange.</p>
<p>On top of this data layer sits the user-facing deliverable: the KNBS Open Data Browser, an interactive web application whose source code is published in the knbs-databrowser repository. The system allows policymakers, researchers, and civil society users to explore census indicators interactively, filter by geographic and demographic dimensions, and retrieve data in structured, reusable form through an application programming interface. That API layer matters as much as the visual interface. A dashboard serves human curiosity; an API serves the entire ecosystem of downstream tools, from statistical software and business intelligence packages to machine-learning pipelines and mobile applications. By making census data available in structured form through a documented interface, the pipeline turns a one-off publication into live infrastructure. The live deployment can be consulted at data.knbs.or.ke, where the 2019 census results are now openly explorable by anyone with an internet connection.</p>
<p>The project did not emerge from a vacuum. It is the product of a long-running Italian cooperation program in statistical capacity building, with documented precedents in Ethiopia, where Istat-supported teams developed metadata-driven monitoring for census data collection, and in Palestine, where similar toolchains supported the country&#8217;s pathway toward the 2030 Agenda for Sustainable Development. Earlier work published by the same Istat group, including a 2025 chapter on open-source innovation in statistical data dissemination, framed the Kenyan case as a proof of concept for a broader approach. The lineage also includes internal Istat modernization efforts such as the Stat2015 programme, and concrete implementations of the Common Statistical Production Architecture, the international blueprint for interoperable statistical production systems. In Kenya, that accumulated experience was distilled into a pipeline small enough to be sustainable and robust enough to carry a national census.</p>
<p>The stakes of this kind of work extend well beyond technical convenience. Morten Jerven&#8217;s influential book Poor Numbers argued that weak African statistical systems distort development policy itself, because decisions worth billions of dollars rest on indicators that are outdated, fragmented, or inaccessible. The Cape Town Global Action Plan for Sustainable Development Data, adopted by the UN Statistical Commission in 2017, explicitly calls for modernized dissemination and better use of open data, and PARIS21&#8217;s 2022 report on the digital transformation of national statistical offices catalogued both the ambition and the shortfall. The World Bank&#8217;s 2023 assessment of the next generation of statistical capacity reached a similar conclusion: transformation fails when it is imposed from outside rather than grown from within. The Kenyan model speaks directly to that diagnosis, because its central design principle is that capacity building succeeds when the beneficiary institution can run, repair, and evolve the system with its own staff.</p>
<p>There is also a governance dimension that the authors take seriously. Open government data research, including the systematic literature review by Wirtz, Weyerer, and Rösch published in Government Information Quarterly, has shown that publication alone does not guarantee use; data must be discoverable, documented, and machine-readable to generate civic and economic value. The Kenyan pipeline addresses this by attaching metadata management to dissemination, following the unified approach to statistical metadata articulated by Signore, Scanu, and Brancato in the Journal of Official Statistics. Structured metadata mean that a user downloading a census table knows what the numbers measure, how they were collected, and what their limitations are. The project also aligns with European interoperability profiles such as StatDCAT-AP, used for describing statistical datasets in data portals, which eases the path toward cross-national discoverability of Kenyan official statistics.</p>
<p>The sustainability argument deserves particular attention, because it is where many digital development projects quietly fail. The Kenyan system carries no licensing cost, which removes the recurring financial burden that so often kills donor-built platforms when funding cycles end. Its components are open source, which means the code can be audited, forked, and reused by any other statistical office facing the same constraints, and the GitHub repositories make that reuse practical rather than theoretical. Its architecture is deliberately lightweight, so it runs on the infrastructure and skills a national office actually possesses rather than on an idealized IT department it does not. And it is embedded in Kenya&#8217;s own strategic planning: the second Kenya Strategy for the Development of Statistics, covering 2023/24 to 2027/28, positions data dissemination modernization within the country&#8217;s broader statistical agenda, giving the pipeline an institutional home rather than a donor deadline.</p>
<p>For other national statistical offices contemplating a similar path, the Kenyan experience offers a concrete checklist rather than an abstract aspiration. Start from the tools already in daily use, in this case CSPro, and build connectors rather than replacements. Adopt lightweight, open standards such as JSON-stat and SDMX so that data remain portable across systems and decades. Publish the dissemination interface as an interactive browser backed by a real API, and release the source code so that peers can replicate it. Fund the work through cooperation frameworks, as AICS did, but structure it so that ownership and operational responsibility transfer fully to the national institution. Measure success not by the launch event but by whether researchers, journalists, and civil society can independently pull disaggregated census data years later. By those measures, Kenya&#8217;s 2019 census has become more than a count of its population; it is now a working demonstration that democratizing official statistics is neither expensive nor dependent on proprietary technology, and that the tools to open up a nation&#8217;s data may already be sitting, free and open source, waiting to be assembled.</p>
<p><strong>Subject of Research:</strong> Open-source dissemination of Kenya&#x27;s 2019 census data through statistical capacity building</p>
<p><strong>Article Title:</strong> Democratizing Access to Official Statistics Data: A Sustainable Model from Kenya’s 2019 Census</p>
<p><strong>Article References:</strong> Bruno, M., Grassia, M., Quaresima, C., Patruno, V., &amp; Zindato, D. (2026). Democratizing Access to Official Statistics Data: A Sustainable Model from Kenya’s 2019 Census. <em>Social Indicators Research, 184</em>(3), Article 51. <a href="https://doi.org/10.1007/s11205-026-03943-4" rel="noopener noreferrer">https://doi.org/10.1007/s11205-026-03943-4</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s11205-026-03943-4" rel="noopener noreferrer">10.1007/s11205-026-03943-4</a></p>
<p><strong>Keywords:</strong> official statistics, open data, Kenya 2019 census, KNBS, Istat, open-source software, statistical dissemination, capacity building, CSPro, JSON-stat, SDGs, data APIs</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">216247</post-id>	</item>
	</channel>
</rss>
