<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>transformation of written weather chronicles into climate data &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/transformation-of-written-weather-chronicles-into-climate-data/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Thu, 08 Oct 2026 12:34:05 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.3</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>transformation of written weather chronicles into climate data &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>AI reads 800 years of storm chronicles to reconstruct Europe&#8217;s thunderstorm past</title>
		<link>https://scienmag.com/ai-reads-800-years-of-storm-chronicles-to-reconstruct-europes-thunderstorm-past/</link>
		
		<dc:creator><![CDATA[Sloane Callahan]]></dc:creator>
		<pubDate>Thu, 08 Oct 2026 12:34:05 +0000</pubDate>
				<category><![CDATA[Climate]]></category>
		<category><![CDATA[AI analysis of medieval weather observations]]></category>
		<category><![CDATA[artificial intelligence in historical climatology]]></category>
		<category><![CDATA[BERT]]></category>
		<category><![CDATA[Central Europe]]></category>
		<category><![CDATA[climate database from 1000 to 1817]]></category>
		<category><![CDATA[Climate of the Past]]></category>
		<category><![CDATA[climate reconstruction]]></category>
		<category><![CDATA[convective weather]]></category>
		<category><![CDATA[digitization of historical weather reports]]></category>
		<category><![CDATA[documentary sources]]></category>
		<category><![CDATA[Ha'il]]></category>
		<category><![CDATA[historical climatology]]></category>
		<category><![CDATA[historical hailstorm and thunderstorm records]]></category>
		<category><![CDATA[Historical storm data reconstruction]]></category>
		<category><![CDATA[long-term climate change and storm patterns]]></category>
		<category><![CDATA[medieval European weather documentation]]></category>
		<category><![CDATA[natural language processing]]></category>
		<category><![CDATA[thunderstorms]]></category>
		<category><![CDATA[transformation of written weather chronicles into climate data]]></category>
		<category><![CDATA[transformer language models]]></category>
		<category><![CDATA[transformer-based language models for storm classification]]></category>
		<category><![CDATA[understanding Europe's thunderstorm history through AI]]></category>
		<category><![CDATA[University of Freiburg]]></category>
		<category><![CDATA[use of AI to analyze centuries-old climate data]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=247646</guid>

					<description><![CDATA[Researchers decoded nearly 7,000 historical Central European reports of thunderstorms and hail dating back to the year 1000 and trained AI language models that classify the events with high accuracy, revealing a seasonal signal that closely matches modern observations.]]></description>
										<content:encoded><![CDATA[<p>Long before rain gauges, barometers and weather radar, the only instruments capable of recording a violent summer hailstorm were the eyes, ears and quills of monks, chroniclers and town clerks. Now, two researchers at the University of Freiburg have shown that those centuries-old written observations can be transformed into quantitative climate data — and that modern artificial intelligence can be taught to do the decoding at scale. In a study published in the journal Climate of the Past, Franck Schätz and Rüdiger Glaser analysed nearly 7,000 historical reports of thunderstorms and hail from Central Europe, spanning more than eight centuries from the year 1000 to 1817, and used them to train transformer-based language models that classify storm intensity with remarkable reliability.</p>
<p>The corpus at the heart of the study draws mainly on HISKLID2, a historical climate database compiled since the 1980s on the research platform tambora.org, originally for reconstructing temperature and precipitation. From this collection, the researchers extracted 6,999 quotations describing convective weather events, drawn from 494 sources. The result is a dataset documenting 6,157 thunderstorm events and 2,006 hail events. Crucially, the corpus is not a cherry-picked archive of disasters: of the more than 51,000 records in the underlying database, roughly 86 percent mention neither thunderstorms nor hail, which structurally rules out a selective bias toward extreme events. Most of the material comes from chronicles, annals, historiographies and administrative records, written by chroniclers, theologians, teachers and officials — the same social profile as the educated observers who would later staff the first systematic meteorological networks of the eighteenth and nineteenth centuries.</p>
<p>Working with such texts poses a formidable linguistic challenge. The quotations span four stages of German — Middle High German, Early New High German, New High German and Contemporary German — with Latin sources translated before normalisation. The team standardised spelling and vocabulary without modernising grammar, preserving idiomatic expressions in their original form to keep the semantics authentic. Obsolete and dialect words were resolved using historical dictionaries, dates recorded in religious or regional calendars were deciphered with standard reference works, and place names were georeferenced so that the events could be mapped across the climatically diverse zones of Central Europe, from temperate oceanic to continental and Mediterranean-influenced regimes.</p>
<p>The classification scheme itself is deliberately formal. Five mutually exclusive thunderstorm classes and four hail classes were defined in advance, anchored in the modern warning systems of the German Weather Service and the Tornado and Storm Research Organisation. Each report was analysed for its associated phenomena — rain, hail, lightning, wind, snow — and for the causal chains, or impact pathways, they triggered: flooding, lightning strikes, damaged crops, destroyed buildings, and in the worst cases fatalities. Hailstone size descriptions, where available, were matched to standardised size ranges, with the critical two-centimetre threshold above which kinetic energy and damage potential rise disproportionately. Intensity classes were then assigned from this structured evidence rather than from impressionistic reading.</p>
<p>Because no independent benchmark exists for verifying a storm classification from the year 1400, the researchers built their own quality-control apparatus. Every word or phrase describing a phenomenon was assigned to one of three evidence classes: C1 for direct, measurable information such as hailstone size or wind strength; C2 for indirect damage indicators like ruined harvests or shattered roofs; and C3 for vague qualitative descriptors such as strong or terrible. From these, a four-level confidence index was derived for each quotation, ranging from no linguistic evidence at all to multiple independent confirmations. The resulting word lists, dubbed silver labels, are published in full, making the entire procedure reproducible.</p>
<p>The physical plausibility check delivered the study&#8217;s most striking result. When the classified events are aggregated by month, they produce a pronounced seasonal cycle — a spring rise, a summer peak between June and August, and a steady autumn decline — that closely mirrors modern observations from German Weather Service reference periods. Spearman rank correlations between the historical and modern seasonal patterns reach 0.66 to 0.78 for thunderstorms and 0.78 to 0.93 for hail, with even stronger agreement in the densely documented window of 1624 to 1654. The pattern also survives stratification: across the four largest language groups the seasonal profiles correlate at 0.82 to 0.94, and the agreement across source types is similarly high. The signal, in other words, is not an artefact of any particular language, genre or era. Independent reconstructions by earlier researchers, including work by Lenke in 1960 and Camuffo and colleagues in 2000, show the same summer-dominated shape, reinforcing the conclusion that the historical texts track real convective processes.</p>
<p>The data also overturn a persistent assumption in the literature: that historical sources record mainly extreme weather. In fact, the weakest thunderstorm class dominates in every month of the year, accounting for roughly 53 to 71 percent of classified events, exactly what a physically realistic frequency distribution would predict. Hail tells a subtler seasonal story, with weak events dominating winter and transitional months while severe hail takes over in summer, when convective energy is highest. One genuine anomaly is an apparent over-representation of winter thunderstorms, which the authors suggest may reflect the fact that medieval and early modern observers found thunder in winter so unusual that they felt compelled to write it down. Source type does matter — chronicles disproportionately report severe hail, while continuous weather records capture weaker events — but statistical testing shows these reporting patterns explain only a modest share of the variance, too little to distort the aggregate signal.</p>
<p>With the validated dataset in hand, the team fine-tuned a multilingual transformer model, mDeBERTa V3 Base, producing two specialised classifiers: ThunderstormBERT and HailBERT. After a two-stage hyperparameter search and stratified data splits, HailBERT achieved a macro-averaged F1 score of 0.93 with 96 percent accuracy, while ThunderstormBERT reached 0.83 with 87 percent accuracy. The gap is instructive: hail intensity can often be pinned down by concrete size descriptions, whereas the moderate thunderstorm class suffers from the vaguest language and the least clear-cut boundaries. Encouragingly, nearly all remaining errors occur between neighbouring intensity classes, indicating that the models have learned a coherent internal ordering of storm severity rather than memorising phrases. Both models converge cleanly within five epochs, show no significant overfitting, and are publicly available for other researchers to apply.</p>
<p>The implications reach well beyond historical linguistics. Instrumental and radar-based records of convective weather span only a few decades, far too short to separate natural variability from the influence of anthropogenic climate change on thunderstorms and hail. A dataset reaching back a millennium, anchored in a pre-industrial climate regime, offers exactly the long baseline needed to contextualise today&#8217;s storm activity. The authors argue that the high quality of these old records is no accident: in agrarian societies, a devastating hailstorm was an existential threat, and it was documented with correspondingly care. The classification method is designed to be language-independent, though it must be rebuilt corpus by corpus for each new language, and the team&#8217;s next target is comparable archives in French and English. If the approach transfers, millions of pages of early modern weather writing could soon become a searchable, quantified archive of extreme weather — read not by armies of archivists, but by machines trained on the careful judgement of historians.</p>
<p><strong>Subject of Research:</strong> Reconstruction of historical thunderstorm and hail activity in Central Europe from documentary sources using transformer-based language models</p>
<p><strong>Article Title:</strong> From manual classification to transformer-based language models: assessing the quality and consistency of historical convective event records</p>
<p><strong>Article References:</strong> Schätz, F., &amp; Glaser, R. (2026). From manual classification to transformer-based language models: assessing the quality and consistency of historical convective event records. <em>Climate of the Past, 22</em>(10), 1833-1861. <a href="https://doi.org/10.5194/cp-22-1833-2026" rel="noopener noreferrer">https://doi.org/10.5194/cp-22-1833-2026</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.5194/cp-22-1833-2026" rel="noopener noreferrer">10.5194/cp-22-1833-2026</a></p>
<p><strong>Keywords:</strong> historical climatology, thunderstorms, hail, transformer language models, BERT, documentary sources, Central Europe, climate reconstruction, natural language processing, convective weather, Climate of the Past, University of Freiburg</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">247646</post-id>	</item>
	</channel>
</rss>
