<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>LightRAG &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/lightrag/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Wed, 07 Oct 2026 06:29:20 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.3</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>LightRAG &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>AI Turns Scattered Rockfall Records Into a Knowledge Graph That Answers Hazard Questions</title>
		<link>https://scienmag.com/ai-turns-scattered-rockfall-records-into-a-knowledge-graph-that-answers-hazard-questions/</link>
		
		<dc:creator><![CDATA[Violet Maxwell]]></dc:creator>
		<pubDate>Wed, 07 Oct 2026 06:29:20 +0000</pubDate>
				<category><![CDATA[Earth Science]]></category>
		<category><![CDATA[AI applications in environmental safety]]></category>
		<category><![CDATA[AI-driven geoscience research]]></category>
		<category><![CDATA[artificial intelligence in geology]]></category>
		<category><![CDATA[Chinese geological hazard studies]]></category>
		<category><![CDATA[earthquake and weather impact on rock stability]]></category>
		<category><![CDATA[GeoGPT]]></category>
		<category><![CDATA[geohazard assessment]]></category>
		<category><![CDATA[geoscience data integration]]></category>
		<category><![CDATA[hazard prediction and risk management]]></category>
		<category><![CDATA[information extraction]]></category>
		<category><![CDATA[knowledge graph]]></category>
		<category><![CDATA[knowledge graph for geohazards]]></category>
		<category><![CDATA[landslide hazards]]></category>
		<category><![CDATA[large language models]]></category>
		<category><![CDATA[LightRAG]]></category>
		<category><![CDATA[natural disaster monitoring]]></category>
		<category><![CDATA[ontology]]></category>
		<category><![CDATA[precarious rock masses]]></category>
		<category><![CDATA[question answering]]></category>
		<category><![CDATA[retrieval-augmented generation]]></category>
		<category><![CDATA[rockfall hazard assessment]]></category>
		<category><![CDATA[seismic activity and erosion analysis]]></category>
		<category><![CDATA[structured knowledge bases for geohazards]]></category>
		<category><![CDATA[Three Gorges]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=243519</guid>

					<description><![CDATA[Chinese researchers have built an ontology-guided LightRAG framework that converts fragmented precarious rock hazard records into a knowledge graph, achieving 88.36 percent F1 in extraction and outperforming general-purpose AI models on expert hazard questions.]]></description>
										<content:encoded><![CDATA[<p>Deep in the mountains above the Three Gorges reservoir region of Chongqing, China, thousands of rock faces hang in a state of uneasy equilibrium. Geologists call these formations precarious rock masses: fractured blocks of stone that have not yet fallen but are slowly being prised loose by weathering, erosion, rainfall and seismic shaking. When one collapses, the consequences can be catastrophic, particularly where roads, schools and farmland crowd the base of a slope. Now a team of Chinese researchers has shown that a carefully engineered combination of artificial intelligence tools can turn decades of fragmented investigation reports about these hazards into a structured, queryable knowledge base that answers expert-level questions with remarkable accuracy.</p>
<p>The study, published in Earth Science Informatics, was led by Ke Li of the Chongqing AI Geological and Mineral Research Institute together with colleagues from the Chongqing Bureau of Geology and Minerals Exploration and partner institutions. Their starting point was a problem familiar to anyone who has worked in applied geoscience: the knowledge needed to assess a dangerous rock slope is scattered across heterogeneous documents, including field investigation reports, engineering records and monitoring archives, each organized differently and none designed for machine reading. This fragmentation complicates three things at once, the authors note: modeling the chains of events that turn a trigger such as a storm into a disaster, extracting reliable facts from messy text, and retrieving both specific entity-level details and broader mechanism-level context when a question is asked.</p>
<p>To solve it, the team built what they call an ontology-guided LightRAG framework, a system with three interlocking components. The first is a domain ontology, a formal specification of the concepts and relationships that matter in precarious rock hazard work. Ontologies have a long pedigree in geoinformatics, tracing back to foundational work on semantic web terminology for earth and environmental science, and they serve as a shared vocabulary that both humans and machines can use. The researchers&#8217; ontology contains 31 entity categories organized across five core evolutionary stages: environmental background, triggering factors, disaster evolutionary processes, mitigation decisions and effectiveness evaluation. In effect, it encodes how geologists think about a rockfall hazard from first causes to final remediation.</p>
<p>The second component is a four-stage conversational extraction process driven by GeoGPT, an open-source geological language model fine-tuned from Qwen2.5-72B-Instruct by its original developers and deployed locally by the team without further modification. In the first stage, the system performs an ontology-aware semantic scan of a document, using a sliding window to pull in surrounding paragraphs so that cross-sentence context is preserved. In the second stage, it anchors entities and resolves co-references, mapping phrases such as &#8216;this rock mass&#8217; or a shorthand label like &#8216;W01&#8217; to a single unique identifier. The third stage involves recursive attribute backfilling, in which the model hunts through the text for missing physical parameters such as volumes, fracture orientations and stability coefficients, and the fourth applies logical verification, checking that numbers and qualitative descriptions agree before stripping away redundant modifiers and emitting structured data.</p>
<p>The authors illustrate the pipeline with a real example from Chaqi Mountain, where a precarious rock unit designated W01 sits on a second-tier cliff above a school housing roughly 4,000 students and staff. From a single passage, the system extracts the unit&#8217;s irregular geometry and volume of about 66.4 cubic meters, its dominant sliding-falling failure mode, the controlling fracture set with its attitude expressed as strike and dip angles, the weak Lower Jurassic mudstone pedestal that softens and forms cavities under weathering, and three stability coefficients, 8.62 under natural conditions, 2.00 under storm conditions and 5.48 under seismic conditions, all confirming a currently stable state. That kind of multi-fact extraction, performed consistently across an entire archive, is precisely what manual digitization struggles to deliver.</p>
<p>The third component is the retrieval architecture. The team adopted LightRAG, a graph-based retrieval-augmented generation approach, and paired it with a hybrid document retrieval strategy. Questions and document chunks are encoded with the BGE-M3 embedding model into 1,024-dimensional dense vectors, while a classic BM25 keyword retriever runs in parallel; the two ranked lists are merged using Reciprocal Rank Fusion with a constant of 60, and the top ten chunks are retained. On the graph side, entities in a question are linked to knowledge-graph nodes by exact, alias or semantic matching, then the system traverses incoming and outgoing relations up to two hops, keeping only simple paths without repeated nodes. The ten best paths, ranked by cosine similarity between question and path embeddings, are serialized as head-relation-tail triplets and handed to the language model alongside the document context. This dual retrieval lets the system answer both pinpoint factual queries and broader questions about failure mechanisms.</p>
<p>The performance numbers are striking. With the complete extraction configuration, the system achieved a precision of 89.00 percent, a recall of 87.74 percent and an F1-score of 88.36 percent against a ground-truth dataset annotated by five professor-level senior and senior engineers, each with more than ten years of experience in engineering geology. Ablation analysis showed that the four-stage conversational extraction process was the primary driver of that performance, with GeoGPT itself contributing a smaller additional benefit under the same extraction setting. The annotation protocol itself was rigorous: two experts independently labeled 1,200 text chunks, from which 212 entities were drawn, reaching an initial agreement rate of 91.8 percent and a Cohen&#8217;s Kappa of 0.894, with a third senior expert adjudicating all disputes. Most disagreements, 73.5 percent of them, clustered at the boundary between environmental background and triggering factors, a reminder that even expert judgment finds some geological distinctions genuinely ambiguous.</p>
<p>For question answering, the team ran a blinded pairwise evaluation across four precarious rock task categories, comparing GeoGPT against two powerful general-purpose models, DeepSeek-V3 accessed through its official API and Qwen2.5-72B-Instruct deployed locally on eight A100 GPUs with BF16 precision and greedy decoding at temperature zero. The GeoGPT-based system achieved higher overall pairwise win rates than both competitors across all four categories, with the largest margins appearing in the most demanding tasks: stability assessment and mitigation planning. A qualitative example shows why. Asked to analyze collapse mechanisms, DeepSeek-V3 assembled a competent descriptive summary of physical parameters, while GeoGPT structured its answer along explicit geological causal chains, incorporating formal failure mode classifications such as sliding, toppling and composite modes, and even geomechanical expressions like the toppling coefficient equation, aligning its output with domain standards for hazard assessment.</p>
<p>The implications reach well beyond one mountain range. Landslides and rockfalls are among the deadliest geological hazards worldwide, and climate change, slope urbanization and aging infrastructure are expected to increase exposure in many regions. Knowledge graphs have already proven their value in geoscience for link prediction and disaster cause analysis, and large language models are being applied to everything from landslide image analysis to earthquake fatality estimation. What this study adds is a template for combining them responsibly: an ontology to constrain what the machine can extract, a staged conversational process to keep extraction faithful to the source text, and graph-aware retrieval to ground answers in verified facts rather than the model&#8217;s own tendencies toward hallucination, a known weakness of generative systems.</p>
<p>There are honest limits. The datasets analyzed in the study cannot be released publicly because of multi-party data-sharing agreements and confidentiality restrictions involving the regional geological survey institutions of the Three Gorges Project, though the authors say data are available upon reasonable request with formal permission from the competent geological authorities. The evaluation also rests on a specific corpus and a specific panel of expert annotators, so the reported win rates describe performance within that setting rather than a universal guarantee. Even so, the framework demonstrates something quietly transformative: that the tacit expertise embedded in thousands of pages of engineering reports can be formalized, verified and made conversational. For the geologists watching rock slopes above schools and highways, an assistant that answers mechanism-level questions in seconds, grounded in their own field&#8217;s ontology, could change how hazard assessments get done.</p>
<p><strong>Subject of Research:</strong> An ontology-guided LightRAG framework for knowledge-graph construction and question answering on precarious rock hazards</p>
<p><strong>Article Title:</strong> An ontology-guided LightRAG framework for knowledge-graph construction and question answering on precarious rock hazards</p>
<p><strong>Article References:</strong> Li, K., Pan, Y., Yi, Z., Guo, S., Liu, P., Dong, L., Wang, L., Wang, Y., Qin, W., Cheng, X., She, N., &amp; Yuan, J. (2026). An ontology-guided LightRAG framework for knowledge-graph construction and question answering on precarious rock hazards. <em>Earth Science Informatics, 19</em>(11), Article 202. <a href="https://doi.org/10.1007/s12145-026-02258-9" rel="noopener noreferrer">https://doi.org/10.1007/s12145-026-02258-9</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s12145-026-02258-9" rel="noopener noreferrer">10.1007/s12145-026-02258-9</a></p>
<p><strong>Keywords:</strong> precarious rock masses, knowledge graph, ontology, LightRAG, retrieval-augmented generation, GeoGPT, large language models, geohazard assessment, Three Gorges, information extraction, question answering, landslide hazards</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">243519</post-id>	</item>
	</channel>
</rss>
