<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>data extraction &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/data-extraction/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Mon, 05 Oct 2026 06:57:47 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.3</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>data extraction &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>AI Turns Tangled Sustainability Studies Into Searchable Knowledge Graphs</title>
		<link>https://scienmag.com/ai-turns-tangled-sustainability-studies-into-searchable-knowledge-graphs/</link>
		
		<dc:creator><![CDATA[Sloane Callahan]]></dc:creator>
		<pubDate>Mon, 05 Oct 2026 06:57:47 +0000</pubDate>
				<category><![CDATA[Climate]]></category>
		<category><![CDATA[AI applications in sustainable product assessment]]></category>
		<category><![CDATA[AI-driven data mining in environmental studies]]></category>
		<category><![CDATA[automated life cycle inventory collection]]></category>
		<category><![CDATA[bio-based chemicals]]></category>
		<category><![CDATA[bio-based plastics environmental impact]]></category>
		<category><![CDATA[data extraction]]></category>
		<category><![CDATA[environmental data extraction from scientific literature]]></category>
		<category><![CDATA[environmental footprint analysis]]></category>
		<category><![CDATA[industrial ecology]]></category>
		<category><![CDATA[industrial ecology research tools]]></category>
		<category><![CDATA[knowledge graph technology in sustainability]]></category>
		<category><![CDATA[knowledge graphs]]></category>
		<category><![CDATA[knowledge graphs for sustainability data]]></category>
		<category><![CDATA[large language models]]></category>
		<category><![CDATA[large language models for environmental research]]></category>
		<category><![CDATA[Life Cycle Assessment]]></category>
		<category><![CDATA[life cycle inventory]]></category>
		<category><![CDATA[natural language processing]]></category>
		<category><![CDATA[Neo4j]]></category>
		<category><![CDATA[retrieval-augmented generation]]></category>
		<category><![CDATA[scholarly archaeology in LCA]]></category>
		<category><![CDATA[Sustainability]]></category>
		<category><![CDATA[sustainability life cycle assessment]]></category>
		<category><![CDATA[text-to-Cypher]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=237144</guid>

					<description><![CDATA[Researchers have built a framework that uses large language models and knowledge graphs to automatically extract, structure, and query life cycle assessment data from scientific literature.]]></description>
										<content:encoded><![CDATA[<p>Every life cycle assessment begins with an act of scholarly archaeology. To quantify the environmental footprint of a product—a bio-based plastic, a fuel, a chemical intermediate—researchers must first assemble a life cycle inventory: a meticulous accounting of every input and output, from raw materials and energy to emissions and waste, across the entire production chain. That data lives scattered across thousands of published papers, locked inside tables, footnotes, and methodological asides written in the idiosyncratic vocabulary of each individual study. A new study published in the Journal of Industrial Ecology shows that large language models, paired with knowledge graph technology, can now do this tedious mining automatically—and answer plain-language questions about the results.</p>
<p>The research, led by Zirui Tang and Qingshi Tu of the Sustainable Bioeconomy Research Group at the University of British Columbia, together with Max Dreger and Kourosh Malek of Forschungszentrum Jülich in Germany and Peijin Jiang of UBC, addresses a bottleneck that has long frustrated the sustainability field. Life cycle assessment, or LCA, is the standard tool for comparing the environmental credentials of products and services, but its reliability hinges on high-quality inventory data. Collecting that data by hand from the literature is slow, labor-intensive, and notoriously difficult to scale, which means that meta-analyses and large-scale syntheses often lag years behind the primary studies they should be aggregating.</p>
<p>The team&#8217;s framework combines two complementary technologies. On the front end, a retrieval-augmented generation pipeline—commonly known as RAG—sends large language models into the full text of published articles to extract three core categories of information: life cycle inventory data, life cycle impact assessment results, and the modeling assumptions that underpin each study. On the back end, the extracted information is normalized and mapped into an ontology-driven knowledge graph built specifically for LCA data and implemented in the Neo4j graph database. The ontology gives every extracted fact a defined place in a shared semantic structure, turning a pile of heterogeneous paper-specific numbers into an interconnected, machine-queryable web of knowledge.</p>
<p>The technical challenge is considerable. Scientific papers do not present their data uniformly; inventory tables appear in different formats, use inconsistent units and system boundaries, and embed crucial context—such as functional units, allocation choices, and geographic scope—in prose rather than structured fields. The pipeline therefore has to parse documents, identify which tables actually contain inventory data, and interpret the surrounding text to understand what each number means. The authors report that the extraction stage achieved F1-scores ranging from 75.18 to 89.50 percent across the three data types, a level of semantic accuracy that suggests the models are genuinely understanding the content rather than merely pattern-matching on keywords.</p>
<p>Extraction alone, however, would only produce a better-organized archive. The second half of the framework is what makes it genuinely interactive: an LLM-based question-answering system that translates natural language queries into executable graph queries. A user can ask, in ordinary English, about the greenhouse gas intensity of a particular bio-based chemical under a given set of assumptions, and the system converts the question into Cypher—the query language of Neo4j—retrieves the relevant nodes and relationships from the knowledge graph, and returns an answer. Crucially, users need no prior knowledge of the graph&#8217;s schema, which has historically been one of the biggest barriers preventing domain scientists from exploiting structured databases directly.</p>
<p>The query system itself was evaluated rigorously, and the results reveal an instructive lesson about how to build such tools. A baseline approach relying on a single strategy achieved an F1-score of only 56.98 percent—hardly reliable enough for scientific work. By combining similarity search, which locates the most relevant portions of the graph for a given question, with text-to-Cypher reasoning, which generates the formal query, the team raised the score to 75.18 percent. The hybrid architecture outperforms either method alone because the two strategies fail in different ways: semantic search is robust to vague phrasing but imprecise about graph structure, while query generation is precise but brittle when a question does not map cleanly onto the schema.</p>
<p>To demonstrate the framework in practice, the researchers applied it to a case study of LCA studies on bio-based chemical production—a domain where the literature has expanded rapidly as industry and academia race to replace fossil-derived chemicals with renewable alternatives. Bio-based chemicals are an ideal testbed because published assessments vary widely in feedstock, conversion technology, and impact categories, making manual synthesis across studies especially painful. The knowledge graph lets researchers compare results across studies while preserving the assumptions that make each figure meaningful, something a simple spreadsheet of extracted numbers could never do.</p>
<p>The broader significance extends beyond one application domain. The study arrives amid a wave of efforts to apply large language models to scientific data extraction, from polymer properties to carbon capture technologies to toxicology reports. What distinguishes this work is the full pipeline: it does not stop at extraction but carries the data through normalization, ontological mapping, and interactive retrieval, addressing the interoperability problem that has long plagued LCA data systems. Previous semantic efforts in the field, including minimal ontology patterns for LCA data proposed over the past decade, laid conceptual groundwork, but the manual effort required to populate such structures remained prohibitive. Automating that population step changes the economics of knowledge synthesis entirely.</p>
<p>There are, of course, caveats. F1-scores in the 75 to 90 percent range mean that a meaningful fraction of extracted facts will still be wrong or missed, and in a field where numbers feed policy decisions and corporate sustainability claims, errors carry consequences. The authors make their ground-truth datasets, extracted results, and pipeline code publicly available through the article&#8217;s supplementary materials and a GitHub repository, which allows the community to audit performance and build improvements. The work was supported by the Natural Sciences and Engineering Research Council of Canada, and the authors declare no competing interests.</p>
<p>Still, the trajectory is clear. As the volume of scientific literature grows faster than any human team can read, the bottleneck in fields like industrial ecology is shifting from generating data to finding and reconciling it. A framework that reads the papers, structures the findings, and answers questions in plain language offers a glimpse of how sustainability science might scale its evidence base to meet the pace of the energy and materials transition—turning decades of scattered assessments into a single, queryable map of environmental knowledge.</p>
<p><strong>Subject of Research:</strong> Automated extraction and retrieval of life cycle assessment data from scientific literature using large language models and knowledge graphs</p>
<p><strong>Article Title:</strong> From literature to knowledge graphs: automated extraction and retrieval of life cycle assessment data with large language models</p>
<p><strong>Article References:</strong> Tang, Z., Dreger, M., Jiang, P., Malek, K., &amp; Tu, Q. (2026). From literature to knowledge graphs: automated extraction and retrieval of life cycle assessment data with large language models. <em>Journal of Industrial Ecology, 30</em>(4), 1985-2000. <a href="https://doi.org/10.1007/s44498-026-00135-8" rel="noopener noreferrer">https://doi.org/10.1007/s44498-026-00135-8</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s44498-026-00135-8" rel="noopener noreferrer">10.1007/s44498-026-00135-8</a></p>
<p><strong>Keywords:</strong> life cycle assessment, large language models, knowledge graphs, retrieval-augmented generation, life cycle inventory, data extraction, Neo4j, text-to-Cypher, sustainability, industrial ecology, natural language processing, bio-based chemicals</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">237144</post-id>	</item>
		<item>
		<title>AI Pipeline Turns Messy Clinical Notes Into Research-Ready Data With Near-Perfect Accuracy</title>
		<link>https://scienmag.com/ai-pipeline-turns-messy-clinical-notes-into-research-ready-data-with-near-perfect-accuracy/</link>
		
		<dc:creator><![CDATA[Ophelia Keating]]></dc:creator>
		<pubDate>Sun, 20 Sep 2026 19:12:50 +0000</pubDate>
				<category><![CDATA[Medicine]]></category>
		<category><![CDATA[AI-powered clinical note extraction]]></category>
		<category><![CDATA[Artificial Intelligence]]></category>
		<category><![CDATA[automated chart review]]></category>
		<category><![CDATA[CLASS pipeline]]></category>
		<category><![CDATA[clinical documentation analysis]]></category>
		<category><![CDATA[clinical notes]]></category>
		<category><![CDATA[Clinical Research]]></category>
		<category><![CDATA[data extraction]]></category>
		<category><![CDATA[electronic health records]]></category>
		<category><![CDATA[esophageal airway treatment surgery]]></category>
		<category><![CDATA[healthcare data accuracy]]></category>
		<category><![CDATA[hospital informatics solutions]]></category>
		<category><![CDATA[Johns Hopkins]]></category>
		<category><![CDATA[Johns Hopkins AI healthcare project]]></category>
		<category><![CDATA[large language models]]></category>
		<category><![CDATA[large language models for medical data]]></category>
		<category><![CDATA[medical informatics]]></category>
		<category><![CDATA[medical record data structuring]]></category>
		<category><![CDATA[natural language processing]]></category>
		<category><![CDATA[natural language processing in healthcare]]></category>
		<category><![CDATA[pediatric surgery]]></category>
		<category><![CDATA[privacy-preserving medical AI]]></category>
		<category><![CDATA[research-ready electronic health records]]></category>
		<category><![CDATA[secure medical data processing]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=201580</guid>

					<description><![CDATA[Researchers at Johns Hopkins All Children's Hospital developed CLASS, a large language model pipeline that extracted pediatric surgical procedure data from unstructured clinical notes with near-perfect concordance to expert review.]]></description>
										<content:encoded><![CDATA[<p>Electronic medical records are often described as gold mines of clinical information, but most of that treasure is buried. While laboratory values, vital signs, and medication orders arrive neatly coded, the richest details of a patient&#8217;s story—the surgical nuances, the procedural variations, the clinical judgment calls—live inside free-text notes written by busy clinicians. For researchers and quality-improvement teams, extracting that information has long meant hours of painstaking manual chart review, a process that is slow, expensive, and nearly impossible to scale. Now, a team at Johns Hopkins All Children&#8217;s Hospital has shown that a carefully engineered large language model pipeline can do much of that work automatically, with accuracy so high that it approaches the ceiling of human agreement.</p>
<p>The system, called CLASS—short for Clinical LLM Abstraction &amp; Structuring System—is described in a new feasibility report published in the Journal of Medical Systems. Led by anesthesiologist and informatics researcher Frederick H. Kuo, the team built CLASS as a modular, Python-based pipeline that runs entirely within a secure institutional computing environment, keeping protected health information behind the hospital&#8217;s own walls rather than sending it to external services. That design choice reflects a growing consensus in medical informatics: the power of large language models can be harnessed for clinical data without compromising patient privacy, provided the infrastructure is configured correctly.</p>
<p>Technically, CLASS rests on three pillars. The first is a set of concept lists curated by subject matter experts—in this case, surgeons and informatics specialists who defined exactly which procedures and clinical details the system should look for. The second is a task-specific prompt suite, a collection of carefully worded instructions that steers the large language model toward consistent, clinically grounded interpretations of each note. The third is a schema-constrained output format, which forces the model to return its findings in a structured, predictable structure rather than free-flowing prose. The results are then exported to an interactive dashboard built for expert review, allowing clinicians to verify outputs, spot errors, and analyze the extracted data at scale.</p>
<p>One of the most innovative features of CLASS is its handling of the unknown. Rather than simply classifying notes against a fixed list of predefined concepts, the pipeline actively flags potential variants or entirely novel concepts that do not fit the existing schema, surfacing them for expert consideration. In specialized fields where terminology evolves quickly and procedures are often described in non-standard ways, this ability to propose expansions to the concept vocabulary could fundamentally change how clinical registries and research databases are built and maintained.</p>
<p>To test the system, the researchers applied CLASS to a retrospective corpus of pediatric esophageal airway treatment surgery, or EATS, operative notes at their single center. EATS is a demanding test case: these operative notes describe complex, highly individualized procedures in children with airway and esophageal abnormalities, and much of the procedural detail is not captured in standard billing codes. The team compared CLASS outputs against adjudication by an experienced surgeon on the twenty longest notes in the corpus, yielding 3,960 individual note-procedure pairs for evaluation.</p>
<p>The results were striking. Across those thousands of judgments, the observed concordance between the automated pipeline and the surgeon&#8217;s adjudication reached an F1 score of 0.9967—a near-perfect measure of precision and recall combined. In practical terms, the model almost never missed a procedure the surgeon identified, and almost never invented one that was not there. For a task as subtle as parsing operative prose about pediatric airway surgery, that level of agreement suggests that large language models, when properly constrained and prompted, can match expert-level abstraction performance in at least some specialized clinical domains.</p>
<p>The novelty-detection results were more nuanced, and arguably more interesting. CLASS proposed 28 candidate procedure variants or additions to the curated concept list, and the clinical team judged 18 of them—64.3 percent—to be genuinely useful. That is a meaningful yield: nearly two out of every three suggestions from the machine were worth a clinician&#8217;s time. At the same time, the surgeon identified 12 additional procedures that CLASS failed to surface, a reminder that the system works best as a collaborator rather than a replacement. The human expert still caught things the machine missed, and the machine still surfaced things the human might not have thought to codify.</p>
<p>The authors are careful to frame the study appropriately. This was an exploratory implementation at a single center, focused on a single surgical service, with evaluation limited to a small set of long notes. Generalizability to other institutions, other note types, and other clinical tasks remains unproven, and the team emphasizes that broader validation is needed before such pipelines could be trusted for high-stakes applications. The full production code and clinical data cannot be released publicly because they involve protected health information and institution-specific infrastructure, but the researchers have shared technical implementation details, template code, pseudocode, and de-identified prompt examples in the online supplementary materials, giving other informatics teams a practical roadmap for building similar systems.</p>
<p>Even with those caveats, the implications are considerable. Clinical research has long been throttled by the bottleneck of manual abstraction: cohort studies that could enroll thousands of patients are often limited to hundreds simply because there are only so many hours in a research coordinator&#8217;s day. Quality-improvement programs face the same constraint, unable to measure surgical outcomes comprehensively when the relevant data must be pulled by hand from narrative notes. If pipelines like CLASS can reliably convert unstructured text into analyzable data within standard institutional infrastructure, the effective sample sizes of clinical research could expand dramatically, and hospitals could monitor the quality of specialized care in near real time.</p>
<p>The study also highlights a shift in how medical informatics teams may work in the coming years. Instead of writing brittle rule-based extraction algorithms or training bespoke machine-learning models on small labeled datasets, teams can now curate expert concept lists, design prompts, and review machine-generated suggestions—a workflow in which clinicians define what matters and the model handles the linguistic heavy lifting. The CLASS experience suggests this human-machine partnership can work: the model performs the extraction with near-perfect fidelity, proposes useful vocabulary expansions, and leaves final judgment to the experts who bear clinical responsibility. As large language models continue to demonstrate their grasp of medical language, studies like this one offer a concrete, privacy-conscious template for turning the narrative richness of the medical record into structured knowledge—without a single chart being pulled by hand.</p>
<p><strong>Subject of Research:</strong> A large language model pipeline for extracting structured data from unstructured clinical notes</p>
<p><strong>Article Title:</strong> Exploratory Implementation and Feasibility Report of CLASS (Clinical LLM Abstraction &amp; Structuring System), A Large Language Model Pipeline for Extracting Unstructured Data From Clinical Notes</p>
<p><strong>Article References:</strong> Exploratory Implementation and Feasibility Report of CLASS (Clinical LLM Abstraction &amp; Structuring System), A Large Language Model Pipeline for Extracting Unstructured Data From Clinical Notes. (n.d.). <a href="https://doi.org/10.1007/s10916-026-02462-6" rel="noopener noreferrer">https://doi.org/10.1007/s10916-026-02462-6</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s10916-026-02462-6" rel="noopener noreferrer">10.1007/s10916-026-02462-6</a></p>
<p><strong>Keywords:</strong> large language models, clinical notes, natural language processing, electronic health records, medical informatics, data extraction, pediatric surgery, artificial intelligence, CLASS pipeline, clinical research, esophageal airway treatment surgery, Johns Hopkins</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">201580</post-id>	</item>
	</channel>
</rss>
