Saturday, October 10, 2026
Science
No Result
View All Result
  • Login
  • HOME
  • SCIENCE NEWS
  • CONTACT US
  • HOME
  • SCIENCE NEWS
  • CONTACT US
No Result
View All Result
Scienmag
No Result
View All Result
Home Science News Technology and Engineering

Fine-Tuned Llama Model Turns Unstructured Farm Data Into a Searchable Knowledge Graph

October 10, 2026
in Technology and Engineering
Alan Morgan
By Alan Morgan Scienmag Editorial Profile - Precision Agriculture
Reading Time: 4 mins read
0
Fine-Tuned Llama Model Turns Unstructured Farm Data Into a Searchable Knowledge Graph

Fine-Tuned Llama Model Turns Unstructured Farm Data Into a Searchable Knowledge Graph

65
SHARES
587
VIEWS
Share on FacebookShare on Twitter
ADVERTISEMENT

Agriculture generates an enormous volume of unstructured text: research reports, extension advisories, news articles, farmers’ feedback and expert recommendations. Most of this information sits locked in documents that machines cannot easily reason over, even though it could power precision agriculture and automated decision support. A new study from COEP Technological University in Pune, India, published in Discover Artificial Intelligence, describes a framework that converts fragmented agronomic text into a structured knowledge graph and then answers farmers’ and researchers’ questions with unusually high precision.

The work, carried out by Rohini Kokare and Sunil Mane, focuses on a niche but important use case: soybean cultivation advisory for Indian agriculture. The researchers scraped unstructured documents from the ICAR-National Soybean Research Institute in Indore, Tamil Nadu Agricultural University and the TNAU Agritech portal, covering crop protection, crop production, farm mechanization, good agricultural practices, variety information and region-wise crop recommendations. From these documents they built a curated instruction dataset of 1,200 samples, each containing exactly one relational triplet, split 90:10 between training and testing with no sentence overlap.

At the heart of the system is a fine-tuned Llama-3.2-1B-Instruct model, a one-billion-parameter language model with multilingual support for eight languages. Rather than retraining the entire network, the team used Low-Rank Adaptation, or LoRA, a parameter-efficient fine-tuning technique that freezes the pre-trained weights and trains only small adapter matrices. For a frozen weight matrix of size 512 by 512, LoRA decomposes the update into two matrices of 512 by 8 and 8 by 512, yielding just 8,192 trainable parameters per adapted matrix. This dramatic reduction in trainable parameters makes the approach computationally affordable while still allowing the model to learn domain-specific behavior.

The researchers chose their target modules carefully. In Llama-style architectures, the feed-forward sub-block contains most of the parameters, so adapting only the attention mechanism would leave much of the model’s representational capacity untouched. The team therefore applied LoRA across all linear modules: the attention projections for query, key, value and output, plus the gated MLP SwiGLU blocks known as gate_proj, up_proj and down_proj. They set the LoRA rank to 8, with an alpha of 16 giving a scaling factor of two, and applied a light dropout of 0.05 on the adapter path to avoid overfitting on the specialized vocabulary of agronomy.

Instruction fine-tuning was central to the approach. Instead of simply exposing the model to domain text, the researchers trained it with explicit instructions such as “Extract entities and relationships as triples,” paired with input sentences and structured outputs. An example from the dataset maps the sentence “Macs 330 controls Bradyrhizobium japonicum” to the triplet consisting of the subject MACS 330, the relation “controls” and the object Bradyrhizobium japonicum. This format teaches the model to follow task-specific guidelines and produce consistent, machine-parseable output rather than free-form text.

Once fine-tuned, the model extracts triplets from unstructured text at inference time, and these subject-predicate-object structures are pushed into Neo4j, a graph database where entities become nodes and relations become edges. The resulting knowledge graph captures multi-relational information about soybean varieties, diseases, pests, treatments and practices in a form that supports structured queries written in the Cypher query language. This graph then serves as the retrieval backbone for a GraphRAG system, a hybrid of graph retrieval and large language model generation.

The GraphRAG retrieval layer combines vector and keyword search over the text properties of document nodes, using embeddings from HuggingFace models matched against embedding properties stored on each node. A user’s natural language query flows directly into the hybrid retriever through a LangChain pipeline without rewriting or entity extraction. Both the retrieved graph context and the original question are then passed to a fixed prompt template, and the fine-tuned Llama model generates the final answer. This design grounds responses in retrieved facts, reducing the hallucination risk that plagues out-of-the-box language models when they answer specialized questions.

The experimental results showed clear gains from the approach. When the team compared fine-tuning configurations, the model trained with all projection modules and instruction fine-tuning outperformed variants trained only on attention layers or only on feed-forward layers, measured by macro-averaged precision, recall and F1-score across the test set. On end-to-end response generation, the fine-tuned GraphRAG pipeline beat conventional vector database retrieval, standard RAG and knowledge-graph-only retrieval on both ROUGE and BLEU precision metrics, which measure how closely generated answers overlap with reference responses at the unigram, bigram and longest-common-subsequence levels.

The authors argue that the framework addresses a genuine gap. Previous research has explored agricultural chatbots and question-answering systems, agricultural knowledge graphs and general-purpose GraphRAG frameworks separately, but no existing system, to their knowledge, had combined LoRA-based generative triple extraction, automated knowledge graph construction and GraphRAG-style retrieval for a low-resource domain like Indian soybean advisory. Conventional pipelines that treat named entity recognition and relation extraction as separate steps are labor-intensive to train and prone to cascading errors, while generalized prompting or full-parameter fine-tuning is computationally inefficient and time-consuming.

The work is not finished. The researchers note that systematic verification of triplet semantics and quality, including comparison of gold versus predicted triples and elimination of noisy extractions, is scheduled as upcoming work. They also plan explicit graph-traversal-based retrieval using relation-path expansion from matched entities, and controlled comparisons against size-matched open-weight models such as Qwen2.5-1.5B-Instruct and DeepSeek small variants, as well as larger frontier models like GPT-4o and Mistral, under identical test questions and decoding settings. Even so, the study demonstrates that a small, efficiently fine-tuned language model can transform scattered agricultural documents into a queryable knowledge network, offering a template for bringing structured, trustworthy AI question-answering to domains where data is abundant but structured knowledge is scarce.

Subject of Research: LoRA-based instruction fine-tuning of large language models for agronomic triple extraction and GraphRAG knowledge retrieval

Article Title: LoRA based instruction fine tuning of large language models for agronomic triple extraction and GraphRAG knowledge retrieval

Article References: Kokare, R., & Mane, S. (2026). LoRA based instruction fine tuning of large language models for agronomic triple extraction and GraphRAG knowledge retrieval. Discover Artificial Intelligence, 6(1), Article 1381. https://doi.org/10.1007/s44163-026-02410-w

Image Credits: AI Generated

DOI: 10.1007/s44163-026-02410-w

Keywords: LoRA, instruction fine-tuning, large language models, Llama, triple extraction, GraphRAG, knowledge graph, Neo4j, agronomy, soybean, retrieval-augmented generation, parameter-efficient fine-tuning

Cite Scienmag News

Alan Morgan. (October 10, 2026). Fine-Tuned Llama Model Turns Unstructured Farm Data Into a Searchable Knowledge Graph. Scienmag. https://scienmag.com/fine-tuned-llama-model-turns-unstructured-farm-data-into-a-searchable-knowledge-graph/

Alan Morgan. "Fine-Tuned Llama Model Turns Unstructured Farm Data Into a Searchable Knowledge Graph." Scienmag, 10 October 2026, https://scienmag.com/fine-tuned-llama-model-turns-unstructured-farm-data-into-a-searchable-knowledge-graph/. Accessed 10 October 2026.

Alan Morgan. "Fine-Tuned Llama Model Turns Unstructured Farm Data Into a Searchable Knowledge Graph." Scienmag. October 10, 2026. https://scienmag.com/fine-tuned-llama-model-turns-unstructured-farm-data-into-a-searchable-knowledge-graph/

Tags: agriculture unstructured text dataagronomic information extractionagronomyAI-powered farm management toolsautomated farm decision supportconverting unstructured farm documentscrop advisory natural language processingfine-tuned Llama model for agricultureGraphRAGIndian agricultural research datainstruction fine-tuningknowledge graphknowledge graph for farm datalarge language modelsLLaMALoRamultilingual agricultural AI modelsNeo4jparameter-efficient fine-tuningprecision agriculture with AIretrieval-augmented generationsoybeansoybean cultivation data analysistriple extraction
Share26Tweet16
Previous Post

Breaking Down the Barriers Holding Back Soybean Contract Farming in Ethiopia

Next Post

Water Scarcity and Political Barriers Squeeze Palestinian Farmers, Study Finds

Related Posts

Magnetic Density Separation Gets a Map of Where It Actually Works
Technology and Engineering

Magnetic Density Separation Gets a Map of Where It Actually Works

October 10, 2026
Coal-Derived Carbon Nanodots Offer a Primer for Atomically Thin Transistors
Technology and Engineering

Coal-Derived Carbon Nanodots Offer a Primer for Atomically Thin Transistors

October 10, 2026
Crop and Forest Waste Could Reshape the Global Energy Map, Study Finds
Technology and Engineering

Crop and Forest Waste Could Reshape the Global Energy Map, Study Finds

October 10, 2026
Front-Electrode Engineering Pushes Perovskite/Silicon Tandem Solar Cells to 33% Efficiency
Technology and Engineering

Front-Electrode Engineering Pushes Perovskite/Silicon Tandem Solar Cells to 33% Efficiency

October 10, 2026
New Tool HiFIseek Hunts Hidden Cancer Mutations in the Genome’s Dark Regulatory Regions
Biology

New Tool HiFIseek Hunts Hidden Cancer Mutations in the Genome’s Dark Regulatory Regions

October 10, 2026
Machine Learning Tool Mines Health Records to Catch Hidden Heart Disease
Medicine

Machine Learning Tool Mines Health Records to Catch Hidden Heart Disease

October 10, 2026
Next Post
Water Scarcity and Political Barriers Squeeze Palestinian Farmers, Study Finds

Water Scarcity and Political Barriers Squeeze Palestinian Farmers, Study Finds

  • Mothers who receive childcare support from maternal grandparents show more optimized

    Mothers who receive childcare support from maternal grandparents show more parental warmth, finds NTU Singapore study

    27656 shares
    Share 11059 Tweet 6912
  • University of Seville Breaks 120-Year-Old Mystery, Revises a Key Einstein Concept

    1061 shares
    Share 424 Tweet 265
  • Bee body mass, pathogens and local climate influence heat tolerance

    682 shares
    Share 273 Tweet 171
  • Researchers record first-ever images and data of a shark experiencing a boat strike

    546 shares
    Share 218 Tweet 137
  • Groundbreaking Clinical Trial Reveals Lubiprostone Enhances Kidney Function

    531 shares
    Share 212 Tweet 133
Science

Embark on a thrilling journey of discovery with Scienmag.com—your ultimate source for cutting-edge breakthroughs. Immerse yourself in a world where curiosity knows no limits and tomorrow’s possibilities become today’s reality!

RECENT NEWS

  • Alumina-Boosted Fly Ash Geopolymers Capture Radioactive Metals More Efficiently
  • Water Scarcity and Political Barriers Squeeze Palestinian Farmers, Study Finds
  • Fine-Tuned Llama Model Turns Unstructured Farm Data Into a Searchable Knowledge Graph
  • Breaking Down the Barriers Holding Back Soybean Contract Farming in Ethiopia

Categories

  • Agriculture
  • Anthropology
  • Archaeology
  • Athmospheric
  • Biology
  • Biotechnology
  • Blog
  • Bussines
  • Cancer
  • Chemistry
  • Climate
  • Earth Science
  • Editorial Policy
  • Marine
  • Mathematics
  • Medicine
  • Pediatry
  • Policy
  • Psychology & Psychiatry
  • Science Education
  • Science News
  • Social Science
  • Space
  • Technology and Engineering

Subscribe to Blog via Email

Enter your email address to subscribe to this blog and receive notifications of new posts by email.

Join 5,150 other subscribers

© 2025 Scienmag - Science Magazine

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • HOME
  • SCIENCE NEWS
  • CONTACT US

© 2025 Scienmag - Science Magazine

Discover more from Science

Subscribe now to keep reading and get access to the full archive.

Continue reading