Agriculture generates an enormous volume of unstructured text: research reports, extension advisories, news articles, farmers’ feedback and expert recommendations. Most of this information sits locked in documents that machines cannot easily reason over, even though it could power precision agriculture and automated decision support. A new study from COEP Technological University in Pune, India, published in Discover Artificial Intelligence, describes a framework that converts fragmented agronomic text into a structured knowledge graph and then answers farmers’ and researchers’ questions with unusually high precision.
The work, carried out by Rohini Kokare and Sunil Mane, focuses on a niche but important use case: soybean cultivation advisory for Indian agriculture. The researchers scraped unstructured documents from the ICAR-National Soybean Research Institute in Indore, Tamil Nadu Agricultural University and the TNAU Agritech portal, covering crop protection, crop production, farm mechanization, good agricultural practices, variety information and region-wise crop recommendations. From these documents they built a curated instruction dataset of 1,200 samples, each containing exactly one relational triplet, split 90:10 between training and testing with no sentence overlap.
At the heart of the system is a fine-tuned Llama-3.2-1B-Instruct model, a one-billion-parameter language model with multilingual support for eight languages. Rather than retraining the entire network, the team used Low-Rank Adaptation, or LoRA, a parameter-efficient fine-tuning technique that freezes the pre-trained weights and trains only small adapter matrices. For a frozen weight matrix of size 512 by 512, LoRA decomposes the update into two matrices of 512 by 8 and 8 by 512, yielding just 8,192 trainable parameters per adapted matrix. This dramatic reduction in trainable parameters makes the approach computationally affordable while still allowing the model to learn domain-specific behavior.
The researchers chose their target modules carefully. In Llama-style architectures, the feed-forward sub-block contains most of the parameters, so adapting only the attention mechanism would leave much of the model’s representational capacity untouched. The team therefore applied LoRA across all linear modules: the attention projections for query, key, value and output, plus the gated MLP SwiGLU blocks known as gate_proj, up_proj and down_proj. They set the LoRA rank to 8, with an alpha of 16 giving a scaling factor of two, and applied a light dropout of 0.05 on the adapter path to avoid overfitting on the specialized vocabulary of agronomy.
Instruction fine-tuning was central to the approach. Instead of simply exposing the model to domain text, the researchers trained it with explicit instructions such as “Extract entities and relationships as triples,” paired with input sentences and structured outputs. An example from the dataset maps the sentence “Macs 330 controls Bradyrhizobium japonicum” to the triplet consisting of the subject MACS 330, the relation “controls” and the object Bradyrhizobium japonicum. This format teaches the model to follow task-specific guidelines and produce consistent, machine-parseable output rather than free-form text.
Once fine-tuned, the model extracts triplets from unstructured text at inference time, and these subject-predicate-object structures are pushed into Neo4j, a graph database where entities become nodes and relations become edges. The resulting knowledge graph captures multi-relational information about soybean varieties, diseases, pests, treatments and practices in a form that supports structured queries written in the Cypher query language. This graph then serves as the retrieval backbone for a GraphRAG system, a hybrid of graph retrieval and large language model generation.
The GraphRAG retrieval layer combines vector and keyword search over the text properties of document nodes, using embeddings from HuggingFace models matched against embedding properties stored on each node. A user’s natural language query flows directly into the hybrid retriever through a LangChain pipeline without rewriting or entity extraction. Both the retrieved graph context and the original question are then passed to a fixed prompt template, and the fine-tuned Llama model generates the final answer. This design grounds responses in retrieved facts, reducing the hallucination risk that plagues out-of-the-box language models when they answer specialized questions.
The experimental results showed clear gains from the approach. When the team compared fine-tuning configurations, the model trained with all projection modules and instruction fine-tuning outperformed variants trained only on attention layers or only on feed-forward layers, measured by macro-averaged precision, recall and F1-score across the test set. On end-to-end response generation, the fine-tuned GraphRAG pipeline beat conventional vector database retrieval, standard RAG and knowledge-graph-only retrieval on both ROUGE and BLEU precision metrics, which measure how closely generated answers overlap with reference responses at the unigram, bigram and longest-common-subsequence levels.
The authors argue that the framework addresses a genuine gap. Previous research has explored agricultural chatbots and question-answering systems, agricultural knowledge graphs and general-purpose GraphRAG frameworks separately, but no existing system, to their knowledge, had combined LoRA-based generative triple extraction, automated knowledge graph construction and GraphRAG-style retrieval for a low-resource domain like Indian soybean advisory. Conventional pipelines that treat named entity recognition and relation extraction as separate steps are labor-intensive to train and prone to cascading errors, while generalized prompting or full-parameter fine-tuning is computationally inefficient and time-consuming.
The work is not finished. The researchers note that systematic verification of triplet semantics and quality, including comparison of gold versus predicted triples and elimination of noisy extractions, is scheduled as upcoming work. They also plan explicit graph-traversal-based retrieval using relation-path expansion from matched entities, and controlled comparisons against size-matched open-weight models such as Qwen2.5-1.5B-Instruct and DeepSeek small variants, as well as larger frontier models like GPT-4o and Mistral, under identical test questions and decoding settings. Even so, the study demonstrates that a small, efficiently fine-tuned language model can transform scattered agricultural documents into a queryable knowledge network, offering a template for bringing structured, trustworthy AI question-answering to domains where data is abundant but structured knowledge is scarce.
Subject of Research: LoRA-based instruction fine-tuning of large language models for agronomic triple extraction and GraphRAG knowledge retrieval
Article Title: LoRA based instruction fine tuning of large language models for agronomic triple extraction and GraphRAG knowledge retrieval
Article References: Kokare, R., & Mane, S. (2026). LoRA based instruction fine tuning of large language models for agronomic triple extraction and GraphRAG knowledge retrieval. Discover Artificial Intelligence, 6(1), Article 1381. https://doi.org/10.1007/s44163-026-02410-w
Image Credits: AI Generated
DOI: 10.1007/s44163-026-02410-w
Keywords: LoRA, instruction fine-tuning, large language models, Llama, triple extraction, GraphRAG, knowledge graph, Neo4j, agronomy, soybean, retrieval-augmented generation, parameter-efficient fine-tuning
Cite Scienmag News
Alan Morgan. (October 10, 2026). Fine-Tuned Llama Model Turns Unstructured Farm Data Into a Searchable Knowledge Graph. Scienmag. https://scienmag.com/fine-tuned-llama-model-turns-unstructured-farm-data-into-a-searchable-knowledge-graph/
Alan Morgan. "Fine-Tuned Llama Model Turns Unstructured Farm Data Into a Searchable Knowledge Graph." Scienmag, 10 October 2026, https://scienmag.com/fine-tuned-llama-model-turns-unstructured-farm-data-into-a-searchable-knowledge-graph/. Accessed 10 October 2026.
Alan Morgan. "Fine-Tuned Llama Model Turns Unstructured Farm Data Into a Searchable Knowledge Graph." Scienmag. October 10, 2026. https://scienmag.com/fine-tuned-llama-model-turns-unstructured-farm-data-into-a-searchable-knowledge-graph/

