Mapping Fine-Tuning and Dataset Curation Practices for Large Language Models in Agriculture and Water Resources: A Scoping Review

Main Article Content

Claudia Lengua-Cantero
María García Medina
Carlos Cohen Manrique
Laudyt Lambraño Pérez
David Acosta Meza

Abstract

The adaptation of Large Language Models (LLMs) to agriculture and water resource management is an emerging field whose fine-tuning strategies and dataset curation practices remain methodologically heterogeneous, with direct implications for Sustainable Development Goals 2, 6, and 13. This scoping review maps the extent, range, and nature of the available literature, identifies dominant models and datasets, and characterises research gaps. Sources of evidence were included if they reported on Large Language Models, foundation models, or transformer-based architectures with an explicit domain-adaptation methodology in agriculture, water resources, hydrology, or irrigation, and described at least one fine-tuning or dataset curation action. A structured search across Scopus, IEEE Xplore, ACM Digital Library, SciELO, and ScienceDirect, with supplementary searches in Nature, MDPI, OpenReview, ACL Anthology, and Frontiers, covered publications from January 2020 to October 2025. Following the methodological framework of Arksey and O’Malley and the PRISMA-ScR reporting guidance, two reviewers independently charted bibliographic metadata, primary domain, AI architecture, adaptation strategy, data strategy, and application focus from the full extracted record of each included source. From 244 compiled records, 91 duplicates and 2 non-verifiable records were removed, yielding 151 verified sources with DOI or stable URL. The LLM/RAG era proper begins in mid- to late 2023, with 76.3% of dated sources published in 2024–2025. Open-source Llama-3 and Mistral families are the most frequently named fine-tuned backbones, while GPT-4 dominates evaluation and hybrid RAG architectures, and dedicated benchmarks and curated datasets signal a methodological shift from model-centric to data-centric approaches. The mapped literature points to dataset quality and domain-specific curation as central engineering concerns for the effective and trustworthy deployment of LLMs in agro-hydrological contexts.

Article Details

How to Cite
Claudia Lengua-Cantero, María García Medina, Carlos Cohen Manrique, Laudyt Lambraño Pérez, & David Acosta Meza. (2026). Mapping Fine-Tuning and Dataset Curation Practices for Large Language Models in Agriculture and Water Resources: A Scoping Review. Waterlines, 44(3s), 271–301. Retrieved from https://papjournals.com/index.php/waterlines/article/view/1092
Section
Articles

Similar Articles

1 2 3 4 5 6 7 8 9 > >> 

You may also start an advanced similarity search for this article.