2025

Oishee Bintey Hoque, Nibir Chandra Mandal, Abhijin Adiga, Samarth Swarup, Sayjro Kossi Nouwakpo, Amanda Wilson, Madhav Marathe
Knowledge-Informed Deep Learning for Irrigation Type Mapping from Remote Sensing Proceedings Article
In: International Joint Conferences on Artificial Intelligence 2025.
Abstract | Links | BibTeX | Tags: Deep learning, Irrigation, Mapping, Remote Sensing
@inproceedings{nokey,
title = {Knowledge-Informed Deep Learning for Irrigation Type Mapping from Remote Sensing},
author = {Oishee Bintey Hoque, Nibir Chandra Mandal, Abhijin Adiga, Samarth Swarup, Sayjro Kossi Nouwakpo, Amanda Wilson, Madhav Marathe},
doi = { https://doi.org/10.48550/arXiv.2505.08302},
year = {2025},
date = {2025-08-22},
urldate = {2025-08-22},
organization = {International Joint Conferences on Artificial Intelligence},
abstract = {Accurate mapping of irrigation methods is crucial for sustainable agricultural practices and food systems. However, existing models that rely solely on spectral features from satellite imagery are ineffective due to the complexity of agricultural landscapes and limited training data, making this a challenging problem. We present Knowledge-Informed Irrigation Mapping (KIIM), a novel Swin-Transformer based approach that uses (i) a specialized projection matrix to encode crop to irrigation probability, (ii) a spatial attention map to identify agricultural lands from non-agricultural
lands, (iii) bi-directional cross-attention to focus complementary information from different modalities, and (iv) a weighted ensemble for combining predictions from images and crop information. Our experimentation on five states in the US shows up to 22.9% (IoU) improvement over baseline with a 71.4% (IoU) improvement for hard-to-classify drip irrigation. In addition, we propose a two-phase transfer learning approach to enhance cross-state irrigation mapping, achieving a 51% IoU boost in a state with limited labeled data. The ability to achieve baseline performance with only 40% of the training data highlights its efficiency, reducing the dependency on extensive manual labeling efforts and making large-scale, automated irrigation mapping more feasible and cost-effective. Code: https://github.com/Nibir088/KIIM},
keywords = {Deep learning, Irrigation, Mapping, Remote Sensing},
pubstate = {published},
tppubtype = {inproceedings}
}
lands, (iii) bi-directional cross-attention to focus complementary information from different modalities, and (iv) a weighted ensemble for combining predictions from images and crop information. Our experimentation on five states in the US shows up to 22.9% (IoU) improvement over baseline with a 71.4% (IoU) improvement for hard-to-classify drip irrigation. In addition, we propose a two-phase transfer learning approach to enhance cross-state irrigation mapping, achieving a 51% IoU boost in a state with limited labeled data. The ability to achieve baseline performance with only 40% of the training data highlights its efficiency, reducing the dependency on extensive manual labeling efforts and making large-scale, automated irrigation mapping more feasible and cost-effective. Code: https://github.com/Nibir088/KIIM
Ranjan Sapkota; Marco Flores-Calero; Rizwan Qureshi; Chetan Badgujar; Upesh Nepal; Alwin Poulose; Peter Zeno; Uday Bhanu Prakash Vaddevolu; Sheheryar Khan; Maged Shoman; Hong Yan; Manoj Karkee
YOLO advances to its genesis: a decadal and comprehensive review of the You Only Look Once (YOLO) series Journal Article
In: Artificial Intelligence Review, vol. 58, no. 9, pp. 274, 2025, ISSN: 1573-7462.
Abstract | Links | BibTeX | Tags: Agriculture, Artificial intelligence, Autonomous vehicles, CNN, Computer vision, Deep learning, Healthcare and medical imaging, Industrial manufacturing, Real-time object detection, Surveillance, Traffic safety, YOLO, YOLO configurations, YOLOv1 to YOLOv12, You Only Look Once
@article{sapkota_yolo_2025,
title = {YOLO advances to its genesis: a decadal and comprehensive review of the You Only Look Once (YOLO) series},
author = {Ranjan Sapkota and Marco Flores-Calero and Rizwan Qureshi and Chetan Badgujar and Upesh Nepal and Alwin Poulose and Peter Zeno and Uday Bhanu Prakash Vaddevolu and Sheheryar Khan and Maged Shoman and Hong Yan and Manoj Karkee},
url = {https://doi.org/10.1007/s10462-025-11253-3},
doi = {10.1007/s10462-025-11253-3},
issn = {1573-7462},
year = {2025},
date = {2025-06-01},
urldate = {2025-06-01},
journal = {Artificial Intelligence Review},
volume = {58},
number = {9},
pages = {274},
abstract = {This review systematically examines the progression of the You Only Look Once (YOLO) object detection algorithms from YOLOv1 to the recently unveiled YOLOv12. Employing a reverse chronological analysis, this study examines the advancements introduced by YOLO algorithms, beginning with YOLOv12 and progressing through YOLO11 (or YOLOv11), YOLOv10, YOLOv9, YOLOv8, and subsequent versions to explore each version’s contributions to enhancing speed, detection accuracy, and computational efficiency in real-time object detection. Additionally, this study reviews the alternative versions derived from YOLO architectural advancements of YOLO-NAS, YOLO-X, YOLO-R, DAMO-YOLO, and Gold-YOLO. Moreover, the study highlights the transformative impact of YOLO models across five critical application areas: autonomous vehicles and traffic safety, healthcare and medical imaging, industrial manufacturing, surveillance and security, and agriculture. By detailing the incremental technological advancements in subsequent YOLO versions, this review chronicles the evolution of YOLO, and discusses the challenges and limitations in each of the earlier versions. The evolution signifies a path towards integrating YOLO with multimodal, context-aware, and Artificial General Intelligence (AGI) systems for the next YOLO decade, promising significant implications for future developments in AI-driven applications.},
keywords = {Agriculture, Artificial intelligence, Autonomous vehicles, CNN, Computer vision, Deep learning, Healthcare and medical imaging, Industrial manufacturing, Real-time object detection, Surveillance, Traffic safety, YOLO, YOLO configurations, YOLOv1 to YOLOv12, You Only Look Once},
pubstate = {published},
tppubtype = {article}
}
Ranjan Sapkota; Rizwan Qureshi; Muhammad Usman Hadi; Syed Zohaib Hassan; Ferhat Sadak; Maged Shoman; Muhammad Sajjad; Fayaz Ali Dharejo; Achyut Paudel; Jiajia Li; Zhichao Meng; John Shutske; Manoj Karkee
Multi-Modal LLMs in Agriculture: A Comprehensive Review Journal Article
In: IEEE Transactions on Automation Science and Engineering, vol. 22, pp. 22510–22540, 2025, ISSN: 1558-3783.
Abstract | Links | BibTeX | Tags: Agriculture, Analytical models, ChatGPT, Computational modeling, Computer vision, Data models, Deep learning, Farming, generative artificial intelligence, Hidden Markov models, Large language models (LLMs), Machine Learning, Precision agriculture, Reviews, Training, Transformers, Translation, Vision-language models
@article{sapkota_multi-modal_2025,
title = {Multi-Modal LLMs in Agriculture: A Comprehensive Review},
author = {Ranjan Sapkota and Rizwan Qureshi and Muhammad Usman Hadi and Syed Zohaib Hassan and Ferhat Sadak and Maged Shoman and Muhammad Sajjad and Fayaz Ali Dharejo and Achyut Paudel and Jiajia Li and Zhichao Meng and John Shutske and Manoj Karkee},
url = {https://ieeexplore.ieee.org/document/11173627},
doi = {10.1109/TASE.2025.3612154},
issn = {1558-3783},
year = {2025},
date = {2025-01-01},
urldate = {2025-01-01},
journal = {IEEE Transactions on Automation Science and Engineering},
volume = {22},
pages = {22510\textendash22540},
abstract = {Given the rapid emergence and applications of Multi-Modal Large Language Models (MM-LLMs) across various scientific fields, insights regarding their applicability in agriculture are still only partially explored. This paper conducts an in-depth review of MM-LLMs in agriculture, focusing on understanding how MM-LLMs can be developed and implemented to optimize agricultural processes, increase efficiency, and reduce costs. Recent studies have explored the capabilities of MM-LLMs in agricultural information processing and decision-making. Despite these advancements, significant gaps persist, particularly in addressing domain-specific challenges such as variable data quality and availability, integration with existing agricultural systems, and the creation of robust training datasets that accurately represent complex agricultural environments. Moreover, a comprehensive understanding of the capabilities, challenges, and limitations of MM-LLMs in agricultural information processing and application is still missing. Exploring these areas is crucial to providing the community with a broader perspective and a clearer understanding of MM-LLMs’ applications, establishing a benchmark for the current state and emerging trends in this field. To bridge this gap, this survey reviews the progress of MM-LLMs and their utilization in agriculture, with an additional focus on 11 key research questions (RQs), where 4 RQs are general and 7 RQs are agriculture focused. By addressing these RQs, this review outlines the current opportunities and challenges, limitations, and future roadmap for MM-LLMs in agriculture. The findings indicate that multi-modal MM-LLMs not only simplify complex agricultural challenges but also significantly enhance decision-making and improve the efficiency of agricultural image processing. These advancements position MM-LLMs as an essential tool for the future of farming. For continued research and understanding, an organized and regularly updated list of papers on MM-LLMs is available at https://github.com/JiajiaLi04/Multi-Modal-LLMs-in-Agriculture Note to Practitioners\textemdashMotivated by the need to optimize agricultural practices, this paper investigates the use of Large Language Models (MM-LLMs) to improve efficiency and decision-making in agriculture. We delve into critical RQs to reveal the capabilities and challenges of MM-LLMs, and their potential applications in the agricultural sector. Looking ahead, our findings suggest a promising future for the integration of MM-LLMs in agriculture, potentially revolutionizing how we manage and operate farms.},
keywords = {Agriculture, Analytical models, ChatGPT, Computational modeling, Computer vision, Data models, Deep learning, Farming, generative artificial intelligence, Hidden Markov models, Large language models (LLMs), Machine Learning, Precision agriculture, Reviews, Training, Transformers, Translation, Vision-language models},
pubstate = {published},
tppubtype = {article}
}
2024
Ranjan Sapkota; Dawood Ahmed; Manoj Karkee
Comparing YOLOv8 and Mask R-CNN for instance segmentation in complex orchard environments Journal Article
In: Artificial Intelligence in Agriculture, vol. 13, pp. 84–99, 2024, ISSN: 2589-7217.
Abstract | Links | BibTeX | Tags: Artificial intelligence, Automation, Deep learning, Machine Learning, Machine vision, Mask R-CNN, Robotics, YOLOv8
@article{sapkota_comparing_2024,
title = {Comparing YOLOv8 and Mask R-CNN for instance segmentation in complex orchard environments},
author = {Ranjan Sapkota and Dawood Ahmed and Manoj Karkee},
url = {https://www.sciencedirect.com/science/article/pii/S258972172400028X},
doi = {10.1016/j.aiia.2024.07.001},
issn = {2589-7217},
year = {2024},
date = {2024-09-01},
urldate = {2024-09-01},
journal = {Artificial Intelligence in Agriculture},
volume = {13},
pages = {84\textendash99},
abstract = {Instance segmentation, an important image processing operation for automation in agriculture, is used to precisely delineate individual objects of interest within images, which provides foundational information for various automated or robotic tasks such as selective harvesting and precision pruning. This study compares the one-stage YOLOv8 and the two-stage Mask R-CNN machine learning models for instance segmentation under varying orchard conditions across two datasets. Dataset 1, collected in dormant season, includes images of dormant apple trees, which were used to train multi-object segmentation models delineating tree branches and trunks. Dataset 2, collected in the early growing season, includes images of apple tree canopies with green foliage and immature (green) apples (also called fruitlet), which were used to train single-object segmentation models delineating only immature green apples. The results showed that YOLOv8 performed better than Mask R-CNN, achieving good precision and near-perfect recall across both datasets at a confidence threshold of 0.5. Specifically, for Dataset 1, YOLOv8 achieved a precision of 0.90 and a recall of 0.95 for all classes. In comparison, Mask R-CNN demonstrated a precision of 0.81 and a recall of 0.81 for the same dataset. With Dataset 2, YOLOv8 achieved a precision of 0.93 and a recall of 0.97. Mask R-CNN, in this single-class scenario, achieved a precision of 0.85 and a recall of 0.88. Additionally, the inference times for YOLOv8 were 10.9 ms for multi-class segmentation (Dataset 1) and 7.8 ms for single-class segmentation (Dataset 2), compared to 15.6 ms and 12.8 ms achieved by Mask R-CNN\'s, respectively. These findings show YOLOv8\'s superior accuracy and efficiency in machine learning applications compared to two-stage models, specifically Mask-R-CNN, which suggests its suitability in developing smart and automated orchard operations, particularly when real-time applications are necessary in such cases as robotic harvesting and robotic immature green fruit thinning.},
keywords = {Artificial intelligence, Automation, Deep learning, Machine Learning, Machine vision, Mask R-CNN, Robotics, YOLOv8},
pubstate = {published},
tppubtype = {article}
}



