您好,欢迎访问广西农业科学院 机构知识库!

CDIP-ChatGLM3: A dual-model approach integrating computer vision and language modeling for crop disease identification and prescription

文献类型: 外文期刊

作者: Yan, Changqing 1 ; Liang, Zeyun 1 ; Cheng, Han 1 ; Li, Shuyang 1 ; Yang, Guangpeng 1 ; Li, Zhiwei 1 ; Yin, Ling 2 ; Qu, Junjie 2 ; Wang, Jing 3 ; Wu, Genghong 4 ; Tian, Qi 4 ; Yu, Qiang 4 ; Zhao, Gang 4 ;

作者机构: 1.Shandong Univ Sci & Technol, Coll Intelligent Equipment, Tai An 271019, Peoples R China

2.Guangxi Acad Agr Sci, Guangxi Crop Genet Improvement & Biotechnol Key La, Nanning 530007, Peoples R China

3.China Agr Univ, Coll Resources & Environm Sci, Beijing 100193, Peoples R China

4.Northwest A&F Univ, Coll Soil & Water Conservat Sci & Engn, Yangling 712100, Peoples R China

关键词: Crop disease; Large language model; Fine-tuning; ChatGLM3; Deep learning; Disease identification; Crop protection

期刊名称:COMPUTERS AND ELECTRONICS IN AGRICULTURE ( 影响因子:8.9; 五年影响因子:9.3 )

ISSN: 0168-1699

年卷期: 2025 年 236 卷

页码:

收录情况: SCI

摘要: Deep learning (DL) models have shown exceptional accuracy in plant disease identification, yet their practical utility for farmers remains limited due to a lack of professional and actionable guidance. To bridge this gap, we developed CDIP-ChatGLM3, an innovative framework that synergizes a state-of-the-art DL-based computer vision model with a fine-tuned large language model (LLM), designed specifically for Crop Disease Identification and Prescription (CDIP). EfficientNet-B2, evaluated among 10 DL models across 48 diseases and 13 crops, achieved top performance with 97.97 % +/- 0.16 % accuracy at a 95 % confidence level. Building on this, we fine-tuned the widely used ChatGLM3-6B LLM using Low-Rank Adaptation (LoRA) and Freeze-tuning, optimizing its ability to deliver precise disease management prescriptions. We compared two training strategies-multi-task learning (MTL) and Dual-stage Mixed Fine-Tuning (DMT)-using a different combination of domain-specific and general datasets. Freeze-tuning with DMT led to substantial performance gains, achieving a 33.16 % improvement in BLEU-4 and a 27.04 % increase in the Average ROUGE F-score, surpassing the original model and state-of-the-art competitors such as Qwen-max, Llama-3.1-405B-Instruct, and GPT-4o. The dual-model architecture of CDIPChatGLM3 leverages the complementary strengths of computer vision for image-based disease detection and LLMs for contextualized, domain-specific text generation, offering unmatched specialization, interpretability, and scalability. Unlike resource-intensive multimodal models that blend modalities, our dual-model approach maintains efficiency while achieving superior performance in both disease identification and actionable prescription generation.

  • 相关文献
作者其他论文 更多>>