[关键词]
[摘要]
目的 构建一种面向中药材鉴别的视觉-文本融合模型HerbVLM,提升中药材智能识别在相似类别区分与跨场景泛化中的性能。方法 基于视觉-语言预训练双编码器框架,在文本侧引入中药知识增强提示模块(traditional Chinese medicine knowledge-enhanced prompt module,TKPM),在图像侧构建分层适配与去混淆模块(hierarchical image adaptation and de-confusion module,HIA),并结合动态融合与联合损失进行优化。在Chinese-Medicine-163开源数据集上开展对比与消融实验,并在自采“新浙八味”外部数据集上进行跨域与小样本适配验证。结果 在测试集上,HerbVLM的首位预测准确率(top-1 accuracy,Top-1 Acc)、宏平均精确率(macro precision,Pmacro)、宏平均召回率(macro recall,Rmacro)、宏平均F1值(macro F1-score,F1macro)分别达到98.69%、98.70%、98.71% 和98.70%,优于残差网络(residual network,ResNet)、上下文优化(context optimization,CoOp)、条件上下文优化(conditional context optimization,CoCoOp)、免训练CLIP适配器(training-free CLIP-Adapter,Tip-Adapter)、多模态提示学习(multi-modal prompt learning,MaPLe)、中文视觉-语言对比预训练模型(contrastive vision-language pretraining in Chinese,Chinese-CLIP)等对比方法。消融实验显示,文本增强、图像适配、去混淆约束和动态融合均带来稳定增益。在自采浙八味数据集外部测试中,已知类别灵芝、覆盆子、铁皮石斛、前胡识别率分别为96.00%、100.00%、100.00%和100.00%;未知类别在1-shot、2-shot、5-shot条件下平均准确率分别约为84.00%、91.54%和98.20%,显示出较好的小样本迁移能力。结论 HerbVLM通过图文知识多模态协同建模显著提升了中药材细粒度识别精度与跨域泛化性能,可为中药材教学、质检辅助等场景提供方法支持。
[Key word]
[Abstract]
Objective To construct HerbVLM, a vision-text fusion model for the identification of traditional Chinese medicinal materials, aiming to improve intelligent recognition performance in distinguishing visually similar categories and generalizing across scenarios. Methods Based on a vision-language pretrained dual-encoder framework, a traditional Chinese medicine knowledge-enhanced prompt module (TKPM) was introduced on the text side, while a hierarchical image adaptation and de-confusion module (HIA) was constructed on the image side. Dynamic fusion and joint loss optimization were further incorporated. Comparative and ablation experiments were conducted on the open-source Chinese-Medicine-163 dataset, and cross-domain and few-shot adaptation validation was performed on a self-collected external dataset of the “new Zhebawei”. Results On the test set, HerbVLM achieved top-1 accuracy (Top-1 Acc), macro precision (Pmacro), macro recall (Rmacro), and macro F1-score (F1macro) of 98.69%, 98.70%, 98.71%, and 98.70%, respectively, outperforming comparative methods including residual network (ResNet), context optimization (CoOp), conditional context optimization (CoCoOp), training-free CLIP-Adapter (Tip-Adapter), multi-modal prompt learning (MaPLe), and contrastive vision-language pretraining in Chinese (Chinese-CLIP). Ablation experiments showed that text enhancement, image adaptation, de-confusion constraints, and dynamic fusion all brought stable performance gains. In external testing on the self-collected Zhejiang Eight Flavors dataset, the recognition rates for known categories Lingzhi (Ganoderma), Fupenzi (Rubi Fructus), Tiepishihu (Dendrobii Officinalis Caulis), and Qianhu (Peucedani Radix) were 96.00%, 100.00%, 100.00%, and 100.00%, respectively. For unknown categories, the average accuracies under 1-shot, 2-shot, and 5-shot settings were approximately 84.00%, 91.54%, and 98.20%, respectively, demonstrating favorable few-shot transfer capability. Conclusion HerbVLM significantly improves the fine-grained recognition accuracy and cross-domain generalization performance of traditional Chinese medicinal materials through multimodal collaborative modeling of image-text knowledge, and can provide methodological support for traditional Chinese medicinal material teaching, quality inspection assistance.
[中图分类号]
TP18;R282.5
[基金项目]
浙江省自然科学基金项目/优青项目(LZYQ25H270001);浙江省自然科学基金项目/探索一般(LY24H270007);浙江省中医药管理局中医药健康服务研究计划项目(2026ZF41);2026年浙江省大学生科技创新活动计划(新苗人才计划)项目(2026R410A003)