农业图书情报学报

• •    

多模态AI赋能非遗科学保护与数智化传承的路径研究

韦韧1,2, 安波1,2, 杨华3()   

  1. 1. 中国社会科学院民族学与人类学研究所,北京 100081
    2. 中国社会科学院中国少数民族语言研究中心,北京 100081
    3. 北京师范大学 地理科学学部,北京 100875
  • 收稿日期:2026-03-05 出版日期:2026-08-20
  • 通讯作者: 杨华
  • 作者简介:

    韦韧(1982- ),女,博士,助理研究员,研究方向为地理语言学、少数民族语言文化

    安波 (1986- ),男,博士,副研究员,硕士生导师,中国社会科学院大学,研究方向为自然语言处理、知识图谱

  • 基金资助:
    中国社会科学院数据库专项“学科交叉视域下青藏高原文化数据库”(2024SJK017)

Path to Empowering the Scientific Protection and Digital-Intelligent Inheritance of Intangible Cultural Heritage through Multimodal AI

WEI Ren1,2, AN Bo1,2, YANG Hua3()   

  1. 1. Institute of Ethnology and Anthropology, Chinese Academy of Social Sciences, Beijing 100081
    2. CASS Research Center for Ethnic Minority Languages, Beijing 100081
    3. Faculty of Geographical Science, Beijing Normal University, Beijing 100875
  • Received:2026-03-05 Online:2026-08-20
  • Contact: YANG Hua

摘要:

[目的/意义] 在国家文化数字化战略与人工智能快速发展背景下,非物质文化遗产的保护与传承正由以记录为主转向面向长期管理与活化利用。多模态人工智能可提升非遗图像、音频、视频与文本资料的整理、关联与利用能力,为过程记录、语境表达与多语言处理提供技术支撑。 [方法/过程] 基于图书情报领域的知识组织与数据治理视角,梳理非遗、科学保护、数智化传承与多模态AI等相关概念与内涵,综合国内外研究进展。选取资源聚合、语音采录和动作记录等代表性案例进行比较分析。 [结果/结论] 研究提出,面向非遗的多模态AI应用宜形成贯通数据采集整理、识别转写与结构化、跨载体关联检索、知识整理与知识库建设、基于资料的问答与内容生成的技术路径。数据建设应从媒体存档扩展到工序步骤、关键动作、工具材料与场域信息等过程资料,并保证来源可追溯与片段可定位。技术应用应强化不同载体信息的对齐与证据汇聚,提升检索与引用效率;生成式应用应依托知识库输出并标明来源,减少失真与误用风险。

关键词: 多模态AI, 非物质文化遗产, 数智技术, 图像识别

Abstract:

[Purpose/Significance] Against the background of the national cultural digitalization strategy and the rapid development of artificial intelligence, the protection and transmission of intangible cultural heritage is shifting from documentation-oriented digital recording to long-term management, knowledge-based organization, and active utilization. Intangible cultural heritage is characterized by living transmission, embodied practice, oral communication, contextual dependence, and community participation. Its key knowledge is often distributed across images, audio recordings, videos, textual documents, field notes, tools, materials, ritual spaces, and the experiences of bearers. Traditional digitization methods, which mainly focus on media storage and platform display, are no longer sufficient to support systematic protection, evidence-based research, or practice-oriented transmission. Multimodal artificial intelligence provides a possible technical route for addressing this challenge. By jointly processing image, audio, video, text, three-dimensional, and contextual data, multimodal AI can improve the organization, association, retrieval, interpretation, and reuse of intangible cultural heritage resources. It is especially valuable for documenting procedural knowledge, representing cultural contexts, processing multilingual and dialectal materials, and supporting knowledge services based on traceable evidence. This study aims to clarify how multimodal AI can empower the scientific safeguarding and digital-intelligent transmission of intangible cultural heritage, and to propose an implementation path that is both technically feasible and culturally appropriate. [Method/Process] From the perspective of knowledge organization and data governance in library and information science, this paper first clarifies the connotations and relationships between intangible cultural heritage, scientific safeguarding, digital-intelligent transmission, multimodal AI, and multimodal knowledge data. On this basis, it reviews relevant research progress concerning image and video recognition, speech transcription, text recognition, multimodal retrieval, knowledge graph construction, and generative AI applications. Instead of merely listing existing platforms, the study selects representative cases of resource aggregation, speech and oral tradition collection, and motion recording practice for comparative analysis. These cases correspond to three important scenarios in intangible cultural heritage digitization: the aggregation and public access of heritage resources, the processing of oral and linguistic materials, and the documentation of embodied and procedural knowledge. Through these cases, the paper examines the applicability and limitations of the proposed implementation logic. Particular attention is paid to the fact that most intangible cultural heritage projects do not possess large-scale annotated datasets. Therefore, the study avoids assuming a conventional "large dataset-model training-automatic recognition" route, and instead emphasizes a gradual mechanism based on minimum viable datasets, pretrained models, human-machine collaborative annotation, expert verification, and continuous feedback. [Results/Conclusions] The study proposes that multimodal AI applications for intangible cultural heritage should follow an integrated technical path covering data collection and organization, recognition and transcription, structural processing, cross-media association and retrieval, knowledge organization and knowledge-base construction, and evidence-based question answering and content generation. Data construction should move beyond the storage of media files and include procedural steps, key actions, tool materials, operational rules, field conditions, ritual contexts, and source information. Such data should be traceable and locatable at the segment level. It should also be reusable for later retrieval, citation, verification, and teaching. In the technical processing stage, image and video recognition can support the extraction of patterns, objects, scenes, actions, and process fragments; speech recognition can transform oral narratives, interviews, chants, and dialect materials into searchable texts; text recognition can convert scanned documents, manuscripts, genealogies, and historical records into structured resources. Multimodal retrieval and knowledge graphs can further align different carriers of information and organize them around persons, places, projects, procedures, tools, and events. Generative AI should not be used as unconstrained content production. Instead, it should be embedded in knowledge-base-supported services, with generated answers, explanations, teaching materials, and public-facing content accompanied by verifiable source locations such as page ranges, timestamps, or database records. This can reduce the risks of distortion, cultural misinterpretation, and unauthorized reuse. The paper further argues that, under low-resource conditions, AI applications in intangible cultural heritage should begin with low-risk auxiliary tasks such as speech transcription, video segmentation, preliminary image classification, similar-resource retrieval, metadata completion, and tag recommendation. Outputs generated by AI systems should be reviewed by heritage experts, community participants, and bearers, and the corrected results should be fed back into datasets for model refinement and rule optimization. In this sense, Multimodal AI does not replace bearers or communities; rather, it provides a set of explainable, traceable, and iterative tools for scientific documentation, knowledge organization, transmission support, and protection management. Future research should further explore data standards for procedural knowledge, evaluation methods for culturally sensitive AI outputs, authorization mechanisms, community participation, and sustainable governance models for digital-intelligent intangible cultural heritage protection.

Key words: multimodal AI, intangible cultural heritage, digital-intelligence technology, image recognition

中图分类号:  G122

引用本文

韦韧, 安波, 杨华. 多模态AI赋能非遗科学保护与数智化传承的路径研究[J/OL]. 农业图书情报学报. https://doi.org/10.13998/j.cnki.issn1002-1248.26-0109.

WEI Ren, AN Bo, YANG Hua. Path to Empowering the Scientific Protection and Digital-Intelligent Inheritance of Intangible Cultural Heritage through Multimodal AI[J/OL]. Journal of library and information science in agriculture. https://doi.org/10.13998/j.cnki.issn1002-1248.26-0109.