农业图书情报学报

• •    

学术文献细粒度方法知识抽取及其学科驱动特征研究——以经管领域为例

徐浩1,2, 范晓虹3, 邓三鸿2, 封柯3, 朱玲玲1   

  1. 1. 南京工程学院 商学院,南京 211167
    2. 南京大学 信息管理学院,南京 210023
    3. 南京工程学院 管理工程学院,南京 211167
  • 收稿日期:2026-06-09 出版日期:2026-08-19
  • 作者简介:

    徐浩(1989- ),男,副教授,博士,研究方向为信息智能处理与检索

    范晓虹(2003- ),女,硕士研究生,研究方向为学术全文本知识挖掘与知识发现

    邓三鸿(1975- ),男,教授,博士,研究方向为智能信息处理与检索

    封柯(1999- ),男,本科,研究方向为信息智能处理

    朱玲玲(1991- ),女,讲师,博士,研究方向为信息服务与用户行为

  • 基金资助:
    国家社科基金重点项目“基于领域集体智慧挖掘的颠覆性范式变革预测研究”(25ATQ008); 江苏高校哲学社会科学研究重大项目“研究方法的跨学科流动路径及其学科驱动力研究”(2024SJZD066); 国家自然科学基金项目“基于算法管理赋能的双向师徒学习机制重构及其对团队创造力的影响路径研究”(72502104); 南京工程学院教学名师工作室资助项目(苏教师函[2025]28号)

Fine-Grained Extraction of Knowledge of Methods from Academic Literature and Analysis of Its Role in Shaping Disciplinary Characteristics: Evidence from Economics and Management

XU Hao1,2, FAN Xiaohong3, DENG Sanhong2, FENG Ke3, ZHU Lingling1   

  1. 1. School of Business, Nanjing Institute of Technology, Nanjing 211167
    2. School of Information Management, Nanjing University, Nanjing 210023
    3. School of Management and Engineering, Nanjing Institute of Technology, Nanjing 211167
  • Received:2026-06-09 Online:2026-08-19

摘要:

[目的/意义] 复杂科学问题的解决需要融合多学科领域的方法知识,准确识别学术文献所采用的方法知识,阐释方法知识所驱动的学科研究特征,有助于科研人员提高解决问题的效率。 [方法/过程] 构建融合关键技术、核心应用及系统支撑3个维度的统一框架,基于聚类思想,按照方法特征词袋构建、密集簇识别、候选实体识别及细粒度分类与编码步骤完成细粒度方法知识抽取,并以经管领域6种权威期刊验证方法驱动的学科特征。 [结果/结论] 基于聚类思想,融合规则及字典的学术文献细粒度方法实体抽取策略能够有效提升人工标注的效率及准确度,方法实体抽取后的知识关联可进一步揭示方法知识驱动的数量特征、共现关系及科研协作网络。经管领域实证研究结果表明:该领域更倾向于使用方法理论类知识,实证研究、问卷调查和双重差分法等研究方法应用较为广泛,体现领域研究对经验材料获取、数据分析和因果识别的重视,方法知识驱动的科研协作网络呈持续发展态势。整体来看,细粒度方法知识抽取可为学术文本知识组织、研究方法推荐和学科领域特征识别提供支持。

关键词: 研究方法, 实体抽取, 学科驱动特征识别, 特征词密集簇

Abstract:

[Purpose/Significance] The ability to solve complex scientific problems is increasingly dependent on knowledge of methods from multiple disciplines. Research methods, theories, software tools, data resources, and analytical models are essential elements of knowledge production, but they are often scattered across academic texts and expressed in diverse forms. Accurately identifying such knowledge can reduce the burden of literature analysis, clarify how methods are selected and combined, and provide structured evidence for knowledge discovery. Existing studies have mainly focused on method-sentence detection, single-type entity recognition, or frequency analysis, with limited integration between fine-grained extraction and the identification of disciplinary research characteristics. This study therefore develops a unified framework incorporating key technologies, system support, and core applications. We extend method entity research from textual recognition to the analysis of disciplinary knowledge production by connecting method entities with documents, authors, and journals. This supports the organization of academic texts, the recommendation of methods, and the identification of methodological preferences in economics and management. [Method/Process] Guided by clustering, this study proposed a fine-grained entity extraction strategy for Chinese academic literature by integrating a domain dictionary, linguistic rules, statistical representation, similarity calculation, and manual verification. The procedure consists of four stages. First, academic texts were pre-processed, stop words were removed, and method-related feature terms were combined with an initial dictionary to construct a method feature-term bag. Synonymous and semantically equivalent expressions were then merged. Second, a dense-cluster procedure locates text segments in which method-related terms are concentrated, reducing the amount of text requiring examination. Third, regular-expression rules, vector-space representation, and cosine similarity were used to identify and recommend candidate entities. Candidates satisfying the specified threshold were submitted for manual verification. Fourth, confirmed entities were classified into four categories: method and theory, software tool, data, and model. Each entity was assigned a unique code and linked to its source document and metadata. Two annotators verified the candidates according to a unified specification, while disagreements were resolved by a third annotator. The empirical dataset comprises 9 858 articles published from 2011 to 2020 in six journals in economics and management. Frequency statistics, co-occurrence analysis, and social network analysis were applied to identify disciplinary characteristics associated with methodological knowledge. [Results/Conclusions] A total of 8 153 method entities were identified from the titles, abstracts, and keywords of the sampled articles. The proposed strategy assists annotators in locating, screening, normalizing, and classifying candidate entities, improving the efficiency and consistency of large-scale annotation. Linking these entities with bibliographic and author metadata further reveals quantitative distributions, co-occurrence relationships, and scientific collaboration networks. Economics and management research relies more heavily on method-and-theory knowledge than on software tools, data resources, or models. Empirical research, questionnaire surveys, and the difference-in-differences method are widely used, reflecting the field's emphasis on empirical evidence, data analysis, policy evaluation, and causal identification. The results also show that research methods and data resources are frequently used in complementary combinations. The method-driven author collaboration network exhibits a complex structure and a tendency toward continuous development, with some stable groups maintained through shared methods and data-collection practices. Overall, fine-grained method entity extraction can support academic knowledge organization, method recommendation, disciplinary preference identification, and research collaboration analysis. Nevertheless, the approach still depends on expert-defined rules, domain dictionaries, manually selected thresholds, and Chinese lexical information, while its empirical scope is limited to six journals and one publication period. Future research should incorporate semi-supervised learning and contextual semantic representations, extend extraction to full texts, broaden the coverage of languages and disciplines, and evaluate cross-domain transferability. These improvements will enhance the generalizability of methods of knowledge extraction and support deeper analysis of disciplinary evolution and interdisciplinary knowledge production.

Key words: research methods, entity extraction, research feature recognition, intensive clusters of methodological feature words

中图分类号:  G350.7

引用本文

徐浩, 范晓虹, 邓三鸿, 封柯, 朱玲玲. 学术文献细粒度方法知识抽取及其学科驱动特征研究——以经管领域为例[J/OL]. 农业图书情报学报. https://doi.org/10.13998/j.cnki.issn1002-1248.26-0392.

XU Hao, FAN Xiaohong, DENG Sanhong, FENG Ke, ZHU Lingling. Fine-Grained Extraction of Knowledge of Methods from Academic Literature and Analysis of Its Role in Shaping Disciplinary Characteristics: Evidence from Economics and Management[J/OL]. Journal of library and information science in agriculture. https://doi.org/10.13998/j.cnki.issn1002-1248.26-0392.