农业图书情报学报

• •    

大语言模型赋能学术论文质量评价研究

孔青青1, 张颖2, 吴文欣2()   

  1. 1.中国社会科学院 图书馆,北京 100732
    2.中国医学科学院北京协和医学院,医学信息研究所/图书馆,北京 100020
  • 收稿日期:2026-05-09 出版日期:2026-09-11
  • 通讯作者: 吴文欣 E-mail:s2024011031@student.pumc.edu.cn
  • 作者简介:孔青青(1983- ),博士,副研究馆员,研究方向为图书情报学、期刊出版
    张颖(1998- ),博士研究生,研究方向为医学人工智能、医学信息学

Research on Large Language Model-Empowered Academic Paper Quality Evaluation

KONG Qingqing1, ZHANG Ying2, WU Wenxin2()   

  1. 1.Chinese Academy of Social Sciences Library, Beijing 100732
    2.Institute of Medical Information/Medical Library, Chinese Academy of Medical Sciences & Peking Union Medical College, Beijing 100020
  • Received:2026-05-09 Online:2026-09-11
  • Contact: WU Wenxin E-mail:s2024011031@student.pumc.edu.cn

摘要:

[目的/意义] 大语言模型(简称“大模型”)正在介入学术论文质量评价,相关研究分散于不同场景与学科,目前尚缺乏系统整合与规律提炼。本文旨在构建跨场景分析框架,明确大模型在学术评价中的能力边界与合理定位,为后续研究提供参考。 [方法/过程] 采用解释性文献综述方法,检索2022年11月至2026年4月国内外相关数据库文献,经滚雪球追溯与双人独立筛选,围绕同行评审、系统综述自动化与文献质量评分3个应用场景展开分析。 [结果/结论] 大模型在清单核验、信息抽取等结构化任务上表现较可靠,但在整体质量判断与深层批判分析方面存在明显局限,此外还存在出版生态公平性风险与提示注入等安全威胁。系统综述自动化集中在文献检索、数据抽取等信息处理环节,但对偏倚评估与证据分级等高风险判断环节覆盖尚不充分。大模型评分与人类判断总体呈正相关,但有效性存在学科差异,同时存在系统性偏高的宽容偏差,降低了评分的有效区分度。进一步归纳为以下几点主要发现:一是“任务粒度效应”,即模型能力随任务由结构化信息处理向深层学术判断推进而递减;二是系统性偏差跨场景相互印证;三是人机协同仍是当前可行的实践路径。大模型更宜定位为辅助工具,而非替代性裁决机制。未来研究可从标准化评测基准建设、高风险环节人机协同机制、多维偏差综合校正、伦理治理与安全防护体系4个方向加以推进。

关键词: 大语言模型, 学术论文质量评价, 同行评审, 文献质量评分, 人机协同

Abstract:

[Purpose/Significance] Academic evaluation is a central mechanism for quality control in scientific knowledge production and communication. However, conventional reviewer-centered systems have long been constrained by shortages of qualified reviewers, prolonged review cycles, uneven review quality, and the rapidly increasing volume of scholarly publications. Since late 2022, the emergence of large language models (LLMs), represented by GPT-series models, Claude, and DeepSeek, has created new possibilities for improving peer review, literature screening, evidence synthesis, and research quality assessment. However, the involvement of LLMs in academic evaluation is not merely a matter of technical efficiency. It also raises fundamental questions concerning reliability, disciplinary applicability, systematic bias, fairness, accountability, auditability, and security. Existing studies remain fragmented across peer review, systematic-review automation, and literature quality scoring, with different datasets, task definitions, and evaluation criteria. Consequently, there is still a lack of an integrated understanding of where LLMs are dependable, where they fail, and what their role should be. To address this gap, this study developed a cross-context analytical framework covering three representative scenarios: LLM-assisted peer review, automation of the information chain in systematic reviews, and LLM-based quality scoring of published literature. Its main contribution is to connect previously isolated research streams through a progressive question chain concerning what LLMs can do, whether their capabilities transfer across evaluation contexts, and whether their judgments are sufficiently stable and valid to be trusted. By distinguishing structured, fine-grained subtasks from holistic and high-risk academic judgments, the study identified common capability boundaries, recurrent bias patterns, and governance challenges, thereby providing a coherent basis for the responsible use of LLMs in scholarly evaluation. [Method/Process] Given the rapid evolution, conceptual openness, and interdisciplinary distribution of this topic, the study adopted a hermeneutic literature review. This method is appropriate because research on LLM-assisted academic evaluation is dispersed across the fields of natural language processing, information science, scientometrics, and medical informatics. Additionally, the relevant concepts, technical approaches, datasets, and validation standards remain in flux. Guided by the hermeneutic circle, the review proceeded through repeated cycles of reading, interpretation, comparison, and conceptual refinement, allowing the analytical framework and search vocabulary to be adjusted as new evidence emerged. A multi-source search was conducted in Web of Science Core Collection, Scopus, ACL Anthology, IEEE Xplore, PubMed, arXiv, CNKI, and Wanfang Data. Search terms combined expressions relating to LLMs with terms relating to peer review, systematic review, scholarly evaluation, and literature quality. The principal search period extended from November 2022 to April 2026, while earlier influential studies were supplemented through citation tracing. Studies were included when they directly addressed quality judgment in at least one focal scenario and reported empirical evaluation, system design and testing, or substantive theoretical analysis. Studies concerning reviewer - manuscript matching, AI-generated-text detection, writing assistance, as well as automated-review studied involving only conventional machine-learning approaches were excluded. Two researchers independently screened and cross-checked the records. The initial search yielded 86 core studies, including 80 in English and 6 in Chinese. Iterative keyword expansion, citation snowballing, and targeted tracking of key authors added 38 studies. The final corpus comprised 124 publications, including 102 English-language and 22 Chinese-language studies. These studies were compared in terms of task structure, judgment risk, and agreement with human evaluation. [Results/Conclusions] The review reveals a consistent "task-granularity effect" across the three application scenarios. LLMs perform relatively well in structured, fine-grained, and verifiable tasks, such as checklist validation, omission detection, information extraction, title-and-abstract screening, and identification of reporting deficiencies. Their usefulness is particularly evident when tasks can be decomposed into explicit criteria and outputs can be checked against source evidence. In peer review, LLM-generated comments may overlap with human reviewers' comments to an extent comparable to the overlap among human reviewers themselves, and they can assist in detecting missing information and generating preliminary feedback. However, performance declines markedly when tasks require holistic assessment of originality, methodological rigor, theoretical contribution, or broader scholarly significance. In systematic reviews, automation remains concentrated on retrieval, screening, data extraction, and summarization. High-risk stages, such as risk-of-bias assessment and certainty-of-evidence grading, are far less automated because they require contextual interpretation, domain expertise, and accountable judgment. Retrieval-augmented generation can improve traceability and reduce unsupported outputs by linking responses to source passages, but it has not resolved the deeper limitations of judgment-intensive tasks. A second finding is that systematic limitations are corroborated across contexts. LLM-based scores are generally positively correlated with human judgments, but validity varies by discipline, task design, input format, and evaluation protocol. Title-and-abstract inputs may sometimes yield better scoring performance than full-text inputs, while repeated scoring and averaging can improve consistency. Nevertheless, a persistent leniency bias appears across peer review and literature scoring: LLMs tend to assign overly favorable scores, cluster ratings in the middle-to-high range, and rarely use the lowest categories. This compresses score distributions and weakens discrimination between publications of different quality levels. Hallucination further undermines factual reliability, while position bias, verbosity bias, self-enhancement bias, and authority bias may interact with leniency bias in ways that remain insufficiently studied. Large-scale adoption may also reproduce inequalities and pose fairness risks to the scholarly publishing ecosystem. Prompt injection creates an additional security risk because concealed instructions embedded in manuscripts may manipulate model-generated reviews or scores. The evidence therefore supports human-AI collaboration rather than substitution. LLMs are best positioned as assistive instruments for structured, low-risk, and auditable subtasks, whereas high-stakes decisions should remain under the control of qualified human experts. Future research should prioritize four areas: constructing multilingual, multidisciplinary, and task-stratified benchmarks; developing evidence-traceable human-AI workflows for high-risk judgments; analyzing and correcting interacting biases; and establishing ethical governance, accountability, audit trails, and prompt-injection defenses. Overall, this review provides a structured synthesis of an emerging field and clarifies the conditions under which LLMs may contribute to academic evaluation without being treated as autonomous arbiters of scholarly quality.

Key words: Large Language Models(LLMs), academic paper quality evaluation, peer review, literature quality scoring, human-machine collaboration

中图分类号:  G350

引用本文

孔青青, 张颖, 吴文欣. 大语言模型赋能学术论文质量评价研究[J/OL]. 农业图书情报学报. https://doi.org/10.13998/j.cnki.issn1002-1248.26-0289.

KONG Qingqing, ZHANG Ying, WU Wenxin. Research on Large Language Model-Empowered Academic Paper Quality Evaluation[J/OL]. Journal of library and information science in agriculture. https://doi.org/10.13998/j.cnki.issn1002-1248.26-0289.