中文    English

Journal of library and information science in agriculture

   

Research on Large Language Model-Empowered Academic Paper Quality Evaluation

KONG Qingqing1, ZHANG Ying2, WU Wenxin2()   

  1. 1.Chinese Academy of Social Sciences Library, Beijing 100732
    2.Institute of Medical Information/Medical Library, Chinese Academy of Medical Sciences & Peking Union Medical College, Beijing 100020
  • Received:2026-05-09 Online:2026-09-11
  • Contact: WU Wenxin E-mail:s2024011031@student.pumc.edu.cn

Abstract:

[Purpose/Significance] Academic evaluation is a central mechanism for quality control in scientific knowledge production and communication. However, conventional reviewer-centered systems have long been constrained by shortages of qualified reviewers, prolonged review cycles, uneven review quality, and the rapidly increasing volume of scholarly publications. Since late 2022, the emergence of large language models (LLMs), represented by GPT-series models, Claude, and DeepSeek, has created new possibilities for improving peer review, literature screening, evidence synthesis, and research quality assessment. However, the involvement of LLMs in academic evaluation is not merely a matter of technical efficiency. It also raises fundamental questions concerning reliability, disciplinary applicability, systematic bias, fairness, accountability, auditability, and security. Existing studies remain fragmented across peer review, systematic-review automation, and literature quality scoring, with different datasets, task definitions, and evaluation criteria. Consequently, there is still a lack of an integrated understanding of where LLMs are dependable, where they fail, and what their role should be. To address this gap, this study developed a cross-context analytical framework covering three representative scenarios: LLM-assisted peer review, automation of the information chain in systematic reviews, and LLM-based quality scoring of published literature. Its main contribution is to connect previously isolated research streams through a progressive question chain concerning what LLMs can do, whether their capabilities transfer across evaluation contexts, and whether their judgments are sufficiently stable and valid to be trusted. By distinguishing structured, fine-grained subtasks from holistic and high-risk academic judgments, the study identified common capability boundaries, recurrent bias patterns, and governance challenges, thereby providing a coherent basis for the responsible use of LLMs in scholarly evaluation. [Method/Process] Given the rapid evolution, conceptual openness, and interdisciplinary distribution of this topic, the study adopted a hermeneutic literature review. This method is appropriate because research on LLM-assisted academic evaluation is dispersed across the fields of natural language processing, information science, scientometrics, and medical informatics. Additionally, the relevant concepts, technical approaches, datasets, and validation standards remain in flux. Guided by the hermeneutic circle, the review proceeded through repeated cycles of reading, interpretation, comparison, and conceptual refinement, allowing the analytical framework and search vocabulary to be adjusted as new evidence emerged. A multi-source search was conducted in Web of Science Core Collection, Scopus, ACL Anthology, IEEE Xplore, PubMed, arXiv, CNKI, and Wanfang Data. Search terms combined expressions relating to LLMs with terms relating to peer review, systematic review, scholarly evaluation, and literature quality. The principal search period extended from November 2022 to April 2026, while earlier influential studies were supplemented through citation tracing. Studies were included when they directly addressed quality judgment in at least one focal scenario and reported empirical evaluation, system design and testing, or substantive theoretical analysis. Studies concerning reviewer - manuscript matching, AI-generated-text detection, writing assistance, as well as automated-review studied involving only conventional machine-learning approaches were excluded. Two researchers independently screened and cross-checked the records. The initial search yielded 86 core studies, including 80 in English and 6 in Chinese. Iterative keyword expansion, citation snowballing, and targeted tracking of key authors added 38 studies. The final corpus comprised 124 publications, including 102 English-language and 22 Chinese-language studies. These studies were compared in terms of task structure, judgment risk, and agreement with human evaluation. [Results/Conclusions] The review reveals a consistent "task-granularity effect" across the three application scenarios. LLMs perform relatively well in structured, fine-grained, and verifiable tasks, such as checklist validation, omission detection, information extraction, title-and-abstract screening, and identification of reporting deficiencies. Their usefulness is particularly evident when tasks can be decomposed into explicit criteria and outputs can be checked against source evidence. In peer review, LLM-generated comments may overlap with human reviewers' comments to an extent comparable to the overlap among human reviewers themselves, and they can assist in detecting missing information and generating preliminary feedback. However, performance declines markedly when tasks require holistic assessment of originality, methodological rigor, theoretical contribution, or broader scholarly significance. In systematic reviews, automation remains concentrated on retrieval, screening, data extraction, and summarization. High-risk stages, such as risk-of-bias assessment and certainty-of-evidence grading, are far less automated because they require contextual interpretation, domain expertise, and accountable judgment. Retrieval-augmented generation can improve traceability and reduce unsupported outputs by linking responses to source passages, but it has not resolved the deeper limitations of judgment-intensive tasks. A second finding is that systematic limitations are corroborated across contexts. LLM-based scores are generally positively correlated with human judgments, but validity varies by discipline, task design, input format, and evaluation protocol. Title-and-abstract inputs may sometimes yield better scoring performance than full-text inputs, while repeated scoring and averaging can improve consistency. Nevertheless, a persistent leniency bias appears across peer review and literature scoring: LLMs tend to assign overly favorable scores, cluster ratings in the middle-to-high range, and rarely use the lowest categories. This compresses score distributions and weakens discrimination between publications of different quality levels. Hallucination further undermines factual reliability, while position bias, verbosity bias, self-enhancement bias, and authority bias may interact with leniency bias in ways that remain insufficiently studied. Large-scale adoption may also reproduce inequalities and pose fairness risks to the scholarly publishing ecosystem. Prompt injection creates an additional security risk because concealed instructions embedded in manuscripts may manipulate model-generated reviews or scores. The evidence therefore supports human-AI collaboration rather than substitution. LLMs are best positioned as assistive instruments for structured, low-risk, and auditable subtasks, whereas high-stakes decisions should remain under the control of qualified human experts. Future research should prioritize four areas: constructing multilingual, multidisciplinary, and task-stratified benchmarks; developing evidence-traceable human-AI workflows for high-risk judgments; analyzing and correcting interacting biases; and establishing ethical governance, accountability, audit trails, and prompt-injection defenses. Overall, this review provides a structured synthesis of an emerging field and clarifies the conditions under which LLMs may contribute to academic evaluation without being treated as autonomous arbiters of scholarly quality.

Key words: Large Language Models(LLMs), academic paper quality evaluation, peer review, literature quality scoring, human-machine collaboration

CLC Number: 

  • G350

Fig.1

Analytical framework for LLM-empowered academic paper quality evaluation"

Table 1

Selected representative studies on LLM-assisted peer review"

研究类别研究主题实验/数据基础主要发现局限
能力边界与人机协同定位评审子任务能力测试[19]多项子任务实验论文清单验证等核查任务准确率较高整体质量判断任务表现不稳定
评审反馈对比分析[9]Nature系列期刊及ICLR会议4 805篇论文大模型与审稿人反馈的重叠率,与审稿人之间的重叠率大体相当方法设计的深层批判存在明显局限
大模型反馈的真实部署[20]ICLR 2025评审流程,超过20 000条意见收到大模型反馈的审稿人中约27%据此更新评审真实部署中接受度存在差异
清单辅助实验[21]NeurIPS 2024超70%的作者认为大模型清单核验反馈有用以主观评价为主,缺乏客观度量
初审筛选与意见生成[22]某教育技术类期刊可提升初审效率,对论文内在逻辑的评估可信度较高创新性分析能力不足
评审意见生成的优化路径评审意见生成覆盖度扩展[23]大规模评审数据资源生成评审维度提示集合,生成更细致、覆盖更广的评审意见静态生成为主,缺乏交互建模
多轮长上下文对话建模[24]基于ICLR等公开评审记录构建的多轮对话语料将评审重构为多轮交互对话,提供动态建模新范式模拟环境与真实评审有差距
多模态信息融合评审[25]多模态评审生成实验多模态信息与外部知识的融合,使生成式评审更接近专家综合多重证据进行判断的认知过程系统复杂度高,落地难度大

可靠性局限与治理挑战

决策能力与可靠性评测[26]RR-MCQ评测基准单选准确率超60%,但“完全正确”率仅约20%长文处理与批判性反馈短板明显
大模型辅助评审生态影响分析[8]配对比较实验处于录用阈值附近的投稿,录用概率出现统计学显著差异基于大模型识别评审来源的方法存在误判可能
评审文本修改痕迹分析[27]ICLR 2024、NeurIPS 2023大规模文本分析6.5%~16.9%的评审文本经大模型实质性修改难以判断修改对实质内容的影响
真实稿件中隐藏提示现象[28]对预印本平台真实稿件中隐藏提示现象的分析隐藏注入可使评分升高样本限于预印本平台,真实审稿场景验证不足

Fig.2

Risk hierarchy of LLM-assisted peer review"

Fig.3

Automation coverage of the systematic review information chain"

Table 2

Systematic bias types in LLM-based scoring"

偏差类别具体偏差表现机制代表性证据影响后果
输出分布偏差宽容偏差(Leniency Bias)系统性给予偏高评分,集中于中高档,几乎不使用最低档在不同任务情境下均观察到一致倾向[26,46]压缩评分区分度,削弱不同质量层次文献的辨别能力
输入形式偏差位置偏差(Position Bias)在成对比较中倾向于特定位置的回答,与内容质量无关在某些模型中,接近半数的判断在交换顺序后被逆转评分对呈现顺序高度敏感,可重复性低
冗长偏差(Verbosity Bias)将文本长度等同于内容质量,倾向于给较长文本更高评分两项研究均发现,即使回答内容未因篇幅增长而实质性提升,大模型仍倾向于给予更长回答更高的评分长文本被系统性高估,可能激励冗余写作
输入身份偏差自我增强偏差(Self-Enhancement Bias)倾向于高估自身或同系列模型生成内容的质量GPT-4o与Claude 3.5 Sonnet均存在显著自我增强偏差同源模型自评导致系统性偏高,存在生态闭环风险
权威偏差(Authority Bias)评分受期刊式引用、URL链接等虚构权威信号干扰证实添加虚构参考文献(含期刊式引用与URL)等干扰手段可影响质量判断评分易被表面权威信号操控,存在被刻意利用的风险
[1] Dance A. Stop the peer-review treadmill. I want to get off[J]. Nature, 2023, 614(7948): 581-583.
[2] Dimitrova D. Evolution and challenges for peer review[J]. Journalism & Mass Communication Quarterly, 2024, 101(1): 5-12.
[3] Hanson M A, Barreiro P G, Crosetto P, et al. The strain on scientific publishing[J]. Quantitative Science Studies, 2024, 5(4): 823-843.
[4] Hosseini M, Horbach S P J M. Fighting reviewer fatigue or amplifying bias? Considerations and recommendations for use of ChatGPT and other large language models in scholarly peer review[J]. Research Integrity and Peer Review, 2023, 8(1): 4.
[5] Guo Daya, Yang Dejian, Zhang Haowei, et al. DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learning[J]. Nature, 2025, 645(8081): 633-638.
[6] van Dis E A M, Bollen J, Zuidema W, et al. ChatGPT: Five priorities for research[J]. Nature, 2023, 614(7947): 224-226.
[7] Donker T. The dangers of using large language models for peer review[J]. The Lancet Infectious Diseases, 2023, 23(7): 781.
[8] Russo G, Horta Ribeiro M, Davidson T R, et al. The AI review lottery: Widespread AI-assisted peer reviews boost paper scores and acceptance rates[J]. Proceedings of the ACM on Human-Computer Interaction, 2025, 9(7): 1-28.
[9] Liang Weixin, Zhang Yuhui, Cao Hancheng, et al. Can large language models provide useful feedback on research papers? A large-scale empirical analysis[J/OL]. NEJM AI, 2024, 1(8): AIoa2400196. DOI:10.1056/AIoa2400196 .
[10] Ye Rui, Pang Xianghe, Chai Jingyi, et al. Are we there yet? Revealing the risks of utilizing large language models in scholarly peer review[PP/OL]. arXiv (2024-12-02)[2026-03-20]. .
[11] Lo Vecchio N. Personal experience with AI-generated peer reviews: A case study[J]. Research Integrity and Peer Review, 2025, 10(1): 4.
[12] 中国人民大学人文社会科学学术成果评价研究中心. 人文社会科学论文质量评估指标体系实施方案(试行)[R/OL]. 北京: 中国人民大学书报资料中心, 2014[2026-05-21]. .
[13] Zhuang Zhenzhen, Chen Jiandong, Xu Hongfeng, et al. Large language models for automated scholarly paper review: A survey[J]. Information Fusion, 2025, 124: 103332.
[14] 张智雄, 于改红, 刘熠, 等. ChatGPT对文献情报工作的影响[J]. 数据分析与知识发现, 2023, 7(3): 36-42.
Zhang Zhixiong, Yu Gaihong, Liu Yi, et al. The influence of ChatGPT on library & information services[J]. Data Analysis and Knowledge Discovery, 2023, 7(3): 36-42.
[15] Thelwall M. Research quality evaluation by AI in the era of large language models: Advantages, disadvantages, and systemic effects-An opinion paper[J]. Scientometrics, 2025, 130(10): 5309-5321.
[16] Boell S K, Cecez-Kecmanovic D. A hermeneutic approach for conducting literature reviews and literature searches[J]. Communications of the Association for Information Systems, 2014, 34: 257-286.
[17] Paré G, Kitsiou S. Methods for literature reviews[M]//Lau F, Kuziemsky C, editors. Handbook of eHealth Evaluation: An Evidence-Based Approach. Victoria (BC): University of Victoria, 2017: 157-180. (2017-11-13)[2026-05-22]. .
[18] Takkinen P. Post-sustainability: A hermeneutic literature review[J]. The Anthropocene Review, 2025, 12(3): 436-464.
[19] Liu R, Shah N B. ReviewerGPT? an exploratory study on using large language models for paper reviewing[PP/OL]. arXiv (2023-06-01)[2026-03-20]. .
[20] Thakkar N, Yuksekgonul M, Silberg J, et al. A large-scale randomized study of large language model feedback in peer review[J]. Nature Machine Intelligence, 2026, 8(3): 326-336.
[21] Goldberg A, Ullah I, Khuong T G H, et al. Usefulness of LLMs as an author checklist assistant for scientific papers: NeurIPS'24 experiment[PP/OL]. V2. arXiv (2024-11-08)[2026-03-20]. .
[22] 焦丽珍, 刘选, 刘俊杰, 等. 走向人机协同:大语言模型赋能学术论文评审的效能与边界[J]. 重庆高教研究, 2025, 13(5): 119-127.
Jiao Lizhen, Liu Xuan, Liu Junjie, et al. Towards human-machine collaboration: Efficacy and boundaries of large language models in empowering academic paper review[J]. Chongqing Higher Education Research, 2025, 13(5): 119-127.
[23] Gao Zhaolin, Brantley K, Joachims T. Reviewer2: Optimizing review generation through prompt generation[PP/OL]. V2. arXiv (2024-12-02)[2026-03-20]. .
[24] Tan Cheng, Dongxin Lyu, Li Siyuan, et al. Peer review as a multi-turn and long-context dialogue with role-based interactions[PP/OL]. arXiv (2024-06-09)[2026-03-20]. .
[25] Taechoyotin P, Wang G, Zeng T, et al. MAMORX: Multi-agent multi-modal scientific review generation with external knowledge[C/OL]//NeurIPS 2024 Workshop on Foundation Models for Science (FM4Science). Vancouver: NeurIPS, 2024[2026-03-20]. .
[26] Zhou Ruiyang, Chen Lu, Yu Kai. Is LLM a reliable reviewer? A comprehensive evaluation of LLM on automatic paper reviewing tasks[C]//Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024). European Language Resources Association (ELRA) and ICCL, 2024: 9340-9351.
[27] Liang Weixin, Izzo Z, Zhang Yaohui, et al. Monitoring AI-modified content at scale: A case study on the impact of ChatGPT on AI conference peer reviews[C]//Proceedings of the 41st International Conference on Machine Learning. Vienna: PMLR, 2024: 29575-29620.
[28] Lin Zhicheng. Hidden prompts in manuscripts exploit AI-assisted peer review[PP/OL]. arXiv (2025-07-08)[2026-03-20]. .
[29] 张重毅, 牛欣悦, 孙君艳, 等. ChatGPT探析: AI大型语言模型下学术出版的机遇与挑战[J]. 中国科技期刊研究, 2023, 34(4): 446-453.
Zhang Chongyi, Niu Xinyue, Sun Junyan, et al. ChatGPT: Opportunities and challenges of large language models for academic publishing[J]. Chinese Journal of Scientific and Technical Periodicals, 2023, 34(4): 446-453.
[30] Scherbakov D, Hubig N, Jansari V, et al. The emergence of large language models as tools in literature reviews: A large language model-assisted systematic review[J]. Journal of the American Medical Informatics Association, 2025, 32(6): 1071-1086.
[31] 陆伟, 刘寅鹏, 石湘, 等. 大模型驱动的学术文本挖掘——推理端指令策略构建及能力评测[J]. 情报学报, 2024, 43(8): 946-959.
Lu Wei, Liu Yinpeng, Shi Xiang, et al. Large language model-driven academic text mining: Construction and evaluation of inference-end prompting strategy[J]. Journal of the China Society for Scientific and Technical Information, 2024, 43(8): 946-959.
[32] 王亮. 检索增强生成(RAG)驱动的知识服务:原理、范式及评估[J]. 科技与出版, 2025(4): 37-46.
Wang Liang. Retrieval-augmented generation (RAG)-driven knowledge service: Principles, paradigms, and evaluation[J]. Science-Technology & Publication, 2025(4): 37-46.
[33] 王婷, 王娜, 崔运鹏, 等. 基于人工智能大模型技术的果蔬农技知识智能问答系统[J]. 智慧农业(中英文), 2023, 5(4): 105-116.
Wang Ting, Wang Na, Cui Yunpeng, et al. Agricultural technology knowledge intelligent question-answering system based on large language model[J]. Smart Agriculture, 2023, 5(4): 105-116.
[34] Lála J, O'Donoghue O, Shtedritski A, et al. PaperQA: Retrieval-augmented generative agent for scientific research[PP/OL]. V2. arXiv (2023-12-14)[2026-03-20]. .
[35] Asai A, He J, Shao Rulin, et al. Synthesizing scientific literature with retrieval-augmented language models[J]. Nature, 2026, 650(8103): 857-863.
[36] 史忠艳, 雷洁, 孙坦, 等. DeepSeek赋能领域知识图谱低成本构建研究[J]. 农业图书情报学报, 2025, 37(3): 4-17.
Shi Zhongyan, Lei Jie, Sun Tan, et al. Research on DeepSeek-empowered low-cost construction of domain-specific knowledge graphs[J]. Journal of Library and Information Science in Agriculture, 2025, 37(3): 4-17.
[37] ELICIT. Elicit: AI for scientific research[EB/OL]. [2026-02-22]. .
[38] ELICIT. How we evaluated Elicit Systematic Review[EB/OL]. (2025-03-18)[2026-02-22]. .
[39] ELICIT. Systematic Literature Reviews[EB/OL]. [2026-02-22]. .
[40] Nicholson J M, Mordaunt M, Lopez P, et al. Scite: A smart citation index that displays the context of citations and classifies their intent using deep learning[J]. Quantitative Science Studies, 2021, 2(3): 882-898.
[41] Apata O E, Kwok O M, Lee Y H. The use of generative artificial intelligence (AI) in academic research: A review of the consensus app[J]. Cureus, 2025, 17(7): e87297.
[42] Thelwall M, Yaghi A. In which fields can ChatGPT detect journal article quality? An evaluation of REF2021 results[J]. Trends in Information Management, 2025, 13(1): 1-29.
[43] 叶继元, 郭卫兵. 生成式人工智能参与学术评价的反思[J]. 中国社会科学评价, 2024(1): 37-48, 158.
Ye Jiyuan, Guo Weibing. Reflections on the participation of generative AI in academic evaluation[J]. China Social Science Review, 2024(1): 37-48, 158.
[44] 程秀峰, 李嘉琦, 杨金庆, 等. 大语言模型对学术论文评价的可利用性探讨[J]. 图书情报工作, 2024, 68(18): 41-49.
Cheng Xiufeng, Li Jiaqi, Yang Jinqing, et al. Exploring the application of large language models in academic paper evaluation[J]. Library and Information Service, 2024, 68(18): 41-49.
[45] Thelwall M. Evaluating research quality with Large Language Models: An analysis of ChatGPT's effectiveness with different settings and inputs[J]. Journal of Data and Information Science, 2025, 10(1): 7-25.
[46] Du Jiangshu, Wang Yibo, Zhao Wenting, et al. LLMs assist NLP researchers: Critique paper (meta-) reviewing[C]//Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. ACL, 2024: 5081-5099.
[1] LIU Fen. Application Scenarios and Efficiency Improvement of DeepSeek in Library Intelligent QA and Service Consultation [J]. Journal of library and information science in agriculture, 2026, 38(7): 70-81.
[2] YANG Xuejie, LIU Jia, WU Qingxiao, WANG Yufei, GU Dongxiao. Big Data Dynamic Aggregation and Intelligent Service Model for Multimodal Healthcare and Eldercare [J]. Journal of library and information science in agriculture, 2025, 37(4): 24-38.
[3] ZHANG Zhixiong, WANG Yuju, ZHAO Yang. Development Trends of International Open Peer Review Platforms and Recommendations for China [J]. Journal of library and information science in agriculture, 2024, 36(5): 14-22.
[4] WNAG Lingfeng, WANG Shenpeng. Key Parameter Optimization Design of Self-organizing Peer Review in National Preprint Publishing Platform Based on Response Surface Analysis [J]. Journal of library and information science in agriculture, 2023, 35(7): 75-84.
[5] WAN Hao, ZHANG Fujun, LV Qianqian. The Validity of Peer Review Results of DEA Based Super Efficiency Projects [J]. Journal of library and information science in agriculture, 2022, 34(2): 88-101.
[6] HUO Zhenxiang, QU Lichun, LI Xiaoping. Resisting Measures and Thinking to Academic Misconduct from the Perspective of Scientific Journal Editor [J]. Journal of library and information science in agriculture, 2018, 30(7): 137-140.
[7] LI Xiangmin, XU Su, CHEN Xin, TANG Xiaoyan, TAO Wenqi, WANG Changqun. Manuscript Handling Based on the Peer Review Comments:Cases Analysis [J]. Journal of library and information science in agriculture, 2017, 29(12): 159-163.
[8] LI Xiangmin, XU Su, CHEN Xin, WANG Changqun, TANG Xiaoyan, TAO Wenqi. Academic Journals Editor Should Appropriately Handle Revised Manuscripts Which Have Been Peer-Reviewed [J]. Journal of library and information science in agriculture, 2017, 29(10): 142-145.
[9] FANG Rui, GAO Chong-sheng, LI Jing. Quality Control of Three-level Peer Review System for Academic Agricultural Sci-tech Periodicals [J]. Journal of library and information science in agriculture, 2016, 28(4): 160-162.
Viewed
Full text


Abstract

Cited

  Shared   
  Discussed   
No Suggested Reading articles found!