中文    English

Journal of library and information science in agriculture

   

Formation Logic, Practical Manifestations, and Governance Strategies of Data Bias in Large Language Model-Assisted Academic Research: A Grounded Theory Study in the Humanities and Social Sciences

SUN Yan1, ZHANG Shanshan2()   

  1. 1.School of Marxism, Beijing Institute of Fashion Technology, Beijing 100029
    2.Library, Beijing Institute of Fashion Technology, Beijing 100029
  • Received:2026-06-03 Online:2026-09-24
  • Contact: ZHANG Shanshan E-mail:zss-002@163.com

Abstract:

[Purpose/Significance] Existing studies of data bias have primarily examined training data, model algorithms, and output-level errors, with limited attention to how scholarly information is selected, reconstructed, confirmed, and reused after large language models (LLMs) are adopted in research. In LLM-assisted research, literature, search results, dialogue records, summaries, and generated texts are repeatedly processed and reintegrated into knowledge production. Bias may therefore arise not only before or within model generation, but also through human–LLM interaction across research tasks. To address this gap, this study adopts a dynamic view of data in knowledge production and extends the analysis of data bias from training data and model outputs to the socio-technical process of academic research. It examines the practical forms, generative conditions, and governance of data bias in the humanities and social sciences, while distinguishing observable bias forms from the technical, behavioral, and contextual conditions that shape them, and highlighting researchers as both potential amplifiers and active correctors. [Method/Process] Grounded theory was adopted to identify recurring patterns and category relations from actual research practice rather than test a predetermined causal model. Empirical materials were collected in two stages from 56 participants, including 8 undergraduates, 18 master's and doctoral students, 26 university teachers, and 4 journal editors. The first stage, conducted from December 2023 to December 2024, used purposive sampling and generated materials from 41 participants. Open coding and preliminary categorization identified initial concepts and relations requiring further elaboration. The second stage, conducted in July 2026, added 15 theoretically sampled participants and focused on underdeveloped relations involving human-LLM trust and use states, researcher capability and task configuration, AI-text detection, academic-community use norms, and disciplinary differences. Second-stage materials were compared item by item with the concepts, properties, and category relations formed in the first stage. Researcher practice records and public online discussions served as supplementary materials for constant comparison. Open, axial, and selective coding produced 86 initial concepts, 12 initial categories, and a theoretical model. Theoretical saturation was reached when later second-stage materials produced no new concepts, categories, or category relations. No fixed numerical frequency threshold was imposed. Relative stability or repeatability was identified through recurrent experiences within the same participant, convergent accounts across participants undertaking similar tasks, or recurrence across different data sources or research stages. Isolated anomalies, including one-off hallucinations, were retained as risk records unless repeated directional patterns emerged. [Results/Conclusions] Four parallel forms of data bias were identified. Knowledge-resource visibility bias concerns recurring differences in which literature, sources, or information become visible or remain omitted during retrieval and AI-assisted reading. Knowledge-reconstruction bias refers to relatively stable or repeatable shifts in information completeness, conceptual boundaries, argumentative strength, or modes of expression during summarization, explanation, translation, or rewriting. Human-LLM interaction feedback bias develops through sustained dialogue when model tendencies interact with researchers' questioning, trust, fatigue, and acceptance, thereby reinforcing particular viewpoints, problem frames, or judgment paths. Knowledge-production recirculation bias emerges when AI-mediated content, expression patterns, detection practices, and community norms are repeatedly incorporated into writing, evaluation, revision, imitation, circulation, and reuse. The four forms do not constitute a fixed sequence. They may enter at different points, interact with one another, and exhibit a possible mechanism of multiple entry points, non-linear formation, and cyclical accumulation. Technical, behavioral, and contextual conditions shape whether such tendencies are accepted, reinforced, corrected, or rejected. Researchers may amplify bias through early delegation of core tasks, accumulated trust, sustained interaction, and fatigue, while source verification, explicit correction, retention of core questions and judgments, and control of human-LLM collaboration boundaries can weaken or interrupt it. Governance should improve academic knowledge-resource provision and traceability, enhance transparency in model-based knowledge processing, strengthen researcher agency and collaboration boundaries, and develop revisable community norms for disclosure, detection, evaluation, and responsibility. The model remains exploratory: the sample is limited to the humanities and social sciences, much of the evidence is retrospective, and the two data-collection stages span a period of rapid LLM development. Future studies should combine longitudinal tracking, behavioral observation, experiments, and quantitative testing to examine boundary conditions and change over time.

Key words: large language models, academic research, data bias, knowledge production, human - LLM interaction, grounded theory

CLC Number: 

  • G203

Table 1

Basic information of interviewees"

受访者类型人数/人编码方式主要特征
合计56-覆盖不同学术角色与使用场景
本科生8B01~B08拥有在课程论文或毕业论文中使用AI的经验
硕博研究生18G01~G18拥有在科研阅读、论文写作中使用AI的经验
高校教师26T01~T26拥有在科研与论文写作中使用AI的经验
期刊编辑4E01~E04具有AI生成论文评审经验

Fig.1

Research process"

Table 2

Representative examples of open coding and category extraction"

初始范畴代表性概念原始资料举例
F1 学术资源获取f11 文献推荐偏好“AI帮我搜索外文文献和翻译还是很方便的,但是它似乎有它的喜好,很多我觉得重要的文献一篇没找出来,所以时常还要返工”(T01)
F1 学术资源获取f15 关键信息遗漏“让AI帮我解读论文确实比较快,但有时候它概括完以后,我会错过论文里很重要的段落”(G02)
F2 学术信息真实性f21 虚假文献生成“让AI帮我写了一段文献综述,后来一查,里面有些文献根本不存在,真的把我吓到了”(G03)
F2 学术信息真实性f23 来源追溯失败“学生用AI搞出来一堆‘百分之多少的人认为’这样的句子,我问她这个百分比从哪里来的,她自己也说不出来”(T05)
F3 知识整理与认知辅助f31 文献脉络梳理“让AI进行文献综述的梳理还是挺有帮助的,至少它可以先把材料大致分析一下,让我知道这些研究在讨论什么”(T04)
F3 知识整理与认知辅助f35 灵感激发“在启发灵感上,它确实有作用。有时候我本来只是模模糊糊有一个感觉,和它聊一聊,问题会慢慢显出来”(T17)
F4 学术表达重构f41 观点稀释“我记性不太好,有些重要的内容会反复说。文章让AI润色的遍数多了以后,我发现自己的观点反而被消解了”(T06)
F4 学术表达重构f43 表达同质化“文章经过它反复润色以后,风格越来越像知网上其他类似的论文。看起来更完整了,但也越来越不像我自己”(T07)
F5 模型生成倾向f51 观点迎合“它有时候特别优秀,让我觉得大开眼界,但是它也经常迎合我的观点。聊到最后,我有时候反而不知道自己的想法到底有没有学术价值”(T09)
F5 模型生成倾向f53 模型框架牵引“有时候我已经有自己的思路了,但AI没有顺着我的问题继续往下推,反而总想让我按照它给出的框架走”(T18)
F6 问题意识与思维路径f62 无意识偏移“被AI带跑本身或许没有那么可怕,真正可怕的是,你根本没有意识到自己的问题已经发生了变化”(T19)
F6 问题意识与思维路径f67 主动拒绝与校正“它出现明显错误或者回答没有深度的时候,是说服不了我的。我会停下来,重新判断这个方向是不是有问题”(T12)
F7 人机信任与使用状态f71 连续成功形成信任“即使一开始很警惕,用过几次以后,只要它一直没有出错,人就会慢慢相信它。它替你做得越多,自己越容易放松”(R02)
F7 人机信任与使用状态f72 疲劳情境顺从“当我很累、又不得不继续干的时候,是最容易被AI带跑的,也是论文最容易缺少人味的时候”(R01)
F8 主体能力与介入边界f81 主体知识约束“AI确实能让我思维活跃,但你真想从它那里套出点什么,前提还是自己得有货。它只能在你的水平上帮助你,很难超过你的判断能力”(T11)
F8 主体能力与介入边界f86 核心构思保留“论文最开始的构思和最后的润色,我觉得还得自己来。尤其是核心问题和真正想表达的东西,不能直接交给AI”(T20)
F9 科研流程与任务分工f91 AI前置写作“我们的论文都是先AI再人改,我觉得这没什么不能说的,很多同学都这样”(B02)
F9 科研流程与任务分工f94 核心框架自主“需要出框架、出逻辑的内容,还是自己去做”(T10)
F10 AI文本识别与检测f102 AI率失真“自己写的AI率有点高,别人用AI写的,再让AI降低AI率,结果反而为零”(N02)
F10 AI文本识别与检测f104 AI降AI循环“有些文章本来就是AI写的,后来又用AI去降AI率”(N03)
F11 学术共同体使用文化f111 AI使用常态化“现在查论文、找资料,我第一反应不是先去百度了,而是先问AI,感觉很多东西它都能给你说一点”(B02)
F11 学术共同体使用文化f112 AI使用隐匿“AI就像房间里的大象,或多或少大家都在使用,但都不说,好像谁用了AI写作,谁就不是在做学术”(T14)
F12 人文社会科学学科情境差异f121 社会学数据优势“社会学还好,CGSS、CFPS这些数据库AI都能处理”(T15)
F12 人文社会科学学科情境差异f123 教育学学理深度不足“教育学写出来都是综而不述、泛而不深,信息密度很低”(T21)

Table 3

Theoretical structure of bias forms, generative conditions, and mitigating mechanisms"

分析层级理论要素及对应编码范畴内涵
偏见形态B1 知识资源可见性偏见:F1 学术资源获取AI介入检索、推荐和辅助阅读,使文献及其内部信息获得不同呈现机会,形成知识可见范围与注意分配的稳定或可重复偏移
偏见形态B2 知识重构偏见:结果范畴——F2 学术信息真实性、F4 学术表达重构;任务场景——F3 知识整理与认知辅助F2中经重复性证据支持的方向性真实性偏移与F4主要呈现偏见结果;单次幻觉仍作风险记录。F3呈现AI参与整理、概括、补充和认知加工的任务场景,其正向辅助本身不构成偏见
偏见形态B3 人机交互反馈偏见:F5 模型生成倾向;F6 问题意识与思维路径模型回应与研究者提问、判断和反馈持续牵引,使已有立场、问题框架或观点强度出现持续或可重复的强化或方向性偏移;研究者的主动修正归入反向调节机制
偏见形态B4 知识生产回流偏见:F10 AI文本识别与检测;F11 学术共同体使用文化内容进入传播与再利用本身不构成偏见;只有特定的信息偏向、表达模式或使用规则在共同体中被反复接受和强化,并持续影响后续知识生产时,才构成B4
生成条件-技术技术条件:模型知识覆盖、检索/内容选择机制、概率生成与上下文处理特征影响学术信息被呈现、加工和回应的方式,为四类偏见整体提供模型侧条件
生成条件-行为行为条件:F7 人机信任与使用状态;F8 主体能力与介入边界;F9 科研流程与任务分工影响偏见被接受、放大、修正或拒绝;信任、疲劳、能力和任务配置本身不直接等同于数据偏见
生成条件-情境情境条件:F12 人文社会科学学科情境差异学科知识结构、材料形态与表达规范影响AI适用范围、偏见表现及共同体评价
反向调节机制主动信息核查;主动拒绝与校正;核心构思与核心判断保留;核心框架自主;协作边界控制研究者通过返回原始材料、保留问题意识与判断责任、限制AI介入核心任务,中断偏见的持续强化

Fig.2

Practice-based formation model of data bias in large language model-assisted academic research"

[1] Ferrara E. Fairness and bias in artificial intelligence: A brief survey of sources, impacts, and mitigation strategies[J]. Sci, 2024, 6(1): 3.
[2] Mehrabi N, Morstatter F, Saxena N, et al. A survey on bias and fairness in machine learning[J]. ACM Computing Surveys, 2022, 54(6): 1-35.
[3] Suresh H, Guttag J. A framework for understanding sources of harm throughout the machine learning life cycle[C]//Equity and Access in Algorithms, Mechanisms, and Optimization. New York: ACM, 2021: 1-9.
[4] Lopez P. Bias does not equal bias: A socio-technical typology of bias in data-based algorithmic systems[J]. Internet Policy Review: Journal on Internet Regulation, 2021, 10(4): 1-29.
[5] 李沛雨. 人工智能数据偏见风险的法律规制[J]. 东南法学, 2025(2): 235-255.
Li Peiyu. Legal regulation of data bias risk of artificial intelligence[J]. Southeast Law Review, 2025(2): 235-255.
[6] Dwivedi Y K, Kshetri N, Hughes L, et al. Opinion Paper: "So what if ChatGPT wrote it?" Multidisciplinary perspectives on opportunities, challenges and implications of generative conversational AI for research, practice and policy[J]. International Journal of Information Management, 2023, 71: 102642.
[7] 蔡芬, 贾枭, 沈文钦. 生成式人工智能在我国研究生学术写作中的应用现状及其影响[J]. 中国高教研究, 2025(1): 75-82.
Cai Fen, Jia Xiao, Shen Wenqin. The application status and impact of artificial intelligence generated content in postgraduate academic writing in China[J]. China Higher Education Research, 2025(1): 75-82.
[8] Ji Ziwei, Lee N, Frieske R, et al. Survey of hallucination in natural language generation[J]. ACM Computing Surveys, 2023, 55(12): 1-38.
[9] Walters W H, Wilder E I. Fabrication and errors in the bibliographic citations generated by ChatGPT[J]. Scientific Reports, 2023, 13: 14045.
[10] Farquhar S, Kossen J, Kuhn L, et al. Detecting hallucinations in large language models using semantic entropy[J]. Nature, 2024, 630(8017): 625-630.
[11] Bender E M, Gebru T, McMillan-Major A, et al. On the dangers of stochastic parrots: Can language models be too big?[C]//Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency. New York: ACM, 2021: 610-623.
[12] 张宁, 袁勤俭. 用户视角下的学术社交网络信息质量影响因素研究——基于扎根理论方法[J]. 图书情报知识, 2018, 35(5): 105-113.
Zhang Ning, Yuan Qinjian. The influence factors of information quality in academic social networks from user' perspective based on grounded theory[J]. Document, Information & Knowledge, 2018, 35(5): 105-113.
[13] Glaser B G, Strauss A L. The Discovery of Grounded Theory: Strategies for Qualitative Research[M]. Chicago: Aldine, 1967.
[14] Corbin J, Strauss A. Basics of Qualitative Research (3rd ed.): Techniques and Procedures for Developing Grounded Theory[M]. 2455 Teller Road, Thousand Oaks, California 91320 United States: SAGE Publications, Inc., 2008.
[15] 张一帆, 陈祖琴, 葛继科, 等. 面向突发事件识别与分类的多模态数据集构建研究[J]. 农业图书情报学报, 2024, 36(10): 76-85.
Zhang Yifan, Chen Zuqin, Ge Jike, et al. Construction of a multimodal dataset for emergency event identification and classification[J]. Journal of Library and Information Science in Agriculture, 2024, 36(10): 76-85.
[16] Sharma M, Tong M, Korbak T, et al. Towards understanding sycophancy in language models[C/OL]//The Twelfth International Conference on Learning Representations (ICLR 2024). Vienna: ICLR, 2024[2026-07-24]. .
[17] Lee J D, See K A. Trust in automation: Designing for appropriate reliance[J]. Human Factors, 2004, 46(1): 50-80.
[18] Buçinca Z, Malaya M B, Gajos K Z. To trust or to think: Cognitive forcing functions can reduce overreliance on AI in AI-assisted decision-making[J]. Proceedings of the ACM on Human-Computer Interaction, 2021, 5(CSCW1): 1-21.
[19] 涂良川, 唐春燕. 人工智能“技术自主性”的哲学叙事[J]. 学术论坛, 2025, 48(5): 1-11.
Tu Liangchuan, Tang Chunyan. The philosophical narrative of "technological autonomy" in artificial intelligence[J]. Academic Forum, 2025, 48(5): 1-11.
[20] Weber-Wulff D, Anohina-Naumeca A, Bjelobaba S, et al. Testing of detection tools for AI-generated text[J]. International Journal for Educational Integrity, 2023, 19(1): 26.
[21] Liang Weixin, Yuksekgonul M, Mao Yining, et al. GPT detectors are biased against non-native English writers[J]. Patterns, 2023, 4(7): 100779.
[22] UNESCO. Guidance for generative AI in education and research[R]. Paris: UNESCO, 2023.
[1] WANG Chao, CHEN Jie, HOU Hui. Human-AI Configuration Differences in Online Reference Services between Chinese and International Academic Libraries: Evidence from "Double First-Class" and U.S. News Top 100 Universities [J]. Journal of library and information science in agriculture, 2026, 38(9): 57-67.
[2] LIU Siyi, LIU Guifeng, LIU Qiong, HAN Muzhe. Construction of a Data Element Value Release Model from the Perspective of Value Co-creation: A Grounded Analysis Based on Typical Application Scenarios of "Data Element ×" [J]. Journal of library and information science in agriculture, 2026, 38(9): 4-15.
[3] QIAN Li, YANG Yanxi, ZHANG Yuanzhe, HU Maodi, CHANG Zhijun. The Impacts and Implications of OpenClaw for Scientific and Technical Literature Intelligence Work [J]. Journal of library and information science in agriculture, 2026, 38(4): 4-12.
[4] HU Anqi. Construction of an Artificial Intelligence Literacy Ability Framework and Training System for College Students [J]. Journal of library and information science in agriculture, 2026, 38(2): 42-55.
[5] REN Fubing, LUO Ya. Evolution Mechanism of User's Network Cluster Behavior from the Perspective of Cognitive Bias [J]. Journal of library and information science in agriculture, 2025, 37(9): 18-31.
[6] ZHANG Li, WANG Bo, JING Shui. Generative AI-Driven Resource Discovery in Public Libraries: Service Optimization Based on a Dynamic Evaluation Model [J]. Journal of library and information science in agriculture, 2025, 37(5): 58-71.
[7] SANG Yuanyuan. Multimodal Learning Technology Aimed at Exploring the Innovative Path of Library Intelligence Service [J]. Journal of library and information science in agriculture, 2025, 37(3): 42-52.
[8] CAI Yiran, HU Zhengyin, LIU Chunjiang. Analysis of Progress in Data Mining of Scientific Literature Using Large Language Models [J]. Journal of library and information science in agriculture, 2025, 37(2): 4-22.
[9] Haoxian WANG, Ziming ZHOU, Feifei DING, Chengfu WEI. Digital Humanities & Large Language Models: Practice and Research in Semantic Retrieval of Ancient Documents [J]. Journal of library and information science in agriculture, 2024, 36(9): 89-101.
[10] Jia XU. The Optimization Path of VR Red Resource Reading Style under the Embodied Cognition: An Exploratory Study Based on Grounded Theory [J]. Journal of library and information science in agriculture, 2024, 36(11): 33-46.
[11] GUO Pengrui, WEN Tingxiao. Research of the Impact of LLMs on Information Retrieval Systems and Users' Information Retrieval Behavior [J]. Journal of library and information science in agriculture, 2023, 35(11): 13-22.
[12] ZHANG Hongping. Promotion Mechanism of User Demand to Library Literature Resource Construction, Management and Service: Based on Grounded Theory [J]. Journal of library and information science in agriculture, 2022, 34(9): 95-103.
[13] GUO Weijia. Influencing Factors of Artificial Intelligence Readiness in Libraries [J]. Journal of library and information science in agriculture, 2022, 34(5): 47-56.
[14] LYU Kun, CHEN Yaoyao, XIANG Minhao. Influencing Factors of Sina Weibo Information Service Quality from the Perspective of Users: Exploratory Analysis Based on Grounded Theory [J]. Journal of library and information science in agriculture, 2022, 34(10): 70-81.
[15] LI Chao, LIU Zijie, LIU Ahui, ZHU Xuekun, FAN Zhenjia. The Modeling of the Formation Mechanism of Consumer Intention in Short Video Marketing [J]. Journal of library and information science in agriculture, 2020, 32(10): 35-46.
Viewed
Full text


Abstract

Cited

  Shared   
  Discussed   
No Suggested Reading articles found!