关怀版

档案信息化2026年第40卷第5期《档案学研究》

档案领域大模型训练中数据污染的治理困境与因应之策

The Challenges and Countermeasures for Addressing Data Contamination in Large Language Model Training for the Archival Domain

邢张睿

XING Zhangrui

苏州大学王健法学院 苏州 215006

出版日期2026-10-28卷期第40卷 第5期页码134-142DOI10.16065/j.cnki.issn1002-1620.2026.05.015浏览次数5 次

摘要

领域大模型的应用正驱动档案工作和档案事业迈向智能化新阶段,但其性能与应用会受训练过程中存在的数据污染的影响。档案领域大模型训练中的数据污染是指在模型开发的训练数据阶段,因原始档案数据存在质量缺陷、标注偏差,人为恶意注入偏见或基准数据不当泄露至训练集等原因,大模型无法准确学习与理解档案信息的原始性、客观性及其上下文关联,最终形成系统性认知偏差的风险。为了应对档案领域大模型的数据污染,可从数据来源、数据质量、数据测试三个层面进行治理,包括建立差异化的知识源审查与众包数据标注机制,提升合成数据透明度并加强重复数据的识别与清洗,以及区分使用公私基准数据集并优化基于记忆规律的数据污染治理策略。

关键词:档案数据领域大模型基准数据数据污染数据训练

Abstract

The application of domain-specific large language model is driving archival work and the archival sector toward a new stage of intelligence. However, their performance and application are constrained by the potential impact of data contamination during the training process. In the context of training archival domain large language models, data contamination refers to issues arising during the training data preparation phase—specifically, defects in the quality of raw archival data, annotation errors, the malicious injection of human biases, or the leakage of benchmark data into the training set. These issues prevent the large language model from accurately learning and comprehending the authenticity, objectivity, and contextual relationships inherent in archival information, thereby creating a risk of systemic cognitive bias. To mitigate data contamination in archival domain large language models, governance strategies should be implemented across three key dimensions: data sources, data quality, and data testing. These strategies include establishing differentiated mechanisms for reviewing knowledge sources and crowdsourced data annotation; enhancing the transparency of synthetic data while strengthening the identification and cleansing of duplicate data; and, finally, distinguishing between and optimizing the use of public versus private benchmark datasets, alongside refining data contamination governance strategies based on memory-centric principles.

Key words: archival data; domain-specific large language models; benchmark data; data contamination; data training

引用格式

邢张睿. 档案领域大模型训练中数据污染的治理困境与因应之策[J]. 档案学研究, 2026, 40(5): 134-142.
XING Zhangrui. The Challenges and Countermeasures for Addressing Data Contamination in Large Language Model Training for the Archival Domain. Archives Science Study, 2026, 40(5): 134-142.

参考文献

展开查看参考文献(43 条)
[1] [7]王毅, 陈文汇, 刘红霞. 基于AI大模型的档案信息资源开发:逻辑机理、应用赋能与推进策略[J]. 山西档案, 2025(8):2-3,1. [2] 李文锋, 刘峰. DeepSeek在工程建设项目档案真实性检查中的应用思考[J]. 陕西档案, 2025(5):24. [3] 中国网. 全国首个辅助决策地方政务大模型落地北京电信助力首都政务档案智慧化跃升[EB/OL].(2025-06-16)[2025-09-25]. https://tech.chinadaily.com.cn/a/202506/16/WS684fb3a4a3102053770383de.html. [4] 刘倩倩, 刘圣婴, 刘炜. 图书情报领域大模型的应用模式和数据治理[J]. 图书馆杂志, 2023 (12):22. [5] 郭之璇, 苏宇. 领域模型的算法透明义务[J]. 人工智能, 2025(4):92. [6] 王建品, 尚照辉. 科技档案管理中垂直大模型应用场景、问题与对策[J]. 档案管理, 2025(4):63. [7] 吴海中. 档案部门应用大语言模型的典型案例分析[J]. 山西档案, 2025(3):152. [8] 王昊贤, 周子茗, 丁菲菲, 等. 数字人文与大语言模型:古文献语义检索实践与探索[J]. 农业图书情报学报, 2024(9):92. [9] 张中泽. 高质量推进红色档案资源保护管理和开发利用的实践与思考[J]. 山西档案, 2025(11):42. [10] Online Computer Library Center. Recommendations of the OCLC Research Library Partnership Web Archiving Metadata Working Group[EB/OL].(2018-02)[2025-10-08]. https://www.oclc.org/research/publications/2018/oclcresearch-descriptive-metadata/recommendations.html. [11] 张军荣. 人工智能生成物的溢出风险及监管规则研究[J]. 行政法学研究, 2025(2):131-132. [12] 刘晓春. 数据功能类型视角下数据污染的治理维度[J]. 人民司法, 2024(10):16. [13] 曾庆醒. 涌现效应下的生成式人工智能数据污染及其治理路径[J]. 法学杂志, 2024(5):53. [14] 陈俊秀, 徐玉琴. 人工智能大语言模型引发的数据污染风险及其规制路径[J]. 大连理工大学学报(社会科学版), 2025(4):55. [15] 刘福元. 电竞人工智能的数据污染风险与治理规则建构[J]. 首都体育学院学报, 2025(2):126-128. [16] Sainz O, Campos J, García-Ferrero I, et al. NLP evaluation in trouble: on the need to measure llm data contamination for each benchmark[C]// Bouamor H, Pino J, Bali K. Findings of the Association for Computational Linguistics: EMNLP 2023. Singapore:Association for Computational Linguistics, 2023: 10776-10787. [17] Pan Y, Pan L, Chen W, et al. On the risk of misinformation pollution with large language models[C]// Bouamor H, Pino J, Bali K. Findings of the association for computational linguistics: EMNLP 2023. Singapore:Association for Computational Linguistics, 2023: 1389-1403. [18] 蒋冠, 刘子焱. 大模型技术与档案编研业务深度融合的障碍与应对策略[J]. 山西档案, 2025(9):3. [19] [29]Lee K, Ippolito D, Nystrom A, et al. Deduplicating training data makes language models better[C]//Muresan S, Nakov P, Villavicencio A. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics. Dublin, Ireland: Association for Computational Linguistics, 2022: 8424-8445. [20] Roberts M, Thakur H, Herlihy C, et al. Data contamination through the lens of time[PP/OL]. arXiv(2023-10-16)[2025-12-10]. https://doi.org/10.48550/arXiv.2310.10628. [21] 安小米, 龙志奇, 邝苗苗. 标准化视角下大模型数据治理的理论框架及其构成要素研究[J]. 情报资料工作, 2024(6):79. [22] 裴宏娇, 郭瑞宇, 杨志和. 图书馆AI语言生成中的信息幻觉:信息素养视角下的风险规避策略[J]. 图书情报导刊, 2024(12):37. [23] Wiggers K. MIT study finds "systematic" labeling errors in popular AI benchmark datasets[EB/OL].(2021-03-29)[2025-11-16]. https://venturebeat.com/ai/mit-study-finds-systematic-labeling-errors-in-popular-ai-benchmark-datasets. [24] Suresh H, Guttag J. A framework for understanding sources of harm throughout the machine learning life cycle[C]// Proceedings of the 1st ACM Conference on Equity and Access in Algorithms, Mechanisms, and Optimization. New York: Association for Computing Machinery, 2021: 1-9. [25] Offenhuber D. Shapes and frictions of synthetic data[J/OL]. Big Data & Society, 2024(2). https://doi.org/10.1177/20539517241249390. [26] Giuffrè M, Shung D L. Harnessing the power of synthetic data in healthcare: innovation, application, and privacy[J]. NPJ digital medicine, 2023(1): 186. [27] 董敏. 论县级综合档案馆重复档案的精简工作[J]. 山东档案, 2020(6):51. [28] Jiang M, Liu K Z, Zhong M, et al. Investigating data contamination for pre-training language models[PP/OL]. arXiv(2024-01-11)[2025-11-15]. https://doi.org/10.48550/arXiv.2401.06059. [29] Brown T, Mann B, Ryder N, et al. Language models are few-shot learners[J]. Advances in neural information processing systems, 2020(33): 1877-1901. [30] OpenAI. GPT-4 Technical Report[EB/OL].(2023-03-27)[2025-11-28]. https://cdn.openai.com/papers/gpt-4.pdf. [31] 赵铁军, 许木璠, 陈安东. 自然语言处理研究综述[J]. 新疆师范大学学报(哲学社会科学版), 2025(2):89-90. [32] Bender E M, Gebru T, McMillan-Major A, et al. On the dangers of stochastic parrots: can language models be too big?[C]// Proceedings of the 2021 ACM conference on fairness, accountability, and transparency. New York: Association for Computing Machinery, 2021: 610-623. [33] Magar I, Schwartz R. Data contamination: from memorization to exploitation[C]// Muresan S, Nakov P, Villavicencio A.Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics. Dublin, Ireland: Association for Computational Linguistics, 2022: 157-165. [34] Jo E S, Gebru T. Lessons from archives: strategies for collecting sociocultural data in machine learning[C]// Proceedings of the 2020 conference on fairness, accountability, and transparency. New York: Association for Computing Machinery, 2020: 306-316. [35] 刘力超, 陈晓珑, 牛力. 大模型驱动档案开放智能审核方法研究:动因、框架与实践[J]. 档案学研究, 2025(2):136. [36] 胡泳, 张文杰. 数据标注治理:可信人工智能的后台风险与治理转向[J]. 云南社会科学, 2024(6):34-35. [37] Jordon J, Szpruch L, Houssiau F, et al. Synthetic data—what,why and how?[PP/OL].[arXiv(2022-05-06)[2025-12-14]. https://doi.org/10.48550/arXiv.2205.03257. [38] Shilov I, Meeus M, de Montjoye Y A. The mosaic memory of large language models[J]. Nature Communications, 2026(1): 2142. [39] Leveling J, Helmer L, Stein B J, et al. Evaluation of Document Deduplication Algorithms for Large Text Corpora[C]//Nicosia G, Ojha V, Giesselbach S, et al. International Conference on Machine Learning, Optimization,and Data Science. Cham: Springer Nature Switzerland, 2024: 390-404. [40] Chandran N, Sitaram S, Gupta D, et al. Private benchmarking to prevent contamination and improve comparative evaluation of llms[PP/OL]. arXiv(2024-06-24)[2025-12-24]. https://doi.org/10.48550/arXiv.2403.00393. [41] 郭硕楠, 吴建华. 档案产业联盟构建:动因、模式与策略[J]. 档案学研究, 2023(4):85. [42] Unstructured Data Management Platform. Forget LLM Memory-Why LLMs Need Adaptive Forgetting[EB/OL].(2025-09-16)[2025-10-24]. https://shelf.io/blog/forget-llm-memory-why-llms-need-adaptive-forgetting/. [43] Blanco-Justicia A, Jebreel N, Manzanares-Salor B, et al. Digital forgetting in large language models: a survey of unlearning methods[J]. Artificial IntelligenceReview, 2025(3): 90.