展开查看参考文献(43 条)
[1] [7]王毅, 陈文汇, 刘红霞. 基于AI大模型的档案信息资源开发:逻辑机理、应用赋能与推进策略[J]. 山西档案, 2025(8):2-3,1.
[2] 李文锋, 刘峰. DeepSeek在工程建设项目档案真实性检查中的应用思考[J]. 陕西档案, 2025(5):24.
[3] 中国网. 全国首个辅助决策地方政务大模型落地北京电信助力首都政务档案智慧化跃升[EB/OL].(2025-06-16)[2025-09-25]. https://tech.chinadaily.com.cn/a/202506/16/WS684fb3a4a3102053770383de.html.
[4] 刘倩倩, 刘圣婴, 刘炜. 图书情报领域大模型的应用模式和数据治理[J]. 图书馆杂志, 2023 (12):22.
[5] 郭之璇, 苏宇. 领域模型的算法透明义务[J]. 人工智能, 2025(4):92.
[6] 王建品, 尚照辉. 科技档案管理中垂直大模型应用场景、问题与对策[J]. 档案管理, 2025(4):63.
[7] 吴海中. 档案部门应用大语言模型的典型案例分析[J]. 山西档案, 2025(3):152.
[8] 王昊贤, 周子茗, 丁菲菲, 等. 数字人文与大语言模型:古文献语义检索实践与探索[J]. 农业图书情报学报, 2024(9):92.
[9] 张中泽. 高质量推进红色档案资源保护管理和开发利用的实践与思考[J]. 山西档案, 2025(11):42.
[10] Online Computer Library Center. Recommendations of the OCLC Research Library Partnership Web Archiving Metadata Working Group[EB/OL].(2018-02)[2025-10-08]. https://www.oclc.org/research/publications/2018/oclcresearch-descriptive-metadata/recommendations.html.
[11] 张军荣. 人工智能生成物的溢出风险及监管规则研究[J]. 行政法学研究, 2025(2):131-132.
[12] 刘晓春. 数据功能类型视角下数据污染的治理维度[J]. 人民司法, 2024(10):16.
[13] 曾庆醒. 涌现效应下的生成式人工智能数据污染及其治理路径[J]. 法学杂志, 2024(5):53.
[14] 陈俊秀, 徐玉琴. 人工智能大语言模型引发的数据污染风险及其规制路径[J]. 大连理工大学学报(社会科学版), 2025(4):55.
[15] 刘福元. 电竞人工智能的数据污染风险与治理规则建构[J]. 首都体育学院学报, 2025(2):126-128.
[16] Sainz O, Campos J, García-Ferrero I, et al. NLP evaluation in trouble: on the need to measure llm data contamination for each benchmark[C]// Bouamor H, Pino J, Bali K. Findings of the Association for Computational Linguistics: EMNLP 2023. Singapore:Association for Computational Linguistics, 2023: 10776-10787.
[17] Pan Y, Pan L, Chen W, et al. On the risk of misinformation pollution with large language models[C]// Bouamor H, Pino J, Bali K. Findings of the association for computational linguistics: EMNLP 2023. Singapore:Association for Computational Linguistics, 2023: 1389-1403.
[18] 蒋冠, 刘子焱. 大模型技术与档案编研业务深度融合的障碍与应对策略[J]. 山西档案, 2025(9):3.
[19] [29]Lee K, Ippolito D, Nystrom A, et al. Deduplicating training data makes language models better[C]//Muresan S, Nakov P, Villavicencio A. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics. Dublin, Ireland: Association for Computational Linguistics, 2022: 8424-8445.
[20] Roberts M, Thakur H, Herlihy C, et al. Data contamination through the lens of time[PP/OL]. arXiv(2023-10-16)[2025-12-10]. https://doi.org/10.48550/arXiv.2310.10628.
[21] 安小米, 龙志奇, 邝苗苗. 标准化视角下大模型数据治理的理论框架及其构成要素研究[J]. 情报资料工作, 2024(6):79.
[22] 裴宏娇, 郭瑞宇, 杨志和. 图书馆AI语言生成中的信息幻觉:信息素养视角下的风险规避策略[J]. 图书情报导刊, 2024(12):37.
[23] Wiggers K. MIT study finds "systematic" labeling errors in popular AI benchmark datasets[EB/OL].(2021-03-29)[2025-11-16]. https://venturebeat.com/ai/mit-study-finds-systematic-labeling-errors-in-popular-ai-benchmark-datasets.
[24] Suresh H, Guttag J. A framework for understanding sources of harm throughout the machine learning life cycle[C]// Proceedings of the 1st ACM Conference on Equity and Access in Algorithms, Mechanisms, and Optimization. New York: Association for Computing Machinery, 2021: 1-9.
[25] Offenhuber D. Shapes and frictions of synthetic data[J/OL]. Big Data & Society, 2024(2). https://doi.org/10.1177/20539517241249390.
[26] Giuffrè M, Shung D L. Harnessing the power of synthetic data in healthcare: innovation, application, and privacy[J]. NPJ digital medicine, 2023(1): 186.
[27] 董敏. 论县级综合档案馆重复档案的精简工作[J]. 山东档案, 2020(6):51.
[28] Jiang M, Liu K Z, Zhong M, et al. Investigating data contamination for pre-training language models[PP/OL]. arXiv(2024-01-11)[2025-11-15]. https://doi.org/10.48550/arXiv.2401.06059.
[29] Brown T, Mann B, Ryder N, et al. Language models are few-shot learners[J]. Advances in neural information processing systems, 2020(33): 1877-1901.
[30] OpenAI. GPT-4 Technical Report[EB/OL].(2023-03-27)[2025-11-28]. https://cdn.openai.com/papers/gpt-4.pdf.
[31] 赵铁军, 许木璠, 陈安东. 自然语言处理研究综述[J]. 新疆师范大学学报(哲学社会科学版), 2025(2):89-90.
[32] Bender E M, Gebru T, McMillan-Major A, et al. On the dangers of stochastic parrots: can language models be too big?[C]// Proceedings of the 2021 ACM conference on fairness, accountability, and transparency. New York: Association for Computing Machinery, 2021: 610-623.
[33] Magar I, Schwartz R. Data contamination: from memorization to exploitation[C]// Muresan S, Nakov P, Villavicencio A.Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics. Dublin, Ireland: Association for Computational Linguistics, 2022: 157-165.
[34] Jo E S, Gebru T. Lessons from archives: strategies for collecting sociocultural data in machine learning[C]// Proceedings of the 2020 conference on fairness, accountability, and transparency. New York: Association for Computing Machinery, 2020: 306-316.
[35] 刘力超, 陈晓珑, 牛力. 大模型驱动档案开放智能审核方法研究:动因、框架与实践[J]. 档案学研究, 2025(2):136.
[36] 胡泳, 张文杰. 数据标注治理:可信人工智能的后台风险与治理转向[J]. 云南社会科学, 2024(6):34-35.
[37] Jordon J, Szpruch L, Houssiau F, et al. Synthetic data—what,why and how?[PP/OL].[arXiv(2022-05-06)[2025-12-14]. https://doi.org/10.48550/arXiv.2205.03257.
[38] Shilov I, Meeus M, de Montjoye Y A. The mosaic memory of large language models[J]. Nature Communications, 2026(1): 2142.
[39] Leveling J, Helmer L, Stein B J, et al. Evaluation of Document Deduplication Algorithms for Large Text Corpora[C]//Nicosia G, Ojha V, Giesselbach S, et al. International Conference on Machine Learning, Optimization,and Data Science. Cham: Springer Nature Switzerland, 2024: 390-404.
[40] Chandran N, Sitaram S, Gupta D, et al. Private benchmarking to prevent contamination and improve comparative evaluation of llms[PP/OL]. arXiv(2024-06-24)[2025-12-24]. https://doi.org/10.48550/arXiv.2403.00393.
[41] 郭硕楠, 吴建华. 档案产业联盟构建:动因、模式与策略[J]. 档案学研究, 2023(4):85.
[42] Unstructured Data Management Platform. Forget LLM Memory-Why LLMs Need Adaptive Forgetting[EB/OL].(2025-09-16)[2025-10-24]. https://shelf.io/blog/forget-llm-memory-why-llms-need-adaptive-forgetting/.
[43] Blanco-Justicia A, Jebreel N, Manzanares-Salor B, et al. Digital forgetting in large language models: a survey of unlearning methods[J]. Artificial IntelligenceReview, 2025(3): 90.