档案信息化2026年第40卷第2期《档案学研究》

档案领域大模型构建的模式、实践及优化建议

Patterns, Practices and Optimization Recommendations for Building Large Language Models in the Archival Domain

刘婧, 吴姗

LIU Jing, WU Shan

1 华中师范大学档案馆 武汉 430079 2 华中师范大学信息管理学院 武汉 430079

出版日期2026-04-28卷期第40卷 第2期页码129-138DOI10.16065/j.cnki.issn1002-1620.2026.02.014

摘要

面对通用大模型在档案领域的应用局限,构建档案领域大模型具有鲜明优势。根据当前行业实践,档案领域大模型构建主要有领域参数内化适配与领域知识外接增强两种模式,两者在构建逻辑、训练过程等方面存在显著差异。目前我国已建有6个具备专业性和多样化服务能力的档案领域大模型。然而,从资源、技术、标准、应用层面分析,当前大模型存在高质量数据获取困难、模型可解释性不足、评估体系滞后、隐私版权风险等问题。据此,本文提出构建高质量与多模态数据集、增强模型可解释性、建立多维评估体系、健全风险防控机制等优化建议,旨在为档案领域大模型的落地应用与数智化转型提供参考。

关键词:档案领域大模型模型微调检索增强生成档案数据集

Abstract

Given the limitations of applying general Large Language Models (LLMs) in the archival field, the development of LLMs tailored to the archival domain has distinct advantages. Based on current practices, the primary approaches to building LLMs in the archival domain include the internalized adaptation mode of domain parameters and the external connection-enhanced mode of domain knowledge,which are significantly different in construction logic and training process. Currently, China has established six LLMs in the archival domain with professional and diversified service capabilities. However, from the perspectives of resources, technology, standards and application, challenges such as difficulties in obtaining high-quality data, insufficient model interpretability, lagging evaluation system, and privacy and copyright risks exist. On these grounds, this paper proposes optimization recommendations including building a high-quality multimodal dataset, enhancing model interpretability, establishing a multi-dimensional evaluation system, and improving the risk prevention and control mechanism, aiming to provide references for the application and digital transformation of LLMs in the archival domain.

Key words: archival LLMs; model fine-tuning; RAG; archive dataset

引用格式

刘婧, 吴姗. 档案领域大模型构建的模式、实践及优化建议[J]. 档案学研究, 2026, 40(2): 129-138.
LIU Jing, WU Shan. Patterns, Practices and Optimization Recommendations for Building Large Language Models in the Archival Domain. Archives Science Study, 2026, 40(2): 129-138.

参考文献

展开查看参考文献(37 条)
[1] SHANAHAN M. Talking about large language models[J]. Communications of the ACM, 2024(2):68-79. [2] 刘璐. 《推进数字档案馆建设实施办法(试行)》解读[J]. 北京档案, 2025(4):44-50. [3] 刘越男, 钱毅, 王平, 等. 挑战与展望:DeepSeek对档案工作的影响及应用前景[J]. 浙江档案, 2025(2):5-13. [4] 牛力, 金持, 黎安润泽. 大模型在档案工作数智转型中的应用:新机遇、新模式和新转变[J]. 档案学通讯, 2024(6):30-38. [5] 刘越男, 张茜雅, 杨建梁. 大语言模型在档案开放审核中的应用框架与路径探究[J]. 档案学通讯, 2025(2):31-38. [6] NGUYEN H D, NGUYEN T H A, NGUYEN T B. A proposed large language model-based smart search for archive system[C]//International Symposium on Information and Communication Technology. Singapore: Springer Nature Singapore, 2024: 210-223. [7] 程媛, 汤舒宁. 生成式AI应用于档案数字叙事的路径研究:技术框架与实施保障[J]. 档案学研究, 2025(4):35-38. [8] 王建品, 尚照辉. 科技档案管理中垂直大模型应用场景、问题与对策[J]. 档案管理, 2025(4):62-66. [9] 刘婧, 欧月. 面向生成式人工智能的档案数智化服务应用场景探索[J]. 档案与建设, 2024(9):83-91. [10] PANDI S P, PARK S, RIKKA P, et al. Can LLMs categorize the specialized documents from web archives in a better way?[C]//Proceedings of the 24th ACM/IEEE Joint Conference on Digital Libraries, 2024:1-11. [11] REUSENS M, ADAMS A, BAESENS B. Large language models to make museum archive collections more accessible[J]. AI & SOCIETY, 2025: 1-13. [12] SUN Y, YANG W, LIU Y. The application of constructing knowledge graph of oral historical archives resources based on LLM-RAG[C]//Proceedings of the 2024 8th International Conference on Information System and Data Mining, 2024: 142-149. [13] 王玉珏, 樊静雅, 侯景瑞, 等. 批判视角下人工智能时代的档案工作:困境反思与职能重构[J]. 档案学研究, 2025(3):4-12. [14] FU Y, SONG J, ZHANG X, et al. Innovative practice of archival data development workflow in the AGI era: a case study of scientist archives project[J]. Information Research an international electronic journal, 2025(1):349-360. [15] 王昊魁. 国家档案局将实施“人工智能+档案”行动[N]. 光明日报,2026-01-16(8). [16] 国家档案局. 2024年度全国档案主管部门和档案馆基本情况摘要(二)[EB/OL].[2026-01-24]. https://www.saac.gov.cn/daj/zhdt/202509/6cac5e5684874efcb454f860a4b94927.shtml. [17] 王云杉. 建设高质量数据集,让人工智能更聪明[N]. 人民日报,2025-05-21(18). [18] 全国信息技术标准化技术委员会(SAC/TC 28). 人工智能大模型第1部分:通用要求:GB/T45288.1—2025[S]. 北京: 中国标准出版社,2025:1. [19] 中国信通院CAICT.北京: 信通院牵头的3项人工智能ITU国际标准正式发布[EB/OL].[2025-09-07]. http://www.cww.net.cn/article?id=600361. [20] 全国信息安全标准化技术委员会(SAC/TC 28). 政务大模型应用安全规范:TC260-004[S]. 北京: 中国标准出版社,2025:3. [21] ZHANG S, PENG S, WANG P, et al. Archives meet GPT: a pilot study on enhancing archival workflows with large language models[J]. iConference 2024 Proceedings, 2024: 1-12. [22] 达观创新性推出大模型档案管理, 全面应用于档案的收、整、管、用[EB/OL].[2025-03-05]. https://mp.weixin.qq.com/s/pICFZMPcIgkUNNZVjrYAEQ. [23] 生成式AI大模型赋能档案管理智慧应用[EB/OL].[2025-03-05]. https://mp.weixin.qq.com/s/gXRWKxJs8mhGMLDFovoQvg. [24] 大模型如何助力档案数智升级—以讯飞星火认知大模型为例[EB/OL].[2025-03-05]. https://mp.weixin.qq.com/s/f6e6CfbJJxQ6bNe3T-kVew. [25] DeepSeek赋能档案未来:“云讯档案大模型”助力政府及各行业高质量发展,共筑数字经济新篇章![EB/OL].[2025-03-05]. https://www.yunwise.cn/new/2/2025-02-27/121.html. [26] 刘刚, 尤健伟. 北京电信赋能首都档案工作智慧化跃升[N]. 人民邮电报,2025-06-17(2). [27] 徐月梅, 胡玲, 赵佳艺, 等. 大语言模型的技术应用前景与风险挑战[J]. 计算机应用, 2024(6):1655-1662. [28] 施敏, 杨海军. 大语言模型数据隐私保护的难点与探索[J]. 大数据, 2024(5):168-176. [29] 容溶, 黄志强. 生成式人工智能在高校档案管理中的应用:伦理挑战与法治监管研究[J]. 法制与经济, 2024(4):54-61. [30] 洪小波. DeepSeek在县区级档案馆的本地化部署与实践—以绍兴市柯桥区档案馆为例[J]. 浙江档案, 2025(3):8-10. [31] TOTH G M, ALBRECHT R, PRUSKI C. Explainable AI, LLM, and digitized archival cultural heritage: a case study of the Grand Ducal Archive of the Medici[J]. AI & SOCIETY, 2025: 1-13. [32] 郑明琪, 陈晓慧, 刘冰, 等. 提示学习中思维链生成和增强方法综述[J]. 计算机科学, 2025(1):56-64. [33] WANG L, XU W, LAN Y, et al. Plan-and-Solve prompting: improving zero-shot chain-of-thought reasoning by large language models[C]//Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics(ACL 2023). 2023: 2609-2634. [34] BOMMASANI R, LIANG P, LEE T. Holistic evaluation of language models[J]. Annals of the New York Academy of Sciences, 2023(1):140-146. [35] 中国信息通信研究院. 大模型基准测试体系研究报告[R/OL].[2025-07-05]. https://www.caict.ac.cn/kxyj/qwfb/ztbg/202407/t20240711_486865.htm. [36] 全国信息技术标准化技术委员会(SAC/TC 28). 人工智能大模型第2部分:评测指标与方法:GB/T45288.2—2025[S]. 北京: 中国标准出版社, 2025:1-20. [37] ZHANG S, HOU J, PENG S, et al. Arcgpt: a large language model tailored for real-world archival applications[EB/OL].[2025-07-05]. https://arxiv.org/abs/2307.14852.