关怀版

档案资源开发2026年第40卷第4期《档案学研究》

基于混合深度学习的满族档案知识重组研究——以《满文老档》汉译本为例

Research on Knowledge Reorganization of Manchu Archives Based on Hybrid Deep Learning: A Case Study of the Chinese Translation of Old Manchu Archives

张文亮, 陈重阳, 姜博文

ZHANG Wenliang, CHEN Chongyang, JIANG Bowen

东北师范大学信息科学与技术学院 长春 130117

出版日期2026-08-28卷期第40卷 第4期页码95-104DOI10.16065/j.cnki.issn1002-1620.2026.04.011通讯作者陈重阳浏览次数0

摘要

针对满族档案知识组织与开发的不足,以《满文老档》为语料来源,结合混合深度学习技术与大语言模型,分别探索《满文老档》满文本知识元识别的可行性、《满文老档》汉译本的知识元识别效果及《满文老档》中复杂语义知识元的抽取。在此基础上,基于知识元抽取结果探索满族档案知识元的重组利用。研究发现:直接开展满文本知识元识别的可行性较差,需要探索面向满文的预训练模型构建;GuwenBERT-BiLSTM-CRF模型在满族档案知识元识别方面效果良好,且在识别职官类和时间类知识元上效果最优;大语言模型在复杂语义知识元识别上效果尚可,但仍需人工进行二次处理以保障质量;满族档案知识图谱能够实现满族档案知识关联、知识检索与知识推理,推动满族档案知识的保护、开发和利用。

关键词:满族档案知识元知识图谱命名实体识别

Abstract

In view of the deficiencies in knowledge organization and development of Manchu archives, the Old Manchu Archives were selected as the corpus source. By integrating hybrid deep learning techniques with large language models, this study explores the the feasibility of identifying knowledge elements in the Manchu text of the Old Manchu Archives, the effectiveness of knowledge element identification in the Chinese translation of the Old Manchu Archives, and the extraction of complex semantic knowledge elements within the Old Manchu Archives. Building upon these findings, the study further investigated the reorganization and utilization of knowledge elements from the Manchu archives based on the results of knowledge element identification.The study found: Directly performing knowledge entity recognition on Manchu text is not feasible, it is necessary to explore the construction of pre-trained models tailored for Manchu text; GuwenBERT-BiLSTM-CRF model has a good effect on knowledge element recognition of Manchu archives, and has the best effect on official position and time categories; Large language model approaches demonstrate acceptable performance in identifying complex semantic knowledge elements, but secondary manual processing remains necessary to obtain high-quality knowledge elements; The knowledge graph of Manchu archives can realize the knowledge association, knowledge retrieval and knowledge reasoning of Manchu archives, advancing the preservation, development, and utilization of the knowledge of Manchu archives.

Key words: Manchu archives; knowledge elements; knowledge graph; named entity recognition

引用格式

张文亮, 陈重阳, 姜博文. 基于混合深度学习的满族档案知识重组研究——以《满文老档》汉译本为例[J]. 档案学研究, 2026, 40(4): 95-104.
ZHANG Wenliang, CHEN Chongyang, JIANG Bowen. Research on Knowledge Reorganization of Manchu Archives Based on Hybrid Deep Learning: A Case Study of the Chinese Translation of Old Manchu Archives. Archives Science Study, 2026, 40(4): 95-104.

参考文献

展开查看参考文献(25 条)
[1] 习近平. 习近平谈治国理政:第一卷[M]. 北京: 外文出版社,2018:161. [2] 关于推进新时代古籍工作的意见[N]. 人民日报,2022-04-12(1). [3] 吴元丰. 满文与满文古籍文献综述[J]. 满族研究, 2008(1):99-113,128. [4] 黄晓捷, 熊回香, 肖兵, 等. 基于“问题—方法”知识元挖掘的学科知识流动研究[J]. 图书情报工作, 2024(8):80-96. [5] 孙成江, 吴正荆. 知识服务战略:创建增值联盟[J]. 情报科学, 2002(10):1028-1029,1035. [6] 李刚. 从传统到现代:中国第一历史档案馆满文档案编纂工作的回顾与展望[J]. 中国档案, 2022(8):40-41. [7] 李健民. 中国第一历史档案馆满文档案全文数据库建设与应用[J]. 历史档案, 2025(1):140-144. [8] 孙文杰. 从满文寄信档看“乌什事变”真相[J]. 云南民族大学学报(哲学社会科学版), 2016(6):128-135. [9] 赵寰熹. 清代满文文献中地理环境用语及其地理观呈现[J]. 满语研究, 2020(1):84-90. [10] 刘金德, 刘荣. 满文文献的保护和利用[J]. 兰台世界, 2010(5):30-31. [11] 郑蕊蕊, 辛守宇, 周瑜, 等. 一种基于N元ECOC的大类别K-shot满文识别方法[J]. 郑州大学学报(理学版), 2021(4):53-60. [12] 孙凯明, 孙磊, 王刚, 等. 面向满文档案图像的手写体满文智能识别软件设计与实现[J]. 自动化技术与应用, 2024(1):91-94. [13] 索传军. 知识转移视角下的学术论文老化与创新研究[J]. 图书情报工作, 2014(5):5-12. [14] 温有奎, 焦玉英. 知识元语义链接模型研究[J]. 图书情报工作, 2010(12):27-31. [15] 索传军, 戎军涛. 知识元理论研究述评[J]. 图书情报工作, 2021(11):133-142. [16] 刘佳, 边俊伊. 基于混合深度学习的藏医古籍命名实体识别研究[J]. 现代情报, 2023(11):37-46. [17] 熊回香, 陈子薇, 肖兵. 数据叙事视角下红色档案资源语义组织研究[J]. 档案学研究, 2025(2):92-101. [18] 张文亮, 丁淼, 魏来. 满文古籍知识库研究与构建[J]. 图书情报工作, 2025(5):126-139. [19] 华林, 冯安仪, 吴皎钰. 基于智慧数据的海疆历史档案知识聚合与服务研究[J]. 档案学研究, 2025(4):59-69. [20] 佚名. 满文老档[M]. 中国第一历史档案馆, 中国社会科学院历史研究所, 译注. 北京: 中华书局,1990. [21] 马诗语, 黄润才. 基于ALBERT与BILSTM的糖尿病命名实体识别[J]. 中国医学物理学杂志, 2021(11):1438-1443. [22] 余丹丹, 黄洁, 党同心, 等. 基于ALBERT的中文简历命名实体识别[J]. 计算机工程与设计, 2024(1):261-267. [23] 舒忠梅, 张萍, 李小霞, 等. 数字记忆视角下重大社会事件档案知识图谱开发—以中山大学抗疫专题档案为例[J]. 浙江档案, 2021(4):46-48. [24] 贾君枝. 资源描述中的词表重用类型与实现方式[J]. 中国图书馆学报, 2021(4):76-85. [25] 赵雪芹, 李天娥, 曾刚. 基于Neo4j的万里茶道数字资源知识图谱构建研究[J]. 情报资料工作, 2022(5):89-97.