| 摘 要: 随着人工智能的发展,大语言模型在自然语言处理领域展现出了卓越的性能,也为机器翻译带来了新的思路。鉴于通用大模型的专业局限性问题,混合领域语料和通用语料对Qwen2-7B大模型进行领域自适应训练,并针对古文到现代文翻译的下游任务进行指令微调。最终,从领域性能和通用性能的角度对模型性能进行了评估。对于大模型的领域性能,从译文的准确性、流畅性和文学性3个维度进行了评估,并针对文学性提出了一套新的评估标准。实验结果表明,结合领域自适应与指令微调的双阶段策略,能够显著提升大模型古文翻译的质量。 |
| 关键词: 大语言模型 机器翻译 领域自适应 指令微调 文学性 |
|
中图分类号: TP391
文献标识码: A
|
| 基金项目: 国家自然科学基金项目资助(61977039) |
|
| Researchon Machine Translation of Ancient Chinese Texts Basedon Large Language Model |
|
TIE Lexin1 , GU Yiran1 , HUANG Liya2
|
(1.College of Automation, Nanjing University of Posts and Telecommunications, Nanjing 210023, China; 2. College of Electronic and Optical Engineering & College of Flexible Electronics (Future Technology), Nanjing University of Posts and Telecommunications, Nanjing 210023, China)
15851888192@163.com; guyr@njupt.edu.cn; huangly@njupt.edu.cn
|
| Abstract: With the development of artificial intelligence, Large Language Models have demonstrated remarkable performance in the field of natural language processing, bringing new insights to machine translation. Given the professional limitations of genera-l purpose large models, this paper conducts domain adaptation training on the Qwen2- 7B large model by mixing domain-specific and general corpora, and performs instruction fine-tuning for the downstream task of translating classical Chinese to modern Chinese. Finally, the model’s performance is evaluated from both domain-specific and general capabilities perspectives. For the domain-specific performance, the evaluation is carried out from three dimensions: accuracy, fluency, and literary quality of the translations with a new set of evaluation criteria proposed specifically for literary quality. Experimental results show that the two-stage strategy combining domain adaptation and instruction fine-tuning can significantly improve the quality of classical Chinese translation by large models. |
| Keywords: Large Language Model machine translation domain adaptation instruction fine-tuning literary quality |