Comparing Chinese–English Translation Performance of Large Language Models: A Multidimensional Evaluation of Financial, Technological, and Political Texts
DOI:
https://doi.org/10.63313/JCSFT.9087Keywords:
Large Language Models, Professional Text Translation, Machine Translation Quality Evaluation, Multidimensional Evaluation, Prompt StrategyAbstract
Large language models (LLMs) are increasingly used for machine translation; however, their reliability in professional Chinese–English translation, particularly for domain-specific texts, remains insufficiently examined. This challenge is especially evident in texts containing specialized terminology, numerical information, complex logical relations, and formal institutional phrasing. To address this issue, this study conducts a controlled comparison of six LLMs: ChatGPT-5.0, Claude Sonnet 4.5, DeepSeek-R1, GLM-4.6, Qwen3-MAX, and Grok 4.0. Three representative Chinese texts from the financial, technological, and political domains were selected and translated under three prompt conditions: zero-shot prompting, domain-enhanced prompting, and glossary-based prompting.
Translation quality was assessed through a multidimensional evaluation framework. Standard reference-based metrics, including BLEU, ChrF++, TER, ROUGE, BERTScore, and COMET, were used to measure automatic translation quality. In addition, lexical richness and syntactic-discourse indicators were incorporated to examine stylistic and structural differences across model outputs.
The results reveal a clear domain effect, indicating that financial, technological, and political texts present different levels and types of translation difficulty for LLMs. By contrast, prompt strategy did not produce significant main effects under the three tested conditions. At the model level, GLM-4.6 achieved the highest COMET-based ranking across all prompt settings, while DeepSeek-R1 showed comparatively stable second-place performance.
Overall, the findings suggest that, within this controlled exploratory dataset, model type and text domain exert a stronger influence on translation performance than prompt modification alone. Rather than offering a definitive ranking of LLM translation capability, this study provides empirical evidence on how selected LLMs handle specialized Chinese–English translation under comparable experimental conditions. Future research should expand the dataset, incorporate human evaluation, and further investigate the role of terminology constraints in improving translation accuracy and stylistic appropriateness.
References
[1] Chan, S.W. (2004) Machine Translation. In: Baker, M. and Saldanha, G., Eds., Routledge Encyclopedia of Translation Studies, Routledge, London, 135.
[2] Fu, Y., Si, S., Mai, L. and Li, X. (2024) FFN: A Fine-Grained Chinese-English Financial Domain Parallel Corpus. arXiv Preprint, arXiv:2406.18856.
[3] Zhao, S., Qiao, L., Luo, K., Zhang, Q.-W., Lu, J. and Yin, D. (2024) SNFinLLM: Systematic and Nuanced Financial Domain Adaptation of Chinese Large Language Models. arXiv Preprint, arXiv:2408.02302.
[4] Kwok, H.L., Shi, Y., Xu, H., Li, D. and Liu, K. (2025) GenAI as a Translation Assistant? A Corpus-Based Study on Lexical and Syntactic Complexity of GPT-Post-Edited Learner Translation. System, 130, Article ID: 103618.
[5] Rei, R., Stewart, C., Farinha, A.C. and Lavie, A. (2020) COMET: A Neural Framework for MT Evaluation. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, Online, 16-20 November 2020, 2685-2702. https://doi.org/10.18653/v1/2020.emnlp-main.213
[6] Zheng, J. (2021) The Importance of a Glossary for Translation Services. Globalization Partners.
[7] Semenov, K., Ailem, M., Jotti, A., España-Bonet, C., Semmar, N., Wanner, L. and Wisniewski, G. (2023) Findings of the WMT 2023 Terminology Shared Task on Machine Translation with Terminologies. Proceedings of the Eighth Conference on Machine Translation.
[8] Gu, W.H. and Leng, B.B. (2024) Four Types of Terminology Mistranslation in the Application of ChatGPT to Scientific and Technological Translation: A Case Study of Mechanical Engineering Terminology. Chinese Science & Technology Translators Journal, 37, 1.
[9] Zheng, X. (2025) A Comparative Study of Translation Quality of Scientific and Technological Texts under Large Language Models: A Case Study of Four Typical Texts. Modern Linguistics.
[10] Yu, L. (2024) A Study on Lexical Diversity and Syntactic Complexity in ChatGPT Translation. Foreign Language Teaching and Research, 56, 2.
[11] Yu, J. (2025) Exploring DeepSeek’s Translation Capability: A Case Study of Literary and Financial Texts. Chinese Translators Journal, 2025, 172-179.
[12] Wen, X. and Tian, Y.L. (2024) The Effectiveness of ChatGPT in Translating Discourse with Chinese Characteristics. Shanghai Journal of Translators, 2024, 27-34.
[13] Zhou, Z.L. (2024) Generative-AI-Based Maritime Translation: Advantages, Challenges and Prospects. Research Report.
[14] Zhao, Y., Zhang, H. and Yang, Y.C. (2024) A Comparative Study of Large Language Models in Text Translation Quality: A Case Study of the Translation of Blossoms. Technology Enhanced Foreign Language Education, 2024, 60-66.
[15] Feng, Q.H. (2025) Innovative Applications of DeepSeek in Translation Teaching and Research. Chinese Translators Journal, 2025, 58-67.
[16] Zhang, S.K. and Zhao, C.Y. (2024) The Applicability of Large Language Models to Literary Translation: A Multi-Metric Evaluation of Multi-Model Translations of Border Town. Research Report.
[17] Liu, S.J. (2024) The Application Effectiveness of Machine Translation in Maritime Translation: A Comprehensive Evaluation Based on BLEU, chrF++ and BERTScore. Research Report.
Downloads
Published
Issue
Section
License
Copyright (c) 2026 by author(s) and Erytis Publishing Limited

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.













